Define one authorization model as code: strong identities, prefixed least-privilege ACLs generated from a central policy, applied through CI, with drift detection and automated tests
Inconsistent authorization is usually a process problem, not a missing feature: ACLs were created by hand, per team, with different conventions and a few wildcards added during incidents. So I redesign in layers. Identity first: every client authenticates with mTLS or SASL (OAUTHBEARER or SCRAM) and maps to a stable, service-level principal, never a shared or personal one. Authorization second: enable the broker authorizer (StandardAuthorizer in KRaft, AclAuthorizer in ZooKeeper-era clusters), set allow.everyone.if.no.acl.found=false so the default is deny, and keep super.users to a tiny platform set. Policy third: define naming conventions that carry ownership, for example team.domain.topic, and grant access with prefixed ACLs on that namespace so a team gets rights over its own prefix and explicit, reviewed grants for anything else. Consumers need Read on the topic and Read on their group; transactional producers need the transactional ID resource; admin operations need explicit cluster or topic permissions.
Then I make the policy the source of truth. Teams declare what they own and what they need in a Git repository (a topic-and-access manifest), a pull request triggers validation (naming, no wildcards, no cluster-level grants without platform approval), and a pipeline applies the result to the cluster using the Admin API or a tool such as Terraform, Strimzi KafkaUser resources or a GitOps tool for Kafka. A scheduled job compares live ACLs to the declared state and alerts on drift, so manual changes are caught. I roll this out without outages: first enable authorizer logging to learn who actually accesses what, generate the policy from observed traffic, review it with owners, apply allow rules alongside the old state, then remove wildcard and orphaned ACLs in stages. Testing is part of the design: automated checks that each service principal can do what it needs and cannot do what it should not, run in CI against an ephemeral cluster and as periodic production audits. Audit logging, secret rotation and certificate expiry alerts complete the loop.
Trade-off: prefixed ACLs make onboarding cheap and consistent but depend on strict naming discipline. Per-topic literal ACLs are precise but do not scale operationally. I use prefixes for ownership and literals for cross-team grants.
Trade-off: central policy review improves safety but can become a bottleneck. Automated validation for routine requests and human review only for exceptions (cross-team, cluster-level, wildcard) keeps teams moving.
Alternative: an external authorizer (OPA, an IAM-backed plugin, or a vendor RBAC layer) gives role-based and attribute-based rules and central audit, at the price of another critical dependency in the broker authorization path. I choose it when ACL volume or compliance needs outgrow native ACLs.
Common mistake: fixing findings by adding wildcard ACLs or User:* to unblock a team. Wildcards defeat the whole review, so the linter should reject them.
Common mistake: forgetting the other resources: consumer groups, transactional IDs, delegation tokens, and cluster-level operations such as creating topics or altering configs. Topic ACLs alone leave gaps.
Common mistake: sharing one principal across many services. You lose attribution in audit logs and cannot revoke one service without breaking the others.
Common mistake: enabling default deny in one step on a live cluster. Observe with authorizer logs, generate the intended policy, and migrate in stages with a rollback path.
Verification: positive and negative access tests per service principal, run in CI and on a schedule, plus drift detection comparing live ACLs to Git.
Version note: StandardAuthorizer in KRaft replaces ZooKeeper-based ACL storage, and ZooKeeper mode is gone in Kafka 4.0. Prefixed ACLs exist since 2.0, but Terraform and operator support, and OAuth or RBAC features, depend on your distribution, so verify what your platform provides.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience