Designing a Multi-Team Kafka Platform
A multi-team Kafka platform must provide five capabilities: self-service topics, schema governance, quotas, security, and observability, plus disaster recovery. Self-service topics mean teams can create and manage their own topics without manual intervention from the platform team. This is typically implemented with a self-service portal or API that enforces naming conventions, ownership, and default configurations. Schema governance means every topic has a schema, registered in a central schema registry, with compatibility policies enforced. Quotas mean each team has limits on produce and consume rates, so one team cannot starve others. Security means authentication (SASL or mTLS) and authorization (ACLs) are enforced, with each team having its own principals and permissions. Observability means metrics, logs, and traces are collected and exposed to teams, with dashboards and alerts for their topics and consumer groups. Disaster recovery means a standby cluster in another region, replicated with MirrorMaker 2, with a tested failover procedure. The trade-off is between central control and team autonomy. The platform team provides the guardrails; the teams operate within them. A platform that is too restrictive slows teams down; a platform that is too permissive is insecure and unreliable.
The mechanism for self-service is a control plane: a service that teams use to request topics. The service validates the request against naming conventions and ownership rules, creates the topic with default configurations (partitions, replication factor, retention, cleanup policy), registers the schema, and sets up ACLs and quotas. The service should be idempotent and auditable. The mechanism for schema governance is a schema registry integrated with the control plane: when a topic is created, a schema subject is created; when a team wants to change the schema, the registry checks compatibility against the policy and rejects breaking changes. The mechanism for quotas is per-principal and per-client-id quotas, configured via the control plane and enforced by the brokers. The mechanism for security is a combination of authentication (SASL/SCRAM, mTLS, or OAuth) and authorization (ACLs). Each team has a principal, and the control plane grants ACLs based on the team's topics and groups. The mechanism for observability is a metrics pipeline that collects broker, topic, and consumer group metrics and exposes them to teams via dashboards. The mechanism for DR is a standby cluster with MirrorMaker 2, monitored for replication lag, with a documented and tested failover procedure. Version note: Confluent Platform and Confluent Cloud provide many of these capabilities out of the box, including RBAC, self-service, and multi-region replication. If you are on open-source Kafka, you build more of this yourself. KRaft improves metadata scalability, which is important for large multi-tenant clusters. Tiered storage can reduce storage costs for long-retention topics.
A common mistake is to give teams too much freedom, which leads to inconsistent configurations, security gaps, and operational surprises. Another mistake is to make the platform too restrictive, which slows teams down and encourages them to bypass the platform. A third mistake is to ignore the operational cost of thousands of topics and consumer groups; the platform must be designed for scale. The trade-off is between centralization and autonomy. A good platform provides sensible defaults, clear guardrails, and self-service within those guardrails. It also provides observability so teams can diagnose their own issues, and a support model for when they cannot. The platform team's role shifts from doing everything to enabling teams to do things safely. Version note: the exact tooling depends on the environment. In a cloud environment, managed Kafka services provide much of this. On-premises, you build more. In all cases, the principles are the same: self-service, governance, quotas, security, observability, and DR. The platform should be treated as a product with its own roadmap and SLOs.
Six capabilities: self-service topics, schema governance, quotas, security, observability, and DR.
Self-service via a control plane that enforces naming, ownership, and defaults.
Schema governance via a central registry with per-subject compatibility.
Quotas per team to prevent resource starvation.
Security via authentication (SASL/mTLS) and authorization (ACLs).
Observability via metrics, logs, and traces exposed to teams.
DR via a standby cluster with MirrorMaker 2 and a tested failover procedure.
Balance central control with team autonomy; provide guardrails, not gates.
0-2 years experience
2-5 years experience
5-8 years experience