04 / 05

Design a multi-tenant Kafka SaaS platform with isolation, quotas, schemas, audit logging, DR and self-service onboarding.

Difficulty: 10/10
Multi-tenant platform

A tiered platform: tenant identity and namespace isolation, enforced quotas and limits, per-tenant schema governance, layered audit logs, explicit DR tiers, and GitOps self-service on a control plane

I start by choosing the isolation model per tenant tier, because it drives everything else. Most tenants share a cluster with logical isolation: a tenant-specific topic prefix, a dedicated principal, prefixed ACLs, and quotas. Large, regulated or noisy tenants get a dedicated cluster (or dedicated brokers) so that failures, upgrades and capacity are independent. The platform has two halves. The data plane is the Kafka clusters, built with rack-aware replication factor 3, min.insync.replicas=2, TLS everywhere and KRaft controllers. The control plane is a self-service system, where a tenant declares what it needs (topics, partitions, retention, schemas, access) in a spec; automated policy checks validate it against the tenant's tier and limits; and a reconciler applies it to Kafka, the schema registry and the identity system. Self-service is the product: if onboarding takes a ticket queue, tenants will bypass the platform.

Isolation has to be enforced, not documented. Security: per-tenant authenticated principal and prefixed ACLs, default deny, no cross-tenant grants without an explicit shared-topic agreement. Resource isolation: produce and fetch byte-rate quotas per tenant principal, request-percentage quotas to protect broker threads, controller-mutation quotas to stop topic-creation storms, plus limits on partitions, topics and retention enforced by a create-topic policy and the control plane, since partitions are the real cost driver. Schemas: per-tenant subject namespaces (or registry contexts) with compatibility mode enforced and schema changes gated in the tenant's CI. Audit logging in layers: authorizer logs for allow and deny decisions, control-plane audit of who changed what and when, and data-access logs for sensitive streams, shipped to tamper-evident storage with retention that meets compliance. DR is a tiered product with honest numbers: best-effort tenants get backups of configuration only; standard tenants get asynchronous cross-cluster replication (MirrorMaker 2 or a managed equivalent) with an RPO above zero and tested offset translation; premium tenants may get a stretched cluster or synchronous options with the latency and quorum costs stated. Finally, capacity and cost: per-tenant metering of bytes in and out, storage and partitions drives chargeback, forecasting and the decision to promote a tenant to a dedicated cluster.

javascript
  1. 1

    Trade-off: shared clusters are cheap and fast to onboard but share failure domains, upgrades and noisy-neighbor risk. Dedicated clusters give hard isolation at much higher cost and operational load. Offer both and promote tenants by measured load and compliance needs.

  2. 2

    Trade-off: topic-per-tenant isolation is clear and simple to meter, but thousands of tenants times many topics creates partition and metadata overhead. For the long tail, consider shared topics keyed by tenant with strict governance, accepting weaker isolation.

  3. 3

    Trade-off: quotas protect everyone but throttle a tenant that bursts. Provide documented limits, burst headroom per tier, and alerting before throttling hurts.

  4. 4

    Common mistake: relying on naming conventions without ACL and quota enforcement, which gives the appearance of isolation but no protection against a misbehaving tenant.

  5. 5

    Common mistake: forgetting that quotas by client-id are spoofable. Use authenticated principals (user quotas) so a tenant cannot dodge limits by changing a string.

  6. 6

    Common mistake: ignoring the noisy resources that are not bytes: partitions, connections, request rate, consumer groups and topic-creation rate. Limit those too.

  7. 7

    Common mistake: promising zero data loss on DR for all tenants. Asynchronous replication has an RPO above zero, and consumer offsets, schemas and ACLs must all be replicated and rehearsed for failover to work.

  8. 8

    Common mistake: building self-service as a ticket form. The spec must be machine-validated and auto-applied, with humans reviewing only exceptions.

  9. 9

    Operational requirements: per-tenant dashboards and alerts, a documented offboarding process with data deletion, tenant-aware upgrade and maintenance windows, and regular DR and noisy-neighbor game days.

  10. 10

    Version note: KRaft is the only mode from Kafka 4.0, and controller-mutation quotas, KIP-848 group protocol and tiered storage maturity depend on your release. The schema registry, RBAC and cluster-linking features come from your distribution or vendor, so state those assumptions.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.