A tiered, governed Kafka platform where workload fit, not enthusiasm, decides what runs on it
I start with requirements before technology: which workloads need replay, fan-out to many consumers, high throughput, or ordering per entity, what the latency and durability expectations are, and what compliance applies. Then I design the platform as a product with a clear contract for teams. The core is a multi-broker KRaft cluster per environment, rack or zone aware with replication factor 3 and min.insync.replicas=2, so an AZ loss does not lose acknowledged data. Around it sit the pieces that make it enterprise-ready: a schema registry with enforced compatibility, Kafka Connect for CDC and sink integration, a stream-processing layer where needed, authN/authZ (mTLS or SASL with least-privilege ACLs), quotas, observability and a self-service topic provisioning workflow.
I would not run everything on one undifferentiated cluster. I define workload tiers with different configs: a critical tier for financial and order events (acks=all, min.insync.replicas=2, unclean leader election disabled, long retention, multi-zone), a high-volume tier for telemetry and clickstream (shorter retention, compression such as zstd, cheaper storage and possibly tiered storage), and a sandbox tier. For multi-region, I would use MirrorMaker 2 or a managed replication offering, and I would be explicit about the RPO and RTO, because cross-cluster replication is asynchronous and offsets are not automatically identical.
Should use Kafka: high-throughput event streams, CDC pipelines, audit and activity logs, data shared by many consumers, stream processing, anything needing replay or rebuildable read models.
Should not use Kafka: synchronous request/response, small low-volume integrations, job queues needing priority or delayed delivery, ad hoc querying, and large binary payloads. Use HTTP/gRPC, a managed queue, a database or object storage.
Governance: naming conventions, topic ownership and data classification, schema compatibility rules (for example backward-compatible by default), PII handling and retention aligned to regulation, and a review path for new topics.
Cost: storage and cross-AZ network traffic dominate. Levers are retention tuning, compression, tiered storage (production-ready in recent releases, but check your version and limits), right-sized partition counts and chargeback by team.
Failure semantics: define delivery guarantees per tier. At-least-once with idempotent consumers is the default, transactions for read-process-write pipelines, a DLQ or retry-topic standard for poison messages, and runbooks for broker loss, under-replicated partitions and consumer lag.
Operating model: managed service versus self-run. Self-managing gives control but needs a platform team, while a managed service trades some cost and control for lower toil. Decide based on team capacity, not preference.
Common mistake: over-partitioning up front. Partitions cannot be reduced, increase broker load and rebalance time, and changing the count breaks key-to-partition mapping for existing data.
Common mistake: one shared topic per company with free-form payloads. Without ownership, contracts and quotas, a noisy tenant degrades everyone.
The trade-off I would state out loud: a central platform gives consistency and economies of scale but can become a bottleneck for teams, so I favor self-service provisioning within guardrails. I would also name what I would validate before committing: load-test realistic peak throughput and failure scenarios, since latency, durability and cost pull against each other. For example, acks=all with cross-zone replication adds latency, while acks=1 is faster but can lose acknowledged data on leader failure.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience