05 / 05

Design an enterprise event-streaming platform and justify which workloads should and should not use Kafka.

Difficulty: 10/10
Core concepts

A tiered, governed Kafka platform where workload fit, not enthusiasm, decides what runs on it

I start with requirements before technology: which workloads need replay, fan-out to many consumers, high throughput, or ordering per entity, what the latency and durability expectations are, and what compliance applies. Then I design the platform as a product with a clear contract for teams. The core is a multi-broker KRaft cluster per environment, rack or zone aware with replication factor 3 and min.insync.replicas=2, so an AZ loss does not lose acknowledged data. Around it sit the pieces that make it enterprise-ready: a schema registry with enforced compatibility, Kafka Connect for CDC and sink integration, a stream-processing layer where needed, authN/authZ (mTLS or SASL with least-privilege ACLs), quotas, observability and a self-service topic provisioning workflow.

I would not run everything on one undifferentiated cluster. I define workload tiers with different configs: a critical tier for financial and order events (acks=all, min.insync.replicas=2, unclean leader election disabled, long retention, multi-zone), a high-volume tier for telemetry and clickstream (shorter retention, compression such as zstd, cheaper storage and possibly tiered storage), and a sandbox tier. For multi-region, I would use MirrorMaker 2 or a managed replication offering, and I would be explicit about the RPO and RTO, because cross-cluster replication is asynchronous and offsets are not automatically identical.

javascript
  1. 1

    Should use Kafka: high-throughput event streams, CDC pipelines, audit and activity logs, data shared by many consumers, stream processing, anything needing replay or rebuildable read models.

  2. 2

    Should not use Kafka: synchronous request/response, small low-volume integrations, job queues needing priority or delayed delivery, ad hoc querying, and large binary payloads. Use HTTP/gRPC, a managed queue, a database or object storage.

  3. 3

    Governance: naming conventions, topic ownership and data classification, schema compatibility rules (for example backward-compatible by default), PII handling and retention aligned to regulation, and a review path for new topics.

  4. 4

    Cost: storage and cross-AZ network traffic dominate. Levers are retention tuning, compression, tiered storage (production-ready in recent releases, but check your version and limits), right-sized partition counts and chargeback by team.

  5. 5

    Failure semantics: define delivery guarantees per tier. At-least-once with idempotent consumers is the default, transactions for read-process-write pipelines, a DLQ or retry-topic standard for poison messages, and runbooks for broker loss, under-replicated partitions and consumer lag.

  6. 6

    Operating model: managed service versus self-run. Self-managing gives control but needs a platform team, while a managed service trades some cost and control for lower toil. Decide based on team capacity, not preference.

  7. 7

    Common mistake: over-partitioning up front. Partitions cannot be reduced, increase broker load and rebalance time, and changing the count breaks key-to-partition mapping for existing data.

  8. 8

    Common mistake: one shared topic per company with free-form payloads. Without ownership, contracts and quotas, a noisy tenant degrades everyone.

The trade-off I would state out loud: a central platform gives consistency and economies of scale but can become a bottleneck for teams, so I favor self-service provisioning within guardrails. I would also name what I would validate before committing: load-test realistic peak throughput and failure scenarios, since latency, durability and cost pull against each other. For example, acks=all with cross-zone replication adds latency, while acks=1 is faster but can lose acknowledged data on leader failure.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.