02 / 05

What is the conceptual difference between KStream and KTable?

Difficulty: 4/10
Stateful stream processing, State stores, Windowing

KStream vs KTable: Event Streams and Changelog Tables

The conceptual difference is that a KStream represents an event stream, where every record is an independent fact, and a KTable represents a changelog table, where every record is an update to the current state for a key. In a KStream, if you receive three records with the same key, you see three events. In a KTable, if you receive three records with the same key, you see only the latest value; the earlier values are superseded. This is the same distinction as an insert-only log versus an upsert log. A KStream is like a sequence of transactions; a KTable is like a database table that is continuously updated. The KTable is materialized from a compacted topic, which is why it retains only the latest value per key.

The mechanism matters because it determines how operations behave. A join between two KStreams is a windowed join, because you need to bound how long you wait for the matching record on the other side. A join between a KStream and a KTable is a stream-table join, where each stream record is enriched with the current value from the table. A join between two KTables is a table-table join, which produces a new table that updates whenever either side updates. Aggregations like count, reduce, and aggregate produce a KTable, because the result is a continuously updated state per key. If you want to turn a KTable back into a KStream, you use toStream(), which emits a record for every update. If you want to turn a KStream into a KTable, you use toTable(), which requires a key and effectively upserts. This duality is one of the most powerful ideas in Kafka Streams, but it is also a common source of confusion.

A common mistake is to treat a KTable as if it were a KStream and expect to see every update. If you do a forEach on a KTable, you only see the latest value per key at the time of processing, not the full history. Another mistake is to use toTable() on a stream with duplicate keys and expect all records to be preserved; they are not, because the table upserts. The trade-off is between history and state. KStream preserves full history and is good for event-driven processing, auditing, and replay. KTable preserves current state and is good for lookups, enrichment, and aggregations. In practice, most topologies use both: a KStream for the raw events and a KTable for the reference data or aggregated state. Version note: the KTable abstraction has been stable since early Kafka Streams versions, but the introduction of GlobalKTable in 0.10.2 added a way to replicate a table to all instances for joins without co-partitioning. In Kafka 3.x, the DSL continues to support these abstractions, and the Processor API exposes the underlying state stores directly.

javascript
  1. 1

    KStream = event stream; every record is an independent fact, full history preserved.

  2. 2

    KTable = changelog table; every record is an update, only latest value per key is retained.

  3. 3

    Stream-stream joins are windowed; stream-table joins enrich events with current state.

  4. 4

    Aggregations produce a KTable because the result is continuously updated state.

  5. 5

    toStream() turns a KTable into a KStream of updates; toTable() upserts a KStream into a KTable.

  6. 6

    Common mistake: expecting a KTable to emit every update; it emits latest per key.

  7. 7

    GlobalKTable (0.10.2+) replicates a table to all instances for joins without co-partitioning.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.