Converters vs Single Message Transforms: Serialization vs Transformation
Converters and Single Message Transforms (SMTs) both operate on records as they flow through Kafka Connect, but they solve completely different problems. A converter is responsible for serialization: it translates between the connector's internal data representation (a ConnectRecord with a schema and a value) and the byte array that Kafka stores. Converters are configured for keys and values separately, and they apply to every record. An SMT is responsible for transformation: it modifies the record's key, value, headers, or even its destination topic. SMTs are optional and are applied in a chain in the order you specify. The key distinction is that converters are about format, while SMTs are about content. You cannot skip converters; every record must be serialized to bytes before being written to Kafka and deserialized after being read. You can skip SMTs entirely.
The mechanism matters because it determines where in the pipeline each operates and what it can see. For a source connector, the flow is: read from external system, apply SMTs to the ConnectRecord, convert to bytes using the converter, write to Kafka. For a sink connector, the flow is: read bytes from Kafka, convert from bytes using the converter, apply SMTs to the ConnectRecord, write to external system. This means a source SMT operates on the connector's native representation of the external data, before it is serialized. A sink SMT operates on the deserialized Kafka record, after it has been converted from bytes. This ordering is why SMTs are the same for both directions but their position in the pipeline differs. Converters, on the other hand, are configured at the worker level and can be overridden per connector. The default converter is usually JSON or Avro, depending on your schema registry setup.
A common mistake is to use SMTs for complex transformation logic. SMTs are intentionally limited to single-record, stateless operations: masking a field, renaming a field, inserting a timestamp, routing to a different topic based on a regex. They cannot do windowed aggregations, joins, or stateful processing. If you find yourself chaining many SMTs to implement business logic, you have outgrown Connect and should move that logic to Kafka Streams or a custom processor. Another mistake is to confuse the converter's role with the connector's serialization. The connector produces a ConnectRecord with a schema; the converter decides how to serialize that schema and value to bytes. If you change the converter, you change the wire format, which can break downstream consumers. Version note: the converter and SMT APIs have been stable for many versions. Kafka Connect includes a set of built-in SMTs (mask, replace, insert, route, etc.), and additional SMTs are available from Confluent and community plugins. The built-in SMTs are documented in the Kafka Connect user guide.
Converters handle serialization between ConnectRecord and byte arrays; they are mandatory.
SMTs handle per-record transformations like masking, renaming, and routing; they are optional.
Source connectors apply SMTs before conversion; sink connectors apply SMTs after conversion.
Converters are configured at worker level and can be overridden per connector.
SMTs are stateless and single-record; they cannot do joins, aggregations, or windowing.
Using many SMTs for business logic is a sign you need a stream processor instead.
Built-in SMTs are available in Kafka Connect; additional SMTs come from plugins.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience