04 / 05

A producer deployment breaks older consumers because of an incompatible schema change. How would you recover?

Difficulty: 8/10
Schema evolution, Compatibility, Schema Registry

Recovering from an Incompatible Schema Change in Production

The first priority is to stop the bleeding. If older consumers are failing, you need to decide whether to roll back the producer or roll forward the consumers. The fastest recovery is usually to roll back the producer to the previous schema version if the change was just deployed. This restores the contract that consumers expect. If rollback is not possible, for example because the producer has already written data with the new schema, you need to deploy a fix to consumers or introduce a compatibility layer. The key insight is that the data already written with the new schema is in the topic, and consumers that read from the beginning or from an offset before the change will encounter both old and new records. This means you cannot simply roll back the producer and assume everything is fine; you must handle the mixed-schema data.

The mechanism for recovery depends on the nature of the incompatibility. If the change was adding a required field with no default, the fix is to make the field optional with a default in a new schema version, then roll out that version. If the change was removing a field that consumers need, the fix is to re-add the field with a default and deprecate it properly. If the change was a type change, the fix is to introduce a new field with the new type and keep the old field for backward compatibility, then migrate consumers over time. In all cases, you should use a schema registry with a compatibility policy that would have caught the change before deployment. The recovery process should also include a post-mortem to understand why the compatibility check did not prevent the issue, and to fix the process so it does not happen again.

A common mistake is to try to fix the problem by editing the schema in the registry or deleting the incompatible version. This does not work because the data is already serialized with the new schema, and consumers that fetch the schema by ID will get the wrong schema if you mutate it. Another mistake is to assume that consumers can be fixed by simply ignoring the new field. If the field is required and the consumer's schema does not have it, deserialization fails. The trade-off is between speed of recovery and correctness. Rolling back the producer is fast but may not be possible if the new schema is already in use. Rolling forward consumers is slower but more correct. In the long term, the fix is process: use a schema registry with FULL_TRANSITIVE compatibility for shared topics, run compatibility checks in CI before deployment, and have a rollback plan for every schema change. Version note: Confluent Schema Registry allows you to check compatibility before registering, and you can configure the registry to reject incompatible schemas. Use this in your deployment pipeline.

javascript
  1. 1

    First priority: stop the bleeding by rolling back the producer or rolling forward consumers.

  2. 2

    Data already written with the new schema is in the topic; you must handle mixed-schema data.

  3. 3

    Fix depends on the incompatibility: make required fields optional, re-add removed fields, or introduce new fields.

  4. 4

    Do not edit or delete schema versions in the registry; the data is already serialized with them.

  5. 5

    Use a schema registry with compatibility checks in CI to prevent the issue.

  6. 6

    Set FULL_TRANSITIVE for shared topics to protect consumers that are many versions behind.

  7. 7

    Post-mortem: understand why the check did not catch it and fix the process.

Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.