← All writing
articleJan 22, 202518 min read

Metadata Management: The Catalogue That Nobody Updates Until the Day They Need It

Schema registries, compatibility modes, ownership as data, and why a metadata system without enforced write paths becomes a documentation site.

Data PlatformSchema RegistryGovernanceArchitecture
Metadata Management: The Catalogue That Nobody Updates Until the Day They Need It cover illustration

Every organisation eventually builds a data catalogue, and most of them end up in the same place: a wiki nobody updates, containing descriptions that are technically plausible and wrong in a way that only surfaces when a dashboard is quietly incorrect.

The reason is not that documentation is hard to maintain. It is that the catalogue is a second source of truth with no enforcement, and second sources of truth without enforcement become lies. The design decisions that fix this are about write paths, compatibility, and making the useful thing the easy thing.

The scale of what is being catalogued

  2,000 services
  50,000 database tables
  3,000 Kafka topics
  20,000 event schemas
  8,000 dbt models
  15,000 dashboards
  4,000 metrics definitions

  ~100,000 registered assets

  without a registry:
    a schema change reaches 50 consumers
    over the next 2 weeks, discovered individually
    -> ~200 compatibility incidents per quarter

The number that justifies the system is the second one. Schema incompatibility is a slow, expensive, recurring cost that is nearly invisible until a consumer breaks hours or days after a change, and the fix is to reject the change at the moment it is proposed rather than to find the consumers afterwards.

The three kinds of metadata, and only one of them is required

  technical metadata
    - field names and types
    - nullability, defaults, keys
    - producers and consumers
    - partition and ordering guarantees
    -> MACHINE-READABLE, ENFORCEABLE
    -> the part the system can guarantee

  business metadata
    - what the field means in the business
    - which team's metric this is
    - the definition of "active user"
    -> HUMAN, NOT ENFORCEABLE
    -> the part that makes it useful

  operational metadata
    - freshness, row counts, update frequency
    - owner, on-call, retention
    - downstream lineage
    -> PARTLY MACHINE-GENERATED
    -> the part that makes it trustworthy

Teams usually start by attempting the middle one and wonder why it does not work. “What does this field mean” cannot be validated by a system; it can only be recorded, requested, and reviewed. It is also the part that determines whether anyone uses the catalogue, because during an incident the engineer needs to know whether amount_cents is gross or net and no amount of type information tells them.

The correct order is: enforce technical metadata first, because the system can actually guarantee it, then generate operational metadata automatically, then solicit business metadata from the teams who own the data, and treat business metadata as a reviewed artefact with an owner rather than as a field someone might fill in.

Compatibility: three modes, and they are not equal

This is the most concrete thing the registry does and the thing teams most often configure wrongly.

  v1                          v2
  {                          {
    "id": "string",             "id": "string",
    "amount": "int"             "amount": "int",     <- new
  }                              "currency": "string" <- new
                              }

Backward compatibility: new consumers can read old data. A new version must be readable by consumers written against the previous version. Adding a field with a default is fine. Removing a field is not. Making an optional field required is not. Narrowing a type is not. Widening is.

Forward compatibility: old consumers can read new data. A new version must be readable by consumers written against the next version — which in practice means consumers that are older. This is the constraint that stops you removing a field: an old consumer expects it, and its absence is a break. Adding a field with a default is fine. Removing is not. Changing a type is not.

Full compatibility: both. Both directions safe, which for a shared event schema means every change is additive-with-defaults. This is the strictest mode and the correct default for anything crossing a service boundary, because you genuinely do not control the consumers’ deploy schedule.

Which mode to use, honestly:

  internal topic, all consumers in one team,
  all deployed together       -> none / backward
  internal topic, consumers
  deploy independently        -> full
  external or partner-facing  -> full, always
  database schema with a
  coordinated migration      -> none, with review

The failure mode is choosing none because “we coordinate our deploys” and then having one team that does not. Compatibility checking is cheap; a compatibility incident is a multi-day investigation with an unknown number of affected consumers. The asymmetry is not close.

One important caveat: compatibility checking is only as good as the schemas you register. If a producer writes a schema to the registry and then publishes something different, the registry is confidently wrong. That is why enforcing registration at the producer — the schema must be registered before the topic is written — is the whole value proposition, and why an unregistered producer is a registry that does not work.

Ownership as a first-class, enforced field

The most valuable metadata field is owner, and the reason it needs enforcement is that ownership decays silently.

  ownership decays like this
    team publishes a table       -> owner: team-a     (correct)
    team-a reorgs                -> owner: team-a     (wrong)
    6 months later               -> still team-a
    an incident                  -> page team-a
    team-a no longer exists      -> nobody responds

  enforcement that prevents it
    - owner is a required field, validated against
      a service catalogue, not free text
    - the owning team is verified on publish
      (they must authenticate as that team)
    - a reorg triggers a revalidation of
      everything they own
    - dashboards and alerts inherit ownership
      from the data they read

The interesting property is the last one. Ownership that lives only on tables is incomplete, because the thing a person needs at 2am is ownership of the dashboard, and the dashboard’s owner should be derived from the data behind it. Deriving it means ownership stays correct as the data changes underneath, which is exactly the case where hand-maintained ownership is always stale.

Lineage has to be captured, not documented

  table: billing_events
    written by:      kafka-consumer #4
    kafka topic:    billing.events.v2
    produced by:    billing-service (producer)
    upstream:       postgres.billing.ledger
                    stripe API (via payments-service)
    consumed by:    dbt model billing_daily
                    -> mart.finance_daily_revenue
                    -> dashboard "Finance Overview"
                    -> alert "Revenue flat"

Manual lineage of that shape is wrong within one quarter and nobody will update it. Machine-captured lineage is available at several points and all of them are better than documentation:

  • Schema registry knows producers and consumers of topics, because they registered.
  • Query engines know the DAG of transformations from the query plan, so a table’s upstream is derivable from the jobs that read and write it.
  • dbt has a full lineage graph already, from static analysis of the model dependencies.
  • Orchestrators know the task DAG of every pipeline.
  • Data quality and observability tools know the runtime relationships from actual query traffic.

The practical approach is to assemble lineage from those sources rather than to introduce a place where people type lineage in. Runtime-derived lineage catches the dynamic dependencies that static analysis misses, and static analysis catches the ones that only happened to run last week.

The lineage feature that earns its cost is impact analysis: given a table, which dashboards, alerts, and models break if I change it. That is the question a schema change raises, and answering it from a lineage graph turns a week of discovery into a query.

Freshness and the trust problem

A catalogue entry with no freshness information is worse than no entry, because it looks authoritative.

  freshness signals, cheapest first
    - declared update frequency in the
      registration (a contract)
    - observed max(updated_at) in the data
    - observed write volume over a window
    - pipeline last-success timestamp

  the staleness signal people forget
    - the pipeline SUCCEEDED and produced
      the WRONG thing
    - freshness is green, correctness is not

The distinction in that last block is the one that makes data observability separate from metadata management. A pipeline that ran successfully an hour ago and has been silently truncating since is fresh and wrong. This is why the catalogue should carry a quality signal alongside a freshness signal, and why the two are displayed differently.

The derived rule for trust: show the last-updated time of the metadata entry itself, next to the last-updated time of the data. A description that says “average order value, excludes tax” where the definition changed eight months ago is actively harmful, and the only defence is making the metadata’s own age visible.

The catalogue as a deploy dependency

The adoption mechanism for schema registries is enforcement at the producer, and that means a deploy can be blocked on a catalogue check.

  deploy pipeline
      |
      v
  +---------------------------+
  |  producer registers the   |  must succeed
  |  schema BEFORE publish    |  -> this is the gate
  +------------+--------------+
               |
      compatibility check against
      the current registered version
               |
        pass  |  fail
              v     v
           deploy  reject with a
                   concrete diff and
                   the list of affected
                   consumers

Two consequences that need planning rather than discovering:

Registry availability is on the critical path of every deploy that touches data. If the registry is down, deploys block. This is a real availability dependency and it needs a considered answer — a cached compatibility snapshot, a local registry mirror, or an explicit degraded mode that allows deploys with a warning and reconciles later. Which you choose depends on how expensive a blocked deploy is relative to a bad deploy.

The error message is the adoption mechanism. A rejected schema that says “incompatible: removed field amount” is a fix in thirty seconds. One that says “COMPATIBILITY_CHECK_FAILED” is a ticket for the data team. The system should return the specific diff, the mode that failed, and the consumers affected, because a rejected deploy that does not explain itself gets a bypass added, and then the registry is advisory.

Failure stories worth testing

Register a schema, then publish something different

Confirm the registry rejects it. If it does not, the registry is documentation and you have lost the main benefit.

Remove a field with none compatibility selected

Confirm the deploy succeeds and consumers break over the following days. Then set the mode to full and confirm it is rejected with a useful message.

Add a field without a default and run full compatibility

Confirm it is rejected. An added field with no default breaks an old consumer that unmarshals strictly, and this is a common miss in otherwise careful schema reviews.

Reorg a team that owns 400 tables

Confirm the revalidation flow updates the owners. Ownership that is not revalidated is the decay path that produces an unpageable dashboard at 2am.

Change a table that nine dashboards read

Confirm impact analysis names all nine before the change is approved. If it cannot, the lineage was not captured from the pipelines.

Make a pipeline produce fresh but wrong data

Confirm the catalogue shows freshness green and quality bad, and that they are displayed separately. This is the failure that makes catalogues untrustworthy.

Take the registry down during a release window

Confirm what happens to in-flight deploys. A blocked deploy is annoying; a release freeze during an incident is not.

Reject a schema with an unhelpful message

Confirm the team’s response. This is the test for whether the error message is doing its job, and the usual outcome is a bypass.

Change a schema’s meaning without changing its structure

Rename nothing, change nothing in the type system, and the meaning changes anyway. Confirm the catalogue distinguishes the two, because only a human review catches the second.

Delete a table that an old version of a dashboard still queries

Confirm the deletion is blocked or warns with consumers. Unreferenced deletion is the one irreversible operation in a catalogue, and it needs a guard.

Have two teams register conflicting schemas for the same topic

Confirm the second registration is rejected. Last-write-wins in a schema registry is a silent correctness problem.

A production-ready architecture

   producers (services, streams, jobs)
        |
        |  register BEFORE publish
        v
  +-------------------------------+
  |  schema registry             |
  |  - versioned schemas          |
  |  - compatibility policy:      |
  |      backward / forward /full |
  |  - rejection returns the      |
  |    specific diff + consumers  |
  |  - owner required, verified   |
  |    against the service        |
  |    catalogue                  |
  +---------------+---------------+
                  |
     compatibility check gates
     every data-touching deploy
                  |
                  v
  +-------------------------------+
  |  catalogue                    |
  |  +------------------+         |
  |  | technical        |  types, keys, owners,
  |  | (enforced)       |  producers, consumers
  |  +------------------+         |
  |  | business         |  meaning, definitions,
  |  | (solicited,      |  reviewed with an owner
  |  |  owned)          |  and a visible entry age
  |  +------------------+         |
  |  | operational      |  freshness, quality,
  |  | (generated)      |  volume, last success
  |  +------------------+         |
  +---------------+---------------+
                  |
   lineage assembled from
   registry + query plans + dbt
   + orchestrator DAGs + runtime
   query traffic
                  |
                  v
  +-------------------------------+
  |  derived graph               |
  |  - upstream / downstream      |
  |  - impact analysis            |
  |  - ownership inherited by     |
  |    dashboards and alerts      |
  +-------------------------------+

A sensible delivery checklist:

  1. Enforce registration at the producer. A schema written but not registered makes every other feature advisory.
  2. Default to full compatibility for anything crossing a service boundary. It is the strictest mode and the only one that respects consumers you do not deploy.
  3. Make the rejection message a concrete diff plus the list of affected consumers. A rejected deploy that does not explain itself gets a bypass.
  4. Make owner required and validate it against the service catalogue rather than accepting free text.
  5. Re-validate ownership on a reorg, automatically, and inherit ownership into dashboards and alerts from the data they read.
  6. Derive lineage from query plans, dbt, orchestrators, and the registry. Do not add a form for typing lineage.
  7. Build impact analysis, not just a graph. “What breaks if I change this table” is the question that makes the catalogue worth running.
  8. Show the entry’s own age next to the data’s age. A stale definition is worse than a missing one.
  9. Show freshness and quality as separate signals, because fresh and wrong is the common and invisible case.
  10. Decide how the registry behaves when it is down, before you need to, and prefer a cached compatibility snapshot to a deploy freeze.
  11. Review business metadata on a schedule, with the owning team, and treat a missing review as a defect.
  12. Guard deletion of registered assets with a lineage check. It is the one irreversible operation and the one most likely to be run in a hurry.

Common mistakes

Mistake What actually happens Better decision
Optional schema registration Producers skip it, registry is documentation Enforce before publish
none compatibility on a shared topic Consumers break days later, slowly full across a service boundary
Rejection message is an error code Teams add a bypass within a week Concrete diff plus affected consumers
owner as free text Unpageable ownership after a reorg Validated against the service catalogue
Ownership hand-maintained per dashboard Stale the moment the data changes Inherit from the data, derive it
Manual lineage Wrong within a quarter, never updated Derive from query plans, dbt, orchestrators
Lineage without impact analysis A graph nobody queries during a change “What breaks if I change this” as a feature
No entry-age display An eight-month-old definition looks authoritative Show metadata age next to data age
Freshness as the only quality signal Fresh and wrong passes every check Separate freshness and correctness signals
Registry down blocks all deploys Release freeze during an incident Cached compatibility snapshot, degraded mode
Business metadata with no review Meaning drifts, definitions conflict Owner review on a schedule
Deleting an asset with no lineage check A dashboard queries a table that is gone Guard deletion with a consumer check
Last-write-wins on schema registration Two teams, two truths, one silent break Reject the conflicting registration
Cataloguing only tables The dashboard a human needs has no entry Catalogue dashboards and alerts too
Technical metadata only Precise and useless during an incident Solicit business meaning with an owner
No schema history Cannot answer “when did this field appear” Full version history, queryable
One catalogue for everything Nobody can find their layer Federate, link, or filter by domain
Metadata as a nightly export Always a day stale, never authoritative Read from the source systems, live

The complete story in one minute

A metadata system fails in a predictable way: it becomes a second source of truth with no enforcement, and second sources of truth without enforcement become lies. The fix is to make the correct write path the path of least resistance — producers must register a schema before they publish, and a deploy that would break an existing consumer is rejected at proposal time rather than discovered days later by the consumer.

Three kinds of metadata need different treatment. Technical metadata is machine-readable and enforceable, so the system can guarantee it. Business metadata — what a field actually means, whether amount_cents is gross — cannot be validated by any system and is the part that makes a catalogue worth opening during an incident, so it needs a named owner and a review. Operational metadata is mostly generated and is what makes the entry trustworthy.

Compatibility is the concrete payoff. Backward protects new consumers from old data, forward protects old consumers from new data, and full is both. Anything crossing a service boundary should be full, because you do not control your consumers’ deploy schedules, and the asymmetry is not close: a compatibility check costs seconds, a compatibility incident costs days with an unknown blast radius.

Lineage must be captured from the pipelines that produce data — query plans, dbt, orchestrator DAGs, registry registrations — because hand-maintained lineage is wrong within a quarter. And the feature that earns the whole system’s cost is impact analysis: given a table, which dashboards, alerts, and models break if I change it.

Finally, the catalogue becomes a deploy dependency the moment you enforce it, which is a real availability dependency. Decide what happens when the registry is down before you need to, and make the rejection message good enough that nobody is tempted to add a bypass.

register before publish, or the registry is a wiki
compatibility: full across service boundaries
lineage: captured from pipelines, never typed
trust: metadata age + data age + freshness + correctness, all visible

The hard part was never storing the schemas. It was making the catalogue the only place the truth lives, instead of the place where someone once wrote down what they believed.

Technical references

Keep reading
Browse everything