← All writing
articleFeb 24, 202518 min read

Data Mesh Architecture: Four Principles and the Organisation Chart Hiding Inside Them

Domain ownership, data as a product, the federated plane, and why data mesh is a reorganisation with an API in front of it.

Data PlatformDomain-Driven DesignGovernanceArchitecture
Data Mesh Architecture: Four Principles and the Organisation Chart Hiding Inside Them cover illustration

Data mesh gets described as an architecture, and it is described that way because the term sounds technical. It is not. It is an organisational model — four principles about who owns what — plus the technical enablers that make the model possible.

The technical part is the easy part and it is well understood. The organisational part is where every data mesh initiative fails, and the failure always looks the same: a central platform team that kept all the power, domains that did not want the responsibility, and a lake full of tables nobody owns.

What the scale changes

  one team, 50 tables
  -> a central analytics team knows everything
  -> a data warehouse with a well-designed
     dimensional model works fine
  -> no reason for mesh

  20 teams, 50,000 tables, 400 pipelines
  -> a central team cannot know who owns
     `customer_touchpoints_final_v2`
  -> every request is a ticket to one team
  -> the queue is the bottleneck
  -> analytics takes 3 weeks to get a number

  mesh claim
  -> ownership distributed to 20 domains
  -> each domain ships its own data products
  -> self-service platform removes the ticket

The trigger is not a table count, it is the ticket queue. A central data team becomes a bottleneck somewhere around the point where a new request takes more than a couple of days, and at that point the options are mesh, hiring more central people, or accepting the bottleneck. Mesh is one of three answers, and the choice depends on organisational facts rather than technical ones.

The four principles, and what each one actually requires

Domain ownership

Data should be owned by the people closest to how it is created and used, and that ownership should be accountable.*

The technical translation is “the producing team owns the table.” The organisational translation, which is the one that matters, is “the producing team is paged when the data is wrong.”

  what domain ownership requires
    - a named team, in a service catalogue,
      not a group alias
    - that team is on the pager for data quality
    - that team is on the pager for freshness
    - they have the authority to change the
      schema and the schema registry enforces
      that consumers are checked
    - they have the authority to fix the
      pipeline without asking the platform team

The third and fourth points are where it falls apart in practice. A domain team that owns a table but cannot change the pipeline without a platform ticket is not the owner; they are a volunteer. And a team that is paged for data quality without having the tooling or authority to fix it will stop answering after a few months, and then the ownership is decorative.

The accountability test is whether the pager rings for them. If the on-call rotation for a data domain is the platform team, mesh has not happened regardless of what the documentation says.

Data as a product

Data should be published, documented, and treated as a product with consumers and a support commitment.*

A data product has a specific shape, and it is not a table.

  a data product, minimally
    - a stable, documented, versioned interface
      (schema, semantics, contract)
    - a stated quality bar
      (completeness, freshness, accuracy, with numbers)
    - an SLA for availability and freshness
    - a support commitment with a channel
      and a response expectation
    - deprecation policy: how consumers get
      told, and how long they get
    - documented known limitations

The quality bar needs numbers. “High quality” is not a quality bar. “99% of rows have a non-null customer_id, refreshed within 15 minutes of the source event, p99 freshness under 30 minutes” is a bar, and it can be monitored, and when it is breached something happens.

The consumer test is the one that exposes fake data products. A data product with no consumers is a cost centre with a schema. Teams routinely publish twelve “data products” that nobody uses because nobody was asked what they needed — and the fix is not better documentation, it is starting from a consumer. The honest sequence is: find out what two teams need, publish that, make it good, and only then generalise.

The federated plane

Data architecture should be a self-serve platform, so that new teams can onboard and ship without a central team involved.*

This is the enabler that makes the other two possible, and it is the part that is most straightforwardly an engineering project.

  what the platform provides
    - data ingestion from every source, templated
    - storage, partitioned, in every format
    - orchestration for batch and streaming
    - transformation with testing and lineage
    - data quality checks, run automatically
    - catalogue, discovery, lineage
    - schema registry and compatibility checks
    - access control, auditing
    - cost visibility per product

  what the domain does
    - writes and versions transformations
    - defines the contract and the quality bar
    - runs the checks
    - responds to consumers
    - names an owner

  what neither does
    - nothing requires the platform team

The test of a federated plane is whether a new domain can ship its first data product in a week without talking to the platform team. If every onboarding needs a ticket, it is a central team with extra steps. This is a large amount of work, and it is the work that has to happen before ownership can be distributed, not after.

Technology as an enabler

Use technology to reduce coupling, not to add it.*

This is the principle most often agreed with and least often applied, and it is a constraint on everything else.

  reduce coupling
    - open table formats (Parquet, Iceberg, Delta)
      so data is not trapped in one engine
    - contracts between domains, enforced by
      a schema registry, not by documentation
    - data movement driven by contracts
      (CDC), not by scheduled batch copies
    - compute separated from storage

  add coupling (avoid)
    - one proprietary engine that only a
      platform team can operate
    - cross-domain joins inside a domain's
      transformation, hidden from the graph
    - a shared "scratch" schema where
      everything is dumped and anyone reads
      - lineage that stops at the platform

The specific thing to avoid is the shared scratch area. A schema where every domain publishes raw dumps and every consumer reads from is not a mesh; it is the centralised model with extra steps and a lot of confused documentation.

What this looks like as an architecture

  +-----------------------------------------------------------------+
  |  domain: payments                                              |
  |                                                                 |
  |   postgres -> CDC -> bronze/silver/gold                        |
  |   contract: payment_events.v3 (schema registry)                |
  |   quality: freshness < 15min, 99.9% valid amounts              |
  |   owner: team-payments (on the pager)                           |
  |   consumers: finance, fraud, exec dashboard                    |
  +----------------------------+------------------------------------+
                               |
                               |  self-serve, no ticket
                               v
  +-----------------------------------------------------------------+
  |  federated plane (shared, self-serve)                           |
  |                                                                 |
  |   ingestion      CDC, batch, streaming, templated pipelines     |
  |   storage         object store + open table formats             |
  |   orchestration   batch DAGs and streaming jobs                 |
  |   transformation  SQL, dbt-like, tested, versioned              |
  |   quality         declarative checks, run per load, per domain  |
  |   contracts       schema registry, compatibility enforced      |
  |   discovery       catalogue, lineage, impact analysis           |
  |   access          per-product permissions, audited               |
  |   cost            per product, per team, visible                |
  +----------------------------+------------------------------------+
                               |
          +--------------------+--------------------+
          |                                         |
          v                                         v
  +------------------+                   +------------------+
  | domain: fraud    |                   | domain: finance  |
  | owns its own     |                   | owns its own     |
  | products         |                   | products         |
  | consumes         |                   | consumes         |
  | payment_events   |                   | payment_events   |
  +------------------+                   +------------------+

  no domain reads another domain's bronze or silver
  cross-domain access is through a published contract

The line at the bottom is the one that gets violated first. A team under deadline reaches into another domain’s raw tables because the contract is not available or is too slow, and once that happens the producing domain can no longer change anything without coordinating with a consumer nobody declared. The escape hatch becomes permanent because breaking it breaks someone.

The way to prevent it is to make the contract the fastest path, not just the sanctioned one. If the published product is easier to query than the raw table, the incentive is right. If it takes a week to get a contract change through, engineers will read the bronze layer and you will find out in a month.

The migration is the hard part

Nobody converts to mesh by announcing it. The sequence that works is unglamorous.

Start with one domain that has an obvious owner and an obvious consumer. A team that already knows their data is well and whose consumers already ask them questions. Prove the model on them, including the quality bar and the pager, before asking anyone else to take on responsibility.

Build the federated plane for that domain, self-serve. The point is not to make one team faster; it is to demonstrate that onboarding does not need the platform team. If it does need the platform team, the model is not proven and the next twenty teams will be right to refuse.

Then hand over existing pipelines with their existing consumers. The hard part is not the new domain, it is the twenty central-team-owned pipelines that three teams depend on. Migrating one means identifying the consumers, publishing a contract, and giving the consumers a deprecation window. This is months of work per domain and it is the actual cost of mesh.

Leave the genuinely cross-cutting work central. Identity, the physical storage layer, the ingestion framework, and cost accounting are platform concerns. Data that is genuinely enterprise-wide — a single customer dimension, a consolidated ledger — may not have a natural domain owner, and forcing one is worse than leaving it central. “Most data is a domain product” is the honest goal, not “all data”.

Do not migrate a domain that is in trouble. A domain with a failing pipeline, an unhappy consumer, and a team understaffed cannot absorb an organisational change. Stabilise it first or leave it where it is.

Failure stories worth testing

Publish a data product with no consumers

Confirm someone notices. A product nobody uses is a cost centre, and the honest response is to either find the consumer or stop publishing.

Announce mesh and ask the central team to keep all approvals

Confirm the domains can change a pipeline without permission. If they cannot, ownership is nominal and the ticket queue is unchanged.

Have a domain read another domain’s raw tables under deadline

Confirm the escape hatch is closed. This is the first rule violation in every mesh migration and it is only prevented by making the contract the fast path.

Break a data quality bar and check who gets paged

If the answer is the platform team, ownership is decorative. The pager is the only honest measure of whether the model changed anything.

Let a domain change a schema that nine consumers read

Confirm the registry rejected it, or that nine consumers broke. Without enforced contracts, “domain owns the data” means “domain breaks other people silently”.

Run a new domain’s onboarding and time it

If it needs a ticket to the platform team, the federated plane does not exist. This is the test of the whole principle.

Force a single owner onto genuinely cross-cutting data

Confirm the owner can actually maintain it. Enterprise-wide data without a natural domain produces a nominal owner who does not do the work.

Migrate a domain with three dependent teams

Confirm the consumers were identified and given a deprecation window. This is the expensive part and the part that determines whether the migration is safe.

Measure the central data team’s ticket queue six months in

If it is unchanged, mesh has not changed anything. That metric is the reason the model was adopted and it should be the one reported.

Leave a scratch schema where everything is dumped

Confirm nothing important depends on it. This is how centralised data architecture survives a mesh migration, entirely undocumented.

Give every domain its own storage and tooling

Confirm the platform can support twenty of them. Uniform tooling is what makes self-serve possible; twenty bespoke stacks is a central team with travel time.

A production-ready architecture

   enterprise concerns (stay central, deliberately)
     identity, access, physical storage,
     ingestion framework, cost accounting,
     genuinely cross-cutting data
        |
        v
  +---------------------------------------------------------+
  |  federated plane: self-serve, templated, no tickets     |
  |  ingestion | storage | orchestration | transformation   |
  |  quality | contracts | discovery | access | cost        |
  +---------------------------------------------------------+
        |              |              |
   +---------+   +---------+   +---------+
   v         v   v         v   v         v
 +--------+ +--------+ +--------+ +--------+
 |payments| |fraud   | |finance | |growth  |
 |domain  | |domain  | |domain  | |domain  |
 +--------+ +--------+ +--------+ +--------+
   |        |        |        |        |
   | publishes a data product:
   |   versioned contract
   |   quality bar with numbers
   |   freshness SLA
   |   owner on the pager
   |   deprecation policy
   |        |
   +--------+--------+--------+
            |  consumers read
            |  ONLY the contract
            v
   other domains, dashboards,
   ML features, exports

   never: another domain's raw
         bronze or silver tables

A sensible delivery checklist:

  1. Prove the model on one domain with an obvious owner and obvious consumers before proposing it to anyone else.
  2. Make onboarding self-serve and time it. A ticket to the platform team means the federated plane does not exist yet.
  3. Put domain teams on the pager for their own data quality and freshness. That is the measurement of whether ownership transferred.
  4. Give domains real authority — over their pipelines, their schemas, and their quality bars — or ownership is a label.
  5. Start every data product from a named consumer, and do not generalise a product nobody asked for.
  6. Write quality bars as numbers with thresholds, and run the checks automatically on every load.
  7. Make the published contract the fastest path to the data, so reading someone else’s raw layer is never the convenient option.
  8. Enforce compatibility at the registry so a domain can change its schema without breaking undeclared consumers.
  9. Leave identity, physical infrastructure, ingestion framework, and cost accounting central, and be honest about cross-cutting data.
  10. Migrate existing pipelines with their consumers identified and a deprecation window agreed. Budget months, not weeks.
  11. Do not migrate a domain that is already in trouble. Stabilise first.
  12. Report the central data team’s ticket queue and time-to-first-data-product. If those do not move, the model is not working regardless of the documentation.

Common mistakes

Mistake What actually happens Better decision
Announce mesh, keep central approvals Domains own nothing, the queue is unchanged Give real authority or do not start
Publish products nobody consumes Cost centres with schemas Start from a named consumer
Platform team stays the on-call for data Ownership is decorative Domains on the pager, with tooling
Read another domain’s raw tables Undeclared coupling, no safe schema change Contract-only access, enforced
One centralised data team, 50,000 tables Requests queue for weeks Mesh, or hire, or accept the bottleneck
No federated plane Every domain needs a platform ticket Build self-serve before distributing ownership
Force an owner onto cross-cutting data Nominal owner, no maintenance Leave genuinely cross-cutting data central
Migrate a domain that is failing The organisational change is absorbed by the fire Stabilise first
Quality bar is “high quality” Not measurable, not monitored Numbers, thresholds, automatic checks
Every domain gets bespoke tooling Central team now supports twenty stacks Uniform templated platform
No deprecation policy Contract changes break consumers silently Stated window and communication
Scratch schema survives the migration The old centralised model lives on, undocumented Audit and remove it
Lake of raw data with no products Storage cost with no consumption Only publish what has a consumer
Central team owns pipelines forever Mesh is 5% adopted and 95% theatre Migrate with consumers and a window
Success measured by domain count Domains created, nothing consumed Measure consumers and time-to-first-product
Skip the business definition Numbers computed three ways Definition owned by the domain, reviewed
Assume mesh is a migration off the warehouse Warehouse and mesh are different tools Keep the warehouse for what it is good at
Ignore cost per product Platform costs grow invisibly Per-product cost visibility from day one
Treat governance as separate from mesh Policy is another central gate Policy as platform capability domains use

The complete story in one minute

Data mesh is not an architecture. It is an organisational model with technical enablers, and the failure is always the same one: the central team kept the power, so nothing changed except documentation. The four principles are domain ownership, data as a product, the federated plane, and technology as an enabler — and each has a testable form.

Domain ownership means being paged for the data’s quality and freshness, with the authority to fix the pipeline without a ticket. If the central team is still on the data on-call rotation, ownership has not moved. Data as a product means a versioned contract, a quality bar written as numbers, a freshness SLA, a support commitment, and a deprecation policy — and the honest test is whether anything consumes it, because a product with no consumers is a cost centre with a schema. The federated plane is the enabler: ingestion, orchestration, transformation, quality, contracts, discovery, and cost, all self-serve, so a new domain can ship in a week without talking to anyone. And technology as an enabler is a constraint — open table formats and enforced contracts, no proprietary engine only a platform team can run, and no shared scratch schema.

The rule that breaks first is cross-domain access to raw tables. An engineer under deadline reads another domain’s bronze layer because the contract is not available yet, and once that happens the producing domain can never change anything without coordinating with a consumer nobody declared. The only real prevention is making the published contract the fastest path, not just the sanctioned one.

The migration is months of unglamorous work per domain, it starts with one domain that has an obvious owner, and it deliberately leaves identity, physical storage, the ingestion framework, and genuinely cross-cutting data central.

domain owns the data, and is paged for it
products start from a consumer, quality bars are numbers
self-serve platform, so no domain needs a ticket
consumers read contracts, never raw tables

The hard part was never choosing open table formats. It was being willing to hand a central team its ticket queue, and then not taking it back.

Technical references

Keep reading
Browse everything