DE Wikiarchitectures / data-mesh

Data Mesh

Data mesh is a decentralized sociotechnical architecture for managing analytical data at scale. Coined by Zhamak Dehghani, it applies domain-driven design and platform-thinking principles to data management, organizing data around business domains rather than a centralized data platform.

Four Principles

1. Domain Ownership

Each business domain owns and manages its data as a product. The marketing team owns marketing data, the finance team owns financial data, etc. This flips the centralized model where a single data team manages all data.

  • Domains are responsible for collecting, cleaning, and serving their data
  • Data is treated as a product with SLAs, documentation, and quality guarantees
  • Each domain has its own data pipeline and storage

2. Data as a Product

Every dataset is treated as a product with defined quality, schema, and access patterns. Data products are discoverable, addressable, trustworthy, and self-serve.

  • Discoverable - Registered in a central catalog
  • Addressable - Has a unique, permanent identifier (URN)
  • Trustworthy - Includes quality metrics and lineage
  • Self-serve - Consumable via well-documented APIs or tables

3. Self-Serve Data Platform

A shared infrastructure platform that makes it easy for domains to build, deploy, and operate their data products. The platform handles:

  • Storage and compute infrastructure (lakehouse, warehouse)
  • Data catalog and discovery
  • Lineage tracking and governance
  • Monitoring and alerting
  • Access control and auditing

4. Federated Computational Governance

Global standards are enforced through automation rather than manual gatekeeping. The platform automates compliance with:

  • Schema standards and evolution policies
  • Data quality thresholds and SLAs
  • Access control rules (RBAC/ABAC)
  • PII detection and anonymization
  • Retention and archival policies

Architecture Diagram (Text)

┌─────────────────────────────────────────────────────┐ │              Data Platform (Infrastructure)          │ │  ┌─────────┐  ┌──────────┐  ┌────────┐  ┌────────┐ │ │  │ Catalog  │  │ Lineage  │  │ Policy │  │Compute │ │ │  │         │  │ Tracking │  │ Engine │  │  &    │ │ │  │         │  │          │  │        │  │Storage │ │ │  └─────────┘  └──────────┘  └────────┘  └────────┘ │ └─────────────────────────────────────────────────────┘ ▲              ▲              ▲ │              │              │ ┌────────┴──┐  ┌────────┴──┐  ┌────────┴──┐ │ Marketing │  │  Finance  │  │  Product   │ │  Domain   │  │  Domain   │  │  Domain    │ │           │  │           │  │           │ │ Data Prod │  │ Data Prod │  │ Data Prod │ │ Customers │  │ Revenue   │  │ Sessions   │ │ Campaigns │  │ Expenses  │  │ Features   │ └───────────┘  └───────────┘  └───────────┘

Data Product Contract

// Data product manifest (YAML) name: customers_current domain: marketing owner: marketing-team@company.com version: 2.1.0 schema: fields: - name: customer_id type: string required: true description: Unique customer identifier - name: email type: string required: true pii: true - name: segment type: string required: false enum: [premium, standard, basic] - name: lifetime_value type: decimal(12,2) required: true output_ports: - type: table location: prod.marketing.customers_current format: delta - type: api endpoint: /api/v1/products/customers auth: oauth2 quality_slas: completeness: 0.999       # 99.9% of fields non-null freshness: 3600           # max 1 hour stale uniqueness: - customer_id lineage: upstream: - domain: crm product: contacts_raw - domain: sales product: transactions downstream: - domain: analytics product: customer_360

When to Adopt Data Mesh

  • Your organization has 5+ data domains with distinct ownership
  • A central data team is becoming a bottleneck
  • Domain teams already have data engineering skills
  • You need to scale data ownership without growing the central team proportionally

Challenges

  • Cultural shift - Requires domain teams to take ownership of data quality and operations
  • Platform investment - Significant upfront investment in the self-serve platform
  • Governance - Balancing domain autonomy with global standards and compliance
  • Cross-domain joins - Joining data across domains requires coordination and consistent identity resolution

Resources

Data Mesh

Data mesh is a decentralized sociotechnical architecture for managing analytical data at scale. Coined by Zhamak Dehghani, it applies domain-driven design and platform-thinking principles to data management, organizing data around business domains rather than a centralized data platform.

Four Principles

1. Domain Ownership

Each business domain owns and manages its data as a product. The marketing team owns marketing data, the finance team owns financial data, etc. This flips the centralized model where a single data team manages all data.

  • Domains are responsible for collecting, cleaning, and serving their data
  • Data is treated as a product with SLAs, documentation, and quality guarantees
  • Each domain has its own data pipeline and storage

2. Data as a Product

Every dataset is treated as a product with defined quality, schema, and access patterns. Data products are discoverable, addressable, trustworthy, and self-serve.

  • Discoverable - Registered in a central catalog
  • Addressable - Has a unique, permanent identifier (URN)
  • Trustworthy - Includes quality metrics and lineage
  • Self-serve - Consumable via well-documented APIs or tables

3. Self-Serve Data Platform

A shared infrastructure platform that makes it easy for domains to build, deploy, and operate their data products. The platform handles:

  • Storage and compute infrastructure (lakehouse, warehouse)
  • Data catalog and discovery
  • Lineage tracking and governance
  • Monitoring and alerting
  • Access control and auditing

4. Federated Computational Governance

Global standards are enforced through automation rather than manual gatekeeping. The platform automates compliance with:

  • Schema standards and evolution policies
  • Data quality thresholds and SLAs
  • Access control rules (RBAC/ABAC)
  • PII detection and anonymization
  • Retention and archival policies

Architecture Diagram (Text)

┌─────────────────────────────────────────────────────┐ │              Data Platform (Infrastructure)          │ │  ┌─────────┐  ┌──────────┐  ┌────────┐  ┌────────┐ │ │  │ Catalog  │  │ Lineage  │  │ Policy │  │Compute │ │ │  │         │  │ Tracking │  │ Engine │  │  &    │ │ │  │         │  │          │  │        │  │Storage │ │ │  └─────────┘  └──────────┘  └────────┘  └────────┘ │ └─────────────────────────────────────────────────────┘ ▲              ▲              ▲ │              │              │ ┌────────┴──┐  ┌────────┴──┐  ┌────────┴──┐ │ Marketing │  │  Finance  │  │  Product   │ │  Domain   │  │  Domain   │  │  Domain    │ │           │  │           │  │           │ │ Data Prod │  │ Data Prod │  │ Data Prod │ │ Customers │  │ Revenue   │  │ Sessions   │ │ Campaigns │  │ Expenses  │  │ Features   │ └───────────┘  └───────────┘  └───────────┘

Data Product Contract

// Data product manifest (YAML) name: customers_current domain: marketing owner: marketing-team@company.com version: 2.1.0 schema: fields: - name: customer_id type: string required: true description: Unique customer identifier - name: email type: string required: true pii: true - name: segment type: string required: false enum: [premium, standard, basic] - name: lifetime_value type: decimal(12,2) required: true output_ports: - type: table location: prod.marketing.customers_current format: delta - type: api endpoint: /api/v1/products/customers auth: oauth2 quality_slas: completeness: 0.999       # 99.9% of fields non-null freshness: 3600           # max 1 hour stale uniqueness: - customer_id lineage: upstream: - domain: crm product: contacts_raw - domain: sales product: transactions downstream: - domain: analytics product: customer_360

When to Adopt Data Mesh

  • Your organization has 5+ data domains with distinct ownership
  • A central data team is becoming a bottleneck
  • Domain teams already have data engineering skills
  • You need to scale data ownership without growing the central team proportionally

Challenges

  • Cultural shift - Requires domain teams to take ownership of data quality and operations
  • Platform investment - Significant upfront investment in the self-serve platform
  • Governance - Balancing domain autonomy with global standards and compliance
  • Cross-domain joins - Joining data across domains requires coordination and consistent identity resolution

Resources