DE Wikiconcepts / harnesses

Harnesses

In data engineering, a harness is an integration framework that connects, transforms, and monitors data flowing between systems. Harnesses abstract away the complexity of individual source and destination APIs, providing a unified interface for data movement and transformation.

What is a Data Harness?

A harness sits between your data sources and your data platform, handling:

  • Connector management - Pre-built connectors for common sources (databases, APIs, SaaS platforms)
  • Schema inference and mapping - Automatically detect schemas and map source fields to target schemas
  • Transformation pipelines - Built-in transform logic for cleaning, enriching, and normalizing data
  • Monitoring and observability - Track data volume, freshness, and quality metrics across all pipelines
  • Error handling - Retry logic, dead-letter queues, and alerting on failures

Popular Harness Tools

Fivetran

Fivetran is a fully managed data integration platform. It provides hundreds of pre-built connectors that handle schema changes, incremental updates, and error recovery automatically. Fivetran uses a harness model where you configure sources and destinations through a UI, and the platform handles the rest.

  • Automatic schema drift detection and propagation
  • Incremental updates via CDC or API pagination
  • dbt Cloud integration for post-load transforms
  • History mode for tracking all changes

Airbyte

Airbyte is an open-source data integration platform. It offers a similar harness model to Fivetran but with the flexibility of self-hosting and a large catalog of community connectors.

  • Open-source with a generous ELT2 license
  • Connector development kit (CDK) for building custom connectors
  • Protocol-guaranteed data delivery with at-least-once semantics
  • PyAirbyte for programmatic connector usage

Stitch (by Talend)

Stitch is a simple, reliable ETL service for streaming data from sources to warehouses. It focuses on simplicity and reliability over extensive transformation capabilities.

Building a Custom Harness

For specialized use cases, you might build a lightweight harness using Python and existing libraries:

from typing import Protocol
import pandas as pd

class Source(Protocol):
    def extract(self, since: str | None = None) -> pd.DataFrame: ...

class Destination(Protocol):
    def load(self, df: pd.DataFrame, schema: str, table: str) -> int: ...

class Transformer(Protocol):
    def transform(self, df: pd.DataFrame) -> pd.DataFrame: ...

class DataHarness:
    def __init__(
        self,
        source: Source,
        destination: Destination,
        transforms: list[Transformer] | None = None,
    ):
        self.source = source
        self.destination = destination
        self.transforms = transforms or []

    def run(self, since: str | None = None) -> dict:
        df = self.source.extract(since)

        for transform in self.transforms:
            df = transform.transform(df)

        rows = self.destination.load(df, "public", "raw_data")

        return {
            "rows_loaded": rows,
            "columns": list(df.columns),
        }

Harness vs Orchestrator

It's important to understand the distinction between a harness and an orchestrator like Airflow or Prefect:

  • Harness - Focuses on the data movement itself: connecting, extracting, transforming, and loading. Handles one pipeline end-to-end.
  • Orchestrator - Manages multiple pipelines, including scheduling, dependencies, and monitoring. Coordinates when and in what order things run.

In practice, you often use both together: an orchestrator triggers harnesses, monitors their completion, and coordinates downstream actions.

Resources

Harnesses

In data engineering, a harness is an integration framework that connects, transforms, and monitors data flowing between systems. Harnesses abstract away the complexity of individual source and destination APIs, providing a unified interface for data movement and transformation.

What is a Data Harness?

A harness sits between your data sources and your data platform, handling:

  • Connector management - Pre-built connectors for common sources (databases, APIs, SaaS platforms)
  • Schema inference and mapping - Automatically detect schemas and map source fields to target schemas
  • Transformation pipelines - Built-in transform logic for cleaning, enriching, and normalizing data
  • Monitoring and observability - Track data volume, freshness, and quality metrics across all pipelines
  • Error handling - Retry logic, dead-letter queues, and alerting on failures

Popular Harness Tools

Fivetran

Fivetran is a fully managed data integration platform. It provides hundreds of pre-built connectors that handle schema changes, incremental updates, and error recovery automatically. Fivetran uses a harness model where you configure sources and destinations through a UI, and the platform handles the rest.

  • Automatic schema drift detection and propagation
  • Incremental updates via CDC or API pagination
  • dbt Cloud integration for post-load transforms
  • History mode for tracking all changes

Airbyte

Airbyte is an open-source data integration platform. It offers a similar harness model to Fivetran but with the flexibility of self-hosting and a large catalog of community connectors.

  • Open-source with a generous ELT2 license
  • Connector development kit (CDK) for building custom connectors
  • Protocol-guaranteed data delivery with at-least-once semantics
  • PyAirbyte for programmatic connector usage

Stitch (by Talend)

Stitch is a simple, reliable ETL service for streaming data from sources to warehouses. It focuses on simplicity and reliability over extensive transformation capabilities.

Building a Custom Harness

For specialized use cases, you might build a lightweight harness using Python and existing libraries:

from typing import Protocol
import pandas as pd

class Source(Protocol):
    def extract(self, since: str | None = None) -> pd.DataFrame: ...

class Destination(Protocol):
    def load(self, df: pd.DataFrame, schema: str, table: str) -> int: ...

class Transformer(Protocol):
    def transform(self, df: pd.DataFrame) -> pd.DataFrame: ...

class DataHarness:
    def __init__(
        self,
        source: Source,
        destination: Destination,
        transforms: list[Transformer] | None = None,
    ):
        self.source = source
        self.destination = destination
        self.transforms = transforms or []

    def run(self, since: str | None = None) -> dict:
        df = self.source.extract(since)

        for transform in self.transforms:
            df = transform.transform(df)

        rows = self.destination.load(df, "public", "raw_data")

        return {
            "rows_loaded": rows,
            "columns": list(df.columns),
        }

Harness vs Orchestrator

It's important to understand the distinction between a harness and an orchestrator like Airflow or Prefect:

  • Harness - Focuses on the data movement itself: connecting, extracting, transforming, and loading. Handles one pipeline end-to-end.
  • Orchestrator - Manages multiple pipelines, including scheduling, dependencies, and monitoring. Coordinates when and in what order things run.

In practice, you often use both together: an orchestrator triggers harnesses, monitors their completion, and coordinates downstream actions.

Resources