Harnesses
In data engineering, a harness is an integration framework that connects, transforms, and monitors data flowing between systems. Harnesses abstract away the complexity of individual source and destination APIs, providing a unified interface for data movement and transformation.
What is a Data Harness?
A harness sits between your data sources and your data platform, handling:
- Connector management - Pre-built connectors for common sources (databases, APIs, SaaS platforms)
- Schema inference and mapping - Automatically detect schemas and map source fields to target schemas
- Transformation pipelines - Built-in transform logic for cleaning, enriching, and normalizing data
- Monitoring and observability - Track data volume, freshness, and quality metrics across all pipelines
- Error handling - Retry logic, dead-letter queues, and alerting on failures
Popular Harness Tools
Fivetran
Fivetran is a fully managed data integration platform. It provides hundreds of pre-built connectors that handle schema changes, incremental updates, and error recovery automatically. Fivetran uses a harness model where you configure sources and destinations through a UI, and the platform handles the rest.
- Automatic schema drift detection and propagation
- Incremental updates via CDC or API pagination
- dbt Cloud integration for post-load transforms
- History mode for tracking all changes
Airbyte
Airbyte is an open-source data integration platform. It offers a similar harness model to Fivetran but with the flexibility of self-hosting and a large catalog of community connectors.
- Open-source with a generous ELT2 license
- Connector development kit (CDK) for building custom connectors
- Protocol-guaranteed data delivery with at-least-once semantics
- PyAirbyte for programmatic connector usage
Stitch (by Talend)
Stitch is a simple, reliable ETL service for streaming data from sources to warehouses. It focuses on simplicity and reliability over extensive transformation capabilities.
Building a Custom Harness
For specialized use cases, you might build a lightweight harness using Python and existing libraries:
from typing import Protocol
import pandas as pd
class Source(Protocol):
def extract(self, since: str | None = None) -> pd.DataFrame: ...
class Destination(Protocol):
def load(self, df: pd.DataFrame, schema: str, table: str) -> int: ...
class Transformer(Protocol):
def transform(self, df: pd.DataFrame) -> pd.DataFrame: ...
class DataHarness:
def __init__(
self,
source: Source,
destination: Destination,
transforms: list[Transformer] | None = None,
):
self.source = source
self.destination = destination
self.transforms = transforms or []
def run(self, since: str | None = None) -> dict:
df = self.source.extract(since)
for transform in self.transforms:
df = transform.transform(df)
rows = self.destination.load(df, "public", "raw_data")
return {
"rows_loaded": rows,
"columns": list(df.columns),
}Harness vs Orchestrator
It's important to understand the distinction between a harness and an orchestrator like Airflow or Prefect:
- Harness - Focuses on the data movement itself: connecting, extracting, transforming, and loading. Handles one pipeline end-to-end.
- Orchestrator - Manages multiple pipelines, including scheduling, dependencies, and monitoring. Coordinates when and in what order things run.
In practice, you often use both together: an orchestrator triggers harnesses, monitors their completion, and coordinates downstream actions.