DE Wikiglossary

Glossary & FAQ

Definitions for key data engineering terms and answers to common questions about data infrastructure, tools, and architecture patterns.

Glossary

TermDefinition
ACIDA set of database transaction properties: Atomicity, Consistency, Isolation, and Durability. Ensures reliable processing of database transactions even during failures.
Batch ProcessingProcessing data in discrete, scheduled jobs on finite datasets. Contrast with stream processing, which processes data continuously.
CDC (Change Data Capture)A technique for tracking and capturing changes made to a database so they can be replicated to other systems in real time or near-real time.
DAG (Directed Acyclic Graph)A graph with directed edges and no cycles. In data engineering, DAGs represent workflow dependencies where tasks are nodes and edges define execution order.
Data LakeA centralized repository that stores raw data in its native format, typically using object storage (S3, ADLS, GCS). Contrast with data warehouse, which stores structured, processed data.
Data WarehouseA system for storing structured, processed data optimized for analytical queries. Typically uses dimensional modeling (star/snowflake schemas).
Delta LakeAn open-source storage layer that brings ACID transactions, schema enforcement, and time travel to data lakes built on Apache Spark and Parquet.
Dimension TableIn dimensional modeling, a table that stores descriptive attributes (e.g., customer name, product category, date). Connected to fact tables via foreign keys.
ELT (Extract, Load, Transform)A data integration pattern where raw data is loaded directly into the target system and transformed in-place using the warehouse's compute power.
ETL (Extract, Transform, Load)A data integration pattern where data is extracted from sources, transformed in a staging area, and then loaded into the target system.
Exactly-Once SemanticsA processing guarantee that each message is processed exactly one time, preventing both data loss and duplication. Achieved through transactional writes and idempotent sinks.
Fact TableIn dimensional modeling, the central table that stores quantitative measurements (facts) and foreign keys to dimension tables. Often the largest table in a warehouse.
IcebergAn open table format for large analytic datasets. Provides ACID transactions, schema evolution, partition evolution, and time travel on object storage.
IdempotencyThe property that running the same operation multiple times produces the same result. Critical for data pipelines that may retry on failure.
Kappa ArchitectureA data processing architecture that uses a single streaming pipeline for both real-time and batch workloads, simplifying the Lambda architecture pattern.
Lambda ArchitectureA data processing architecture with two parallel pipelines: a speed layer for real-time data and a batch layer for comprehensive, accurate results.
LakehouseAn architecture combining data lake flexibility with warehouse ACID guarantees. Built on open table formats (Delta Lake, Iceberg, Hudi) and object storage.
MCP (Multi-Channel Platform)A unified data ingestion and routing layer handling multiple data channels (streaming, batch, API, CDC) with consistent processing guarantees.
OLAP (Online Analytical Processing)A class of database systems optimized for complex analytical queries and aggregations over large datasets. Typically uses columnar storage and denormalized schemas.
OLTP (Online Transaction Processing)A class of database systems optimized for handling many small, concurrent transactions. Typically uses normalized schemas and row-oriented storage.
OrchestratorA system that manages, schedules, and monitors data workflows. Examples: Apache Airflow, Prefect, Dagster. Handles task dependencies, retries, and logging.
ParquetAn open-source columnar storage format for Hadoop and Spark. Provides efficient compression and encoding schemes, making it ideal for analytical workloads.
PartitioningDividing a dataset into smaller, manageable chunks based on a partition key (e.g., date). Enables partition pruning for faster queries and incremental processing.
RAG (Retrieval-Augmented Generation)An AI pattern that combines information retrieval with text generation. Documents are chunked, embedded, and stored in a vector database for retrieval at inference time.
SCD (Slowly Changing Dimension)A dimension that changes slowly over time. SCD Type 1 overwrites the old value; Type 2 tracks history by adding new rows; Type 3 uses a previous-value column.
Schema EvolutionThe ability to change a table's schema (add, drop, rename columns) without rewriting existing data. Supported by Delta Lake, Iceberg, and Hudi.
Stream ProcessingProcessing data continuously as it arrives, rather than in fixed batches. Tools: Apache Kafka, Flink, Spark Streaming, Kafka Streams.
Time TravelA feature of table formats (Delta Lake, Iceberg) that allows querying and restoring data as of a specific version or timestamp.

Frequently Asked Questions

What is the difference between a data lake and a data warehouse?

A data lake stores raw data in its native format (files, JSON, Parquet) on object storage, while a data warehouse stores structured, transformed data optimized for SQL analytics. Data lakes are flexible and cheap; warehouses are performant and governed. The lakehouse pattern combines both.

When should I use Airflow vs Prefect vs Dagster?

Airflow is the most mature with the largest ecosystem, best for organizations that need extensive operator support and existing community knowledge. Prefect offers a smoother developer experience with Python-native flows and automatic retries. Dagster is best for asset-oriented pipelines where data quality and lineage are top priorities. All three are valid; choose based on your team's workflow preference and requirements.

What is the difference between ETL and ELT?

ETL transforms data before loading into the target system, using a dedicated transformation server. ELT loads raw data first, then transforms it in the warehouse. ELT is preferred for cloud warehouses with elastic compute (Snowflake, BigQuery); ETL is better when the target has limited compute or strict governance requirements.

What is exactly-once processing and why does it matter?

Exactly-once processing guarantees that each message is processed exactly one time, preventing both data loss and duplication. It matters for accuracy-sensitive applications like financial reconciliation, fraud detection, and inventory management. It's harder to achieve than at-least-once (which allows duplicates) or at-most-once (which may drop data).

What is a data mesh and when should I adopt it?

Data mesh is a decentralized architecture that organizes data by business domain, with each domain owning its data as a product. Adopt it when your organization has 5+ distinct data domains, a central data team is becoming a bottleneck, and domain teams have the skills to manage their own data. It requires significant cultural and platform investment.

What is the difference between Lambda and Kappa architecture?

Lambda architecture uses two separate pipelines (batch and streaming) that produce results merged in a serving layer. Kappa architecture uses a single streaming pipeline for everything - batch is handled by replaying the stream from the beginning. Kappa is simpler but requires the streaming platform to retain the full event history.

What is a harness in data engineering?

A harness is an integration framework that connects data sources to destinations, handling connector management, schema inference, transformation, error handling, and monitoring. Examples: Fivetran, Airbyte. Harnesses handle data movement; orchestrators (Airflow, Prefect) manage when and in what order things run.

What is an MCP (Multi-Channel Platform)?

A Multi-Channel Platform unifies data ingestion across streaming, batch, API, and CDC channels into a single routing and processing layer. It handles schema normalization, exactly-once deduplication, and channel prioritization. Kafka + Connect is the most common MCP foundation.

What is the difference between batch and stream processing?

Batch processes finite datasets on a schedule (hourly, daily), trading latency for completeness and accuracy. Stream processes data continuously as it arrives, providing low-latency results that may be approximate. Many systems use both: stream for real-time dashboards and batch for accurate historical reporting.

How do I choose between Delta Lake, Iceberg, and Hudi?

Delta Lake is the most mature with the best Spark integration and is the default on Databricks. Iceberg has the broadest engine support (Spark, Flink, Trino, Presto, Hive) with a strong community. Hudi specializes in incremental processing and upsert-heavy workloads. All are open source and provide ACID on data lakes.

Still Have Questions?

This wiki is maintained as a living document. If there's a term or question you'd like to see covered, each guide page includes links to external documentation and community resources for further exploration.

Glossary & FAQ

Definitions for key data engineering terms and answers to common questions about data infrastructure, tools, and architecture patterns.

Glossary

TermDefinition
ACIDA set of database transaction properties: Atomicity, Consistency, Isolation, and Durability. Ensures reliable processing of database transactions even during failures.
Batch ProcessingProcessing data in discrete, scheduled jobs on finite datasets. Contrast with stream processing, which processes data continuously.
CDC (Change Data Capture)A technique for tracking and capturing changes made to a database so they can be replicated to other systems in real time or near-real time.
DAG (Directed Acyclic Graph)A graph with directed edges and no cycles. In data engineering, DAGs represent workflow dependencies where tasks are nodes and edges define execution order.
Data LakeA centralized repository that stores raw data in its native format, typically using object storage (S3, ADLS, GCS). Contrast with data warehouse, which stores structured, processed data.
Data WarehouseA system for storing structured, processed data optimized for analytical queries. Typically uses dimensional modeling (star/snowflake schemas).
Delta LakeAn open-source storage layer that brings ACID transactions, schema enforcement, and time travel to data lakes built on Apache Spark and Parquet.
Dimension TableIn dimensional modeling, a table that stores descriptive attributes (e.g., customer name, product category, date). Connected to fact tables via foreign keys.
ELT (Extract, Load, Transform)A data integration pattern where raw data is loaded directly into the target system and transformed in-place using the warehouse's compute power.
ETL (Extract, Transform, Load)A data integration pattern where data is extracted from sources, transformed in a staging area, and then loaded into the target system.
Exactly-Once SemanticsA processing guarantee that each message is processed exactly one time, preventing both data loss and duplication. Achieved through transactional writes and idempotent sinks.
Fact TableIn dimensional modeling, the central table that stores quantitative measurements (facts) and foreign keys to dimension tables. Often the largest table in a warehouse.
IcebergAn open table format for large analytic datasets. Provides ACID transactions, schema evolution, partition evolution, and time travel on object storage.
IdempotencyThe property that running the same operation multiple times produces the same result. Critical for data pipelines that may retry on failure.
Kappa ArchitectureA data processing architecture that uses a single streaming pipeline for both real-time and batch workloads, simplifying the Lambda architecture pattern.
Lambda ArchitectureA data processing architecture with two parallel pipelines: a speed layer for real-time data and a batch layer for comprehensive, accurate results.
LakehouseAn architecture combining data lake flexibility with warehouse ACID guarantees. Built on open table formats (Delta Lake, Iceberg, Hudi) and object storage.
MCP (Multi-Channel Platform)A unified data ingestion and routing layer handling multiple data channels (streaming, batch, API, CDC) with consistent processing guarantees.
OLAP (Online Analytical Processing)A class of database systems optimized for complex analytical queries and aggregations over large datasets. Typically uses columnar storage and denormalized schemas.
OLTP (Online Transaction Processing)A class of database systems optimized for handling many small, concurrent transactions. Typically uses normalized schemas and row-oriented storage.
OrchestratorA system that manages, schedules, and monitors data workflows. Examples: Apache Airflow, Prefect, Dagster. Handles task dependencies, retries, and logging.
ParquetAn open-source columnar storage format for Hadoop and Spark. Provides efficient compression and encoding schemes, making it ideal for analytical workloads.
PartitioningDividing a dataset into smaller, manageable chunks based on a partition key (e.g., date). Enables partition pruning for faster queries and incremental processing.
RAG (Retrieval-Augmented Generation)An AI pattern that combines information retrieval with text generation. Documents are chunked, embedded, and stored in a vector database for retrieval at inference time.
SCD (Slowly Changing Dimension)A dimension that changes slowly over time. SCD Type 1 overwrites the old value; Type 2 tracks history by adding new rows; Type 3 uses a previous-value column.
Schema EvolutionThe ability to change a table's schema (add, drop, rename columns) without rewriting existing data. Supported by Delta Lake, Iceberg, and Hudi.
Stream ProcessingProcessing data continuously as it arrives, rather than in fixed batches. Tools: Apache Kafka, Flink, Spark Streaming, Kafka Streams.
Time TravelA feature of table formats (Delta Lake, Iceberg) that allows querying and restoring data as of a specific version or timestamp.

Frequently Asked Questions

What is the difference between a data lake and a data warehouse?

A data lake stores raw data in its native format (files, JSON, Parquet) on object storage, while a data warehouse stores structured, transformed data optimized for SQL analytics. Data lakes are flexible and cheap; warehouses are performant and governed. The lakehouse pattern combines both.

When should I use Airflow vs Prefect vs Dagster?

Airflow is the most mature with the largest ecosystem, best for organizations that need extensive operator support and existing community knowledge. Prefect offers a smoother developer experience with Python-native flows and automatic retries. Dagster is best for asset-oriented pipelines where data quality and lineage are top priorities. All three are valid; choose based on your team's workflow preference and requirements.

What is the difference between ETL and ELT?

ETL transforms data before loading into the target system, using a dedicated transformation server. ELT loads raw data first, then transforms it in the warehouse. ELT is preferred for cloud warehouses with elastic compute (Snowflake, BigQuery); ETL is better when the target has limited compute or strict governance requirements.

What is exactly-once processing and why does it matter?

Exactly-once processing guarantees that each message is processed exactly one time, preventing both data loss and duplication. It matters for accuracy-sensitive applications like financial reconciliation, fraud detection, and inventory management. It's harder to achieve than at-least-once (which allows duplicates) or at-most-once (which may drop data).

What is a data mesh and when should I adopt it?

Data mesh is a decentralized architecture that organizes data by business domain, with each domain owning its data as a product. Adopt it when your organization has 5+ distinct data domains, a central data team is becoming a bottleneck, and domain teams have the skills to manage their own data. It requires significant cultural and platform investment.

What is the difference between Lambda and Kappa architecture?

Lambda architecture uses two separate pipelines (batch and streaming) that produce results merged in a serving layer. Kappa architecture uses a single streaming pipeline for everything - batch is handled by replaying the stream from the beginning. Kappa is simpler but requires the streaming platform to retain the full event history.

What is a harness in data engineering?

A harness is an integration framework that connects data sources to destinations, handling connector management, schema inference, transformation, error handling, and monitoring. Examples: Fivetran, Airbyte. Harnesses handle data movement; orchestrators (Airflow, Prefect) manage when and in what order things run.

What is an MCP (Multi-Channel Platform)?

A Multi-Channel Platform unifies data ingestion across streaming, batch, API, and CDC channels into a single routing and processing layer. It handles schema normalization, exactly-once deduplication, and channel prioritization. Kafka + Connect is the most common MCP foundation.

What is the difference between batch and stream processing?

Batch processes finite datasets on a schedule (hourly, daily), trading latency for completeness and accuracy. Stream processes data continuously as it arrives, providing low-latency results that may be approximate. Many systems use both: stream for real-time dashboards and batch for accurate historical reporting.

How do I choose between Delta Lake, Iceberg, and Hudi?

Delta Lake is the most mature with the best Spark integration and is the default on Databricks. Iceberg has the broadest engine support (Spark, Flink, Trino, Presto, Hive) with a strong community. Hudi specializes in incremental processing and upsert-heavy workloads. All are open source and provide ACID on data lakes.

Still Have Questions?

This wiki is maintained as a living document. If there's a term or question you'd like to see covered, each guide page includes links to external documentation and community resources for further exploration.