Lambda Architecture
Lambda architecture is a data processing pattern that maintains two parallel pipelines: a speed layer (streaming) for real-time data and a batch layer for comprehensive, accurate results. A serving layer merges outputs from both for querying.
Architecture Layers
1. Batch Layer
The batch layer stores all incoming raw data and computes comprehensive views. It provides the ground truth: accurate, complete results at the cost of higher latency (minutes to hours).
- Stores the complete immutable dataset (master dataset)
- Runs periodic batch jobs (Spark, Hive) to compute batch views
- Results are accurate but delayed
2. Speed Layer
The speed layer processes data in real-time to compensate for the batch layer's latency. It provides approximate results immediately using stream processing.
- Processes data as it arrives (Kafka, Flink, Spark Streaming)
- Produces real-time views with low latency (seconds)
- Results may be approximate (trade accuracy for speed)
3. Serving Layer
The serving layer merges batch views and real-time views, presenting a unified query interface. When batch results become available, they replace the corresponding speed layer results.
Implementation Example
-- Batch layer: hourly aggregation (Spark)
# spark_batch.py
def compute_hourly_metrics():
df = spark.read.parquet("s3://data-lake/raw-events/")
hourly = (
df.groupBy("event_type", window("event_time", "1 hour"))
.agg(count("*").alias("count"), sum("value").alias("total"))
)
hourly.write.mode("overwrite").parquet(
"s3://data-lake/batch-views/events_hourly/"
)
-- Speed layer: real-time aggregation (Spark Streaming)
# spark_streaming.py
stream = (
spark
.readStream
.format("kafka")
.option("subscribe", "events")
.load()
.selectExpr("CAST(value AS STRING) as json")
.select(from_json("json", schema).alias("data"))
.select("data.*")
)
real_time = (
stream
.groupBy("event_type",
window("event_time", "5 minutes", "1 minute"))
.agg(count("*").alias("count"), sum("value").alias("total"))
)
# Write speed layer results
real_time.writeStream .outputMode("complete") .option("checkpointLocation", "s3://checkpoints/speed/") .table("speed_layer.events_5min") .start()
-- Serving layer: merge batch and speed (Presto/Trino)
SELECT
event_type,
window_start,
total
FROM batch_views.events_hourly
UNION ALL
SELECT
event_type,
window_start,
total
FROM speed_layer.events_5min
WHERE window_start >= NOW() - INTERVAL '1' HOUR;When to Use Lambda Architecture
- Your application needs both real-time and historical analytics
- Streaming approximations are acceptable for current data
- You need reprocessing capabilities (batch layer can recompute from raw data)
- You cannot afford to lose data (batch layer serves as backup)
Criticism and Alternatives
Lambda architecture has been criticized for maintaining two separate codebases (batch and speed) that must produce the same result. This leads to:
- Code duplication and increased maintenance burden
- Subtle differences between batch and streaming results
- Complex deployment and debugging
These issues led to the development of the Kappa architecture, which uses a single streaming pipeline for both real-time and batch workloads.