DE Wikicompanies / reducto

Reducto

Reducto is an AI-native data processing platform focused on extracting, chunking, and embedding unstructured document data for retrieval-augmented generation (RAG) and AI pipelines. It turns raw documents (PDFs, HTML, images, Office files) into clean, structured, AI-ready data.

Core Product: Document Processing API

Reducto's primary product is a document processing API that handles the entire document-to-embedding pipeline:

1. Ingestion

  • Multi-format parsing - PDF (with OCR fallback), HTML, DOCX, PPTX, XLSX, images, markdown
  • Layout preservation - Maintains reading order, table structure, and heading hierarchy
  • Bulk processing - Batch upload via API or cloud storage integration (S3, GCS)

2. Chunking

Reducto's smart chunking uses document structure (headings, sections, tables) rather than fixed token windows to create semantically meaningful chunks.

// Reducto API: chunk a document
const response = await fetch("https://api.reducto.ai/v1/chunk", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    document_url: "s3://bucket/reports/q4-2024.pdf",
    chunking_strategy: "semantic",
    max_chunk_size: 1000,   // tokens
    overlap: 100,           // token overlap between chunks
    extract_tables: true,
    extract_images: false,
  }),
});

const { chunks } = await response.json();
// Each chunk has: id, text, metadata (page, heading, section), embedding

3. Embedding

  • Supports OpenAI, Cohere, and open-source embedding models (BGE, E5, Instructor)
  • Batching for high-throughput embedding at lower cost
  • Metadata injection into vector database entries

Architecture

Reducto's architecture is built on a microservices model with specialized workers for each processing stage:

  • Parser workers - Run document parsing (PyMuPDF, LibreOffice, Tesseract OCR) in isolated containers
  • Chunk workers - Apply chunking strategies (fixed, semantic, document-aware)
  • Embedding workers - Batch embedding requests to model APIs with caching and rate limiting
  • Queue - SQS/RabbitMQ for async processing with priority queuing
  • Storage - S3 for raw documents and processed chunks; Postgres for job metadata

Case Study: Large-Scale RAG Pipeline

A financial services company used Reducto to process 2 million PDF pages of regulatory filings into a RAG system. Key results:

  • 99.2% extraction accuracy on complex financial tables
  • 3x reduction in chunk count vs fixed-token chunking
  • 42% improvement in retrieval precision (due to better chunk boundaries)
  • Processing throughput of 500 pages/second at peak

Integration Patterns

With Airflow/Prefect

# Airflow DAG for document processing pipeline

from airflow import DAG
from airflow.providers.http.operators.http import HttpOperator
from airflow.operators.python import PythonOperator

with DAG("document_pipeline", ...) as dag:

    chunk_docs = HttpOperator(
        task_id="chunk_documents",
        http_conn_id="reducto_api",
        endpoint="/v1/chunk_batch",
        data={"document_urls": urls, "chunking_strategy": "semantic"},
        method="POST",
    )

    embed_chunks = HttpOperator(
        task_id="embed_chunks",
        http_conn_id="reducto_api",
        endpoint="/v1/embed",
        data={"chunk_ids": "{{ task_instance.xcom_pull('chunk_docs') }}"},
        method="POST",
    )

    load_to_pinecone = PythonOperator(
        task_id="load_to_vector_db",
        python_callable=load_embeddings,
        op_kwargs={
            "chunks": "{{ task_instance.xcom_pull('embed_chunks') }}",
        },
    )

    chunk_docs >> embed_chunks >> load_to_pinecone

Resources

Reducto

Reducto is an AI-native data processing platform focused on extracting, chunking, and embedding unstructured document data for retrieval-augmented generation (RAG) and AI pipelines. It turns raw documents (PDFs, HTML, images, Office files) into clean, structured, AI-ready data.

Core Product: Document Processing API

Reducto's primary product is a document processing API that handles the entire document-to-embedding pipeline:

1. Ingestion

  • Multi-format parsing - PDF (with OCR fallback), HTML, DOCX, PPTX, XLSX, images, markdown
  • Layout preservation - Maintains reading order, table structure, and heading hierarchy
  • Bulk processing - Batch upload via API or cloud storage integration (S3, GCS)

2. Chunking

Reducto's smart chunking uses document structure (headings, sections, tables) rather than fixed token windows to create semantically meaningful chunks.

// Reducto API: chunk a document
const response = await fetch("https://api.reducto.ai/v1/chunk", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    document_url: "s3://bucket/reports/q4-2024.pdf",
    chunking_strategy: "semantic",
    max_chunk_size: 1000,   // tokens
    overlap: 100,           // token overlap between chunks
    extract_tables: true,
    extract_images: false,
  }),
});

const { chunks } = await response.json();
// Each chunk has: id, text, metadata (page, heading, section), embedding

3. Embedding

  • Supports OpenAI, Cohere, and open-source embedding models (BGE, E5, Instructor)
  • Batching for high-throughput embedding at lower cost
  • Metadata injection into vector database entries

Architecture

Reducto's architecture is built on a microservices model with specialized workers for each processing stage:

  • Parser workers - Run document parsing (PyMuPDF, LibreOffice, Tesseract OCR) in isolated containers
  • Chunk workers - Apply chunking strategies (fixed, semantic, document-aware)
  • Embedding workers - Batch embedding requests to model APIs with caching and rate limiting
  • Queue - SQS/RabbitMQ for async processing with priority queuing
  • Storage - S3 for raw documents and processed chunks; Postgres for job metadata

Case Study: Large-Scale RAG Pipeline

A financial services company used Reducto to process 2 million PDF pages of regulatory filings into a RAG system. Key results:

  • 99.2% extraction accuracy on complex financial tables
  • 3x reduction in chunk count vs fixed-token chunking
  • 42% improvement in retrieval precision (due to better chunk boundaries)
  • Processing throughput of 500 pages/second at peak

Integration Patterns

With Airflow/Prefect

# Airflow DAG for document processing pipeline

from airflow import DAG
from airflow.providers.http.operators.http import HttpOperator
from airflow.operators.python import PythonOperator

with DAG("document_pipeline", ...) as dag:

    chunk_docs = HttpOperator(
        task_id="chunk_documents",
        http_conn_id="reducto_api",
        endpoint="/v1/chunk_batch",
        data={"document_urls": urls, "chunking_strategy": "semantic"},
        method="POST",
    )

    embed_chunks = HttpOperator(
        task_id="embed_chunks",
        http_conn_id="reducto_api",
        endpoint="/v1/embed",
        data={"chunk_ids": "{{ task_instance.xcom_pull('chunk_docs') }}"},
        method="POST",
    )

    load_to_pinecone = PythonOperator(
        task_id="load_to_vector_db",
        python_callable=load_embeddings,
        op_kwargs={
            "chunks": "{{ task_instance.xcom_pull('embed_chunks') }}",
        },
    )

    chunk_docs >> embed_chunks >> load_to_pinecone

Resources