How Pathway's Multimodal RAG Pipeline Extracts and Indexes Tables and Charts from PDFs Using GPT-4o

Pathway's multimodal RAG pipeline uses GPT-4o's vision capabilities within a DoclingParser component to rasterize PDF pages, extract structured data from tables and charts, and convert visual elements into searchable text chunks that feed into a vector index for precise retrieval-augmented generation.

The pathwaycom/llm-app repository provides a production-ready implementation of this architecture in the multimodal_rag template. Unlike traditional RAG systems that only process extracted text, this Pathway multimodal RAG pipeline treats visual data—tables, charts, and figures—as first-class citizens by leveraging GPT-4o's vision-language capabilities during the ingestion phase.

Four-Stage Pipeline Architecture

The system operates through four tightly coupled stages defined in templates/multimodal_rag/app.yaml. The architecture diagram in the template's README highlights the two critical GPT-4o touchpoints: table parsing and image parsing.

Stage 1: PDF Data Ingestion

The pipeline monitors a local directory for new files. In app.yaml, the $sources block (lines 12-18) configures the !pw.io.fs.read operator to ingest binary PDF files from the data/ folder. This operator streams raw PDF bytes into the Pathway runtime, triggering downstream processing whenever new documents appear.

Stage 2: Multimodal Parsing with GPT-4o

The DoclingParser ($parser block, lines 72-78) handles the core extraction. It identifies table and image regions within PDF pages and routes them to a vision-language model for interpretation:

parser:
  multimodal_llm: $parsing_llm
  image_parsing_strategy: "llm"
  table_parsing_strategy: "llm"

The $parsing_llm configuration (lines 64-70) specifies model: "gpt-4o". GPT-4o receives rasterized page images, applies visual reasoning to extract structured table data and chart descriptions, and returns JSON-like representations. The pipeline treats these outputs as standard text chunks, making previously inaccessible visual data fully searchable.

Stage 3: Vector Embedding and Indexing

Parsed chunks flow through a TokenCountSplitter (optional) before reaching the OpenAIEmbedder ($embedder), which generates dense vectors using text-embedding-3-small. These vectors populate a UsearchKnnFactory index ($retriever_factory, lines 85-90), enabling efficient K-nearest neighbors search.

The DocumentStore ($document_store, lines 92-96) orchestrates this flow, binding sources, parser, splitter, and retriever into a unified ingestion graph that automatically re-indexes when source files change.

Stage 4: Retrieval-Augmented Question Answering

The BaseRAGQuestionAnswerer (question_answerer, lines 99-104) handles query execution. It retrieves top-k relevant chunks from the USearch index and passes them to an answer generation LLM ($llm, typically configured as "gpt-4.1-mini"). Because GPT-4o already converted tables and charts into rich textual context during ingestion, the answer model can cite precise numeric values from the original PDFs, such as specific financial figures from quarterly reports.

Running the Pipeline Locally

To execute this multimodal RAG system against your own documents:


# Install dependencies

pip install -r templates/multimodal_rag/requirements.txt -U

# Configure OpenAI API key

echo "OPENAI_API_KEY=sk-..." > templates/multimodal_rag/.env

# Launch the service

python templates/multimodal_rag/app.py

The entry point (templates/multimodal_rag/app.py) instantiates a QASummaryRestServer (lines 34-38) and starts Pathway's persistent runtime via pw.run (lines 63-67), exposing REST endpoints for real-time interaction.

Once running, query the indexed documents:


# List available documents

curl -X POST http://0.0.0.0:8000/v2/list_documents \
     -H "accept: */*" -H "Content-Type: application/json"

# Query specific table values

curl -X POST http://0.0.0.0:8000/v2/answer \
     -H "accept: */*" -H "Content-Type: application/json" \
     -d '{"prompt":"How much was Operating lease cost in 2021?"}'

The response contains accurate figures extracted from tables because GPT-4o already processed the visual data into searchable text during the indexing phase.

For containerized deployment:

cd templates/multimodal_rag
docker build -t pathway-multimodal-rag .
docker run -v "$(pwd)/data:/app/data" -p 8000:8000 pathway-multimodal-rag

Summary

  • Pathway's multimodal RAG pipeline ingests binary PDFs via pw.io.fs.read configured in app.yaml.
  • GPT-4o serves as the vision parser within DoclingParser, extracting structured data from tables and charts using image_parsing_strategy: "llm" and table_parsing_strategy: "llm".
  • Parsed content embeds via OpenAIEmbedder (using text-embedding-3-small) and indexes in UsearchKnnFactory for high-performance vector retrieval.
  • The BaseRAGQuestionAnswerer retrieves these rich multimodal chunks to answer precise questions about visual data, such as financial metrics from the sample 20230203_alphabet_10K.pdf.

Frequently Asked Questions

What is the difference between the parsing LLM and the answer LLM?

The parsing LLM ($parsing_llm, configured as GPT-4o) operates during document ingestion to convert visual PDF elements into structured text. The answer LLM ($llm, often a lighter chat model) generates final responses using the retrieved context. This separation allows the pipeline to use expensive vision capabilities only during indexing while keeping per-query latency and costs low.

How does GPT-4o handle complex table layouts that traditional OCR might misparse?

According to the app.yaml configuration, GPT-4o receives rasterized page images rather than raw OCR text. With table_parsing_strategy: "llm", the model applies visual reasoning to understand spatial relationships, merged cells, and multi-row headers, outputting structured JSON-like representations that preserve the semantic hierarchy of the original table.

Why does the pipeline use USearch specifically for the vector index?

The UsearchKnnFactory ($retriever_factory, lines 85-90) provides an efficient K-nearest neighbors implementation optimized for the dense vectors generated by text-embedding-3-small. This component stores embeddings produced by the OpenAIEmbedder and enables fast similarity search during the retrieval phase of the RAG cycle.

Can I substitute GPT-4o with other multimodal models for the parsing stage?

While the app.yaml template explicitly sets model: "gpt-4o" in the $parsing_llm block (lines 64-70), the variable-based architecture ($parsing_llm) suggests the parser accepts any Pathway-compatible LLM connector. To use alternative vision models, you would define a new LLM configuration with appropriate image input support and reference it in the $parser.multimodal_llm field.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →