# How Pathway's Multimodal RAG Pipeline Extracts and Indexes Tables and Charts from PDFs Using GPT-4o

> Discover how Pathway's multimodal RAG pipeline leverages GPT-4o to extract and index tables and charts from PDFs. Learn about its advanced data extraction and retrieval capabilities.

- Repository: [Pathway/llm-app](https://github.com/pathwaycom/llm-app)
- Tags: how-to-guide
- Published: 2026-03-07

---

**Pathway's multimodal RAG pipeline uses GPT-4o's vision capabilities within a DoclingParser component to rasterize PDF pages, extract structured data from tables and charts, and convert visual elements into searchable text chunks that feed into a vector index for precise retrieval-augmented generation.**

The `pathwaycom/llm-app` repository provides a production-ready implementation of this architecture in the `multimodal_rag` template. Unlike traditional RAG systems that only process extracted text, this **Pathway multimodal RAG pipeline** treats visual data—tables, charts, and figures—as first-class citizens by leveraging GPT-4o's vision-language capabilities during the ingestion phase.

## Four-Stage Pipeline Architecture

The system operates through four tightly coupled stages defined in [`templates/multimodal_rag/app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/templates/multimodal_rag/app.yaml). The architecture diagram in the template's README highlights the two critical GPT-4o touchpoints: table parsing and image parsing.

### Stage 1: PDF Data Ingestion

The pipeline monitors a local directory for new files. In [`app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/app.yaml), the `$sources` block (lines 12-18) configures the `!pw.io.fs.read` operator to ingest binary PDF files from the `data/` folder. This operator streams raw PDF bytes into the Pathway runtime, triggering downstream processing whenever new documents appear.

### Stage 2: Multimodal Parsing with GPT-4o

The **DoclingParser** (`$parser` block, lines 72-78) handles the core extraction. It identifies table and image regions within PDF pages and routes them to a vision-language model for interpretation:

```yaml
parser:
  multimodal_llm: $parsing_llm
  image_parsing_strategy: "llm"
  table_parsing_strategy: "llm"

```

The `$parsing_llm` configuration (lines 64-70) specifies `model: "gpt-4o"`. GPT-4o receives rasterized page images, applies visual reasoning to extract structured table data and chart descriptions, and returns JSON-like representations. The pipeline treats these outputs as standard text chunks, making previously inaccessible visual data fully searchable.

### Stage 3: Vector Embedding and Indexing

Parsed chunks flow through a **TokenCountSplitter** (optional) before reaching the **OpenAIEmbedder** (`$embedder`), which generates dense vectors using `text-embedding-3-small`. These vectors populate a **UsearchKnnFactory** index (`$retriever_factory`, lines 85-90), enabling efficient K-nearest neighbors search.

The **DocumentStore** (`$document_store`, lines 92-96) orchestrates this flow, binding sources, parser, splitter, and retriever into a unified ingestion graph that automatically re-indexes when source files change.

### Stage 4: Retrieval-Augmented Question Answering

The **BaseRAGQuestionAnswerer** (`question_answerer`, lines 99-104) handles query execution. It retrieves top-k relevant chunks from the USearch index and passes them to an answer generation LLM (`$llm`, typically configured as "gpt-4.1-mini"). Because GPT-4o already converted tables and charts into rich textual context during ingestion, the answer model can cite precise numeric values from the original PDFs, such as specific financial figures from quarterly reports.

## Running the Pipeline Locally

To execute this **multimodal RAG** system against your own documents:

```bash

# Install dependencies

pip install -r templates/multimodal_rag/requirements.txt -U

# Configure OpenAI API key

echo "OPENAI_API_KEY=sk-..." > templates/multimodal_rag/.env

# Launch the service

python templates/multimodal_rag/app.py

```

The entry point ([`templates/multimodal_rag/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/multimodal_rag/app.py)) instantiates a `QASummaryRestServer` (lines 34-38) and starts Pathway's persistent runtime via `pw.run` (lines 63-67), exposing REST endpoints for real-time interaction.

Once running, query the indexed documents:

```bash

# List available documents

curl -X POST http://0.0.0.0:8000/v2/list_documents \
     -H "accept: */*" -H "Content-Type: application/json"

# Query specific table values

curl -X POST http://0.0.0.0:8000/v2/answer \
     -H "accept: */*" -H "Content-Type: application/json" \
     -d '{"prompt":"How much was Operating lease cost in 2021?"}'

```

The response contains accurate figures extracted from tables because GPT-4o already processed the visual data into searchable text during the indexing phase.

For containerized deployment:

```bash
cd templates/multimodal_rag
docker build -t pathway-multimodal-rag .
docker run -v "$(pwd)/data:/app/data" -p 8000:8000 pathway-multimodal-rag

```

## Summary

- **Pathway's multimodal RAG pipeline** ingests binary PDFs via `pw.io.fs.read` configured in [`app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/app.yaml).
- **GPT-4o** serves as the vision parser within `DoclingParser`, extracting structured data from tables and charts using `image_parsing_strategy: "llm"` and `table_parsing_strategy: "llm"`.
- Parsed content embeds via **OpenAIEmbedder** (using `text-embedding-3-small`) and indexes in **UsearchKnnFactory** for high-performance vector retrieval.
- The **BaseRAGQuestionAnswerer** retrieves these rich multimodal chunks to answer precise questions about visual data, such as financial metrics from the sample `20230203_alphabet_10K.pdf`.

## Frequently Asked Questions

### What is the difference between the parsing LLM and the answer LLM?

The parsing LLM (`$parsing_llm`, configured as GPT-4o) operates during document ingestion to convert visual PDF elements into structured text. The answer LLM (`$llm`, often a lighter chat model) generates final responses using the retrieved context. This separation allows the pipeline to use expensive vision capabilities only during indexing while keeping per-query latency and costs low.

### How does GPT-4o handle complex table layouts that traditional OCR might misparse?

According to the [`app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/app.yaml) configuration, GPT-4o receives rasterized page images rather than raw OCR text. With `table_parsing_strategy: "llm"`, the model applies visual reasoning to understand spatial relationships, merged cells, and multi-row headers, outputting structured JSON-like representations that preserve the semantic hierarchy of the original table.

### Why does the pipeline use USearch specifically for the vector index?

The **UsearchKnnFactory** (`$retriever_factory`, lines 85-90) provides an efficient K-nearest neighbors implementation optimized for the dense vectors generated by `text-embedding-3-small`. This component stores embeddings produced by the **OpenAIEmbedder** and enables fast similarity search during the retrieval phase of the RAG cycle.

### Can I substitute GPT-4o with other multimodal models for the parsing stage?

While the [`app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/app.yaml) template explicitly sets `model: "gpt-4o"` in the `$parsing_llm` block (lines 64-70), the variable-based architecture (`$parsing_llm`) suggests the parser accepts any Pathway-compatible LLM connector. To use alternative vision models, you would define a new LLM configuration with appropriate image input support and reference it in the `$parser.multimodal_llm` field.