# TextFlow Two-Stage Architecture vs End-to-End VLM: A Technical Deep Dive for Flowchart Understanding

> Compare TextFlow two-stage architecture with end-to-end VLMs for flowchart understanding. Learn how TextFlow separates visual extraction and logical reasoning for precise results.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: deep-dive
- Published: 2026-03-05

---

**TextFlow's two-stage architecture explicitly separates visual structure extraction from logical reasoning, converting flowchart images into deterministic graph representations before applying language model inference, while end-to-end vision-language models attempt to simultaneously perceive and reason over raw pixels in a single forward pass.**

The `junyiye/textflow` repository implements a specialized **two-stage architecture for flowchart understanding** that challenges the monolithic design of modern vision-language models. Unlike end-to-end approaches that feed raw images directly into transformer stacks, TextFlow first extracts an explicit symbolic graph using the `Textualizer` class, then performs targeted reasoning via the `Reasoner` module.

## Understanding TextFlow's Two-Stage Pipeline

TextFlow processes flowcharts through a deterministic pipeline that decouples perception from cognition. This separation allows for explicit error checking and modular component upgrades without retraining entire systems.

### Stage 1: Structured Extraction with the Textualizer

The first stage converts raw flowchart images into a formal graph representation. In [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), the `Textualizer` class orchestrates OCR and shape detection to identify nodes and edges.

The output is a `Flowchart` object defined in [`src/flowchart.py`](https://github.com/junyiye/textflow/blob/main/src/flowchart.py), which stores the diagram as a structured graph with explicit node identifiers and edge relationships. This object provides deterministic methods like `to_dict()` that serialize the flowchart into JSON format for inspection or downstream processing.

```python
from textualizer import Textualizer
from flowchart import Flowchart

# Convert image to explicit graph structure

txtzr = Textualizer()
graph = txtzr.image_to_flowchart("data/flowvqa/images/instruct00180.png")

# graph is a Flowchart object with deterministic nodes/edges

# Verify intermediate representation

json_repr = graph.to_dict()
print(json_repr)  # {"nodes": {"A": {"id":"A", "description":"Start", ...}, ...}}

```

### Stage 2: Graph-Based Reasoning and VQA

The second stage operates on the extracted graph rather than pixels. In [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), the `Reasoner` class performs structural operations such as shortest path calculation, successor identification, and connectivity checks.

The `VQA` module in [`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py) orchestrates this process, feeding graph traversals to language models via [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py). Because the visual parsing is already complete, the language model receives clean symbolic input, reducing the risk of hallucinations caused by misinterpreting visual layouts.

```python
from reasoner import Reasoner

# Initialize reasoner with extracted graph

reasoner = Reasoner(graph)

# Answer structural questions using graph traversal

answer = reasoner.answer_question(
    "What is the shortest path length from 'Start' to 'Decision'?"
)
print(answer)  # "2"

```

## How End-to-End VLM Approaches Handle Flowcharts

End-to-end vision-language models, such as GPT-4V or LLaVA, process flowchart understanding as a single multimodal task. These systems feed raw bitmaps directly into transformer architectures that simultaneously handle visual feature extraction and linguistic reasoning.

In this paradigm, the model must implicitly learn to detect shapes, read text via OCR, identify connecting lines, and perform logical reasoning within a single forward pass. There is no explicit intermediate graph representation; instead, structural relationships are encoded as latent activations distributed across the model's parameters.

## TextFlow Two-Stage Architecture vs End-to-End VLM: Critical Differences

The architectural divergence between TextFlow and monolithic VLMs creates distinct operational characteristics for flowchart understanding tasks.

### Modularity and Debugging

TextFlow's separation of concerns allows independent optimization of visual extraction and reasoning components. Developers can inspect the intermediate JSON output from `src/flowchart.py::to_dict()` to verify that node detection succeeded before any language model calls occur. End-to-end VLMs offer no such inspection points; errors in visual parsing are inseparable from reasoning failures.

### Error Propagation and Correction

In TextFlow, Stage 1 errors (e.g., missed edges) are visible in the explicit graph structure and can be corrected via rule-based post-processing before Stage 2 executes. End-to-end systems propagate visual misinterpretations directly into the reasoning pathway, often resulting in confident hallucinations about non-existent connections.

### Data Efficiency and Token Usage

TextFlow's symbolic graph provides a concise description of flowchart structure, requiring significantly fewer tokens for the language model to comprehend logical relationships. End-to-end VLMs must process high-resolution image tokens alongside text, demanding larger context windows and more training data to achieve comparable structural understanding.

## Implementation Walkthrough: Key Source Files

The `junyiye/textflow` repository organizes its two-stage pipeline across specialized modules that handle distinct aspects of flowchart understanding.

| File | Role | Key Components |
|------|------|----------------|
| [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) | Stage 1 implementation | `Textualizer` class, `image_to_flowchart()` method |
| [`src/flowchart.py`](https://github.com/junyiye/textflow/blob/main/src/flowchart.py) | Graph representation | `Flowchart` class, `to_dict()` serialization |
| [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) | Stage 2 reasoning | `Reasoner` class, graph traversal algorithms |
| [`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py) | Pipeline orchestration | VQA module coordinating extraction and reasoning |
| [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py) | LLM interface | API wrappers for language model calls |
| [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py) | Backend management | Model loading for local or remote LLMs |

## Summary

TextFlow's two-stage architecture offers a deterministic alternative to end-to-end vision-language models for flowchart understanding tasks. Key advantages include:

- **Explicit graph representation**: The `Textualizer` converts images to structured `Flowchart` objects before reasoning begins
- **Modular error handling**: Visual parsing errors in Stage 1 can be detected and corrected via `to_dict()` inspection before Stage 2 executes
- **Efficient reasoning**: The `Reasoner` operates on compact symbolic graphs rather than high-resolution pixel tensors, reducing token consumption and hallucination risks
- **Component independence**: Each stage can be upgraded independently (e.g., swapping OCR engines or LLM back-ends) without retraining the entire pipeline

## Frequently Asked Questions

### What is the main advantage of TextFlow's two-stage architecture over end-to-end VLMs?

The primary advantage is **error visibility and correction**. TextFlow's Stage 1 produces an explicit graph representation via [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) that developers can inspect using `Flowchart.to_dict()`. If the OCR misses a node or an edge is misidentified, you can apply rule-based corrections before Stage 2 reasoning begins. End-to-end VLMs lack this inspection point; visual errors propagate directly into the reasoning layer, often producing confident but incorrect answers about flowchart structure.

### How does the Textualizer convert flowchart images into structured data?

The `Textualizer` class in [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) orchestrates a pipeline of computer vision techniques including OCR and shape detection to identify flowchart elements. It processes the input image and constructs a `Flowchart` object (defined in [`src/flowchart.py`](https://github.com/junyiye/textflow/blob/main/src/flowchart.py)) containing explicit node and edge representations. This object provides methods like `to_dict()` to serialize the graph into JSON format, creating a deterministic intermediate representation that Stage 2 components consume for reasoning tasks.

### Can TextFlow's reasoning components work with different language models?

Yes, the architecture decouples the reasoning layer from specific LLM implementations. The [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py) file provides wrappers for various language model APIs, while [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py) handles the loading of both local and remote back-ends (such as GPT-4 or Claude). The `Reasoner` class in [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) and the VQA module in [`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py) interact with these abstracted interfaces, allowing you to upgrade the language model without modifying the graph extraction or reasoning logic.

### Is TextFlow more data-efficient than training an end-to-end VLM for flowcharts?

TextFlow is significantly more data-efficient for specialized flowchart understanding tasks. Because Stage 1 extracts a compact symbolic graph representation, Stage 2 language models process concise textual descriptions rather than high-resolution image tokens. This reduces the context window requirements and the amount of training data needed to learn structural reasoning patterns. End-to-end VLMs must simultaneously learn visual feature extraction and logical reasoning from pixel data, requiring substantially larger datasets and compute resources to achieve comparable performance on flowchart-specific queries.