TextFlow Two-Stage Architecture vs End-to-End VLM: A Technical Deep Dive for Flowchart Understanding

TextFlow's two-stage architecture explicitly separates visual structure extraction from logical reasoning, converting flowchart images into deterministic graph representations before applying language model inference, while end-to-end vision-language models attempt to simultaneously perceive and reason over raw pixels in a single forward pass.

The junyiye/textflow repository implements a specialized two-stage architecture for flowchart understanding that challenges the monolithic design of modern vision-language models. Unlike end-to-end approaches that feed raw images directly into transformer stacks, TextFlow first extracts an explicit symbolic graph using the Textualizer class, then performs targeted reasoning via the Reasoner module.

Understanding TextFlow's Two-Stage Pipeline

TextFlow processes flowcharts through a deterministic pipeline that decouples perception from cognition. This separation allows for explicit error checking and modular component upgrades without retraining entire systems.

Stage 1: Structured Extraction with the Textualizer

The first stage converts raw flowchart images into a formal graph representation. In src/textualizer.py, the Textualizer class orchestrates OCR and shape detection to identify nodes and edges.

The output is a Flowchart object defined in src/flowchart.py, which stores the diagram as a structured graph with explicit node identifiers and edge relationships. This object provides deterministic methods like to_dict() that serialize the flowchart into JSON format for inspection or downstream processing.

from textualizer import Textualizer
from flowchart import Flowchart

# Convert image to explicit graph structure

txtzr = Textualizer()
graph = txtzr.image_to_flowchart("data/flowvqa/images/instruct00180.png")

# graph is a Flowchart object with deterministic nodes/edges

# Verify intermediate representation

json_repr = graph.to_dict()
print(json_repr)  # {"nodes": {"A": {"id":"A", "description":"Start", ...}, ...}}

Stage 2: Graph-Based Reasoning and VQA

The second stage operates on the extracted graph rather than pixels. In src/reasoner.py, the Reasoner class performs structural operations such as shortest path calculation, successor identification, and connectivity checks.

The VQA module in src/vqa.py orchestrates this process, feeding graph traversals to language models via src/models/api_models.py. Because the visual parsing is already complete, the language model receives clean symbolic input, reducing the risk of hallucinations caused by misinterpreting visual layouts.

from reasoner import Reasoner

# Initialize reasoner with extracted graph

reasoner = Reasoner(graph)

# Answer structural questions using graph traversal

answer = reasoner.answer_question(
    "What is the shortest path length from 'Start' to 'Decision'?"
)
print(answer)  # "2"

How End-to-End VLM Approaches Handle Flowcharts

End-to-end vision-language models, such as GPT-4V or LLaVA, process flowchart understanding as a single multimodal task. These systems feed raw bitmaps directly into transformer architectures that simultaneously handle visual feature extraction and linguistic reasoning.

In this paradigm, the model must implicitly learn to detect shapes, read text via OCR, identify connecting lines, and perform logical reasoning within a single forward pass. There is no explicit intermediate graph representation; instead, structural relationships are encoded as latent activations distributed across the model's parameters.

TextFlow Two-Stage Architecture vs End-to-End VLM: Critical Differences

The architectural divergence between TextFlow and monolithic VLMs creates distinct operational characteristics for flowchart understanding tasks.

Modularity and Debugging

TextFlow's separation of concerns allows independent optimization of visual extraction and reasoning components. Developers can inspect the intermediate JSON output from src/flowchart.py::to_dict() to verify that node detection succeeded before any language model calls occur. End-to-end VLMs offer no such inspection points; errors in visual parsing are inseparable from reasoning failures.

Error Propagation and Correction

In TextFlow, Stage 1 errors (e.g., missed edges) are visible in the explicit graph structure and can be corrected via rule-based post-processing before Stage 2 executes. End-to-end systems propagate visual misinterpretations directly into the reasoning pathway, often resulting in confident hallucinations about non-existent connections.

Data Efficiency and Token Usage

TextFlow's symbolic graph provides a concise description of flowchart structure, requiring significantly fewer tokens for the language model to comprehend logical relationships. End-to-end VLMs must process high-resolution image tokens alongside text, demanding larger context windows and more training data to achieve comparable structural understanding.

Implementation Walkthrough: Key Source Files

The junyiye/textflow repository organizes its two-stage pipeline across specialized modules that handle distinct aspects of flowchart understanding.

File Role Key Components
src/textualizer.py Stage 1 implementation Textualizer class, image_to_flowchart() method
src/flowchart.py Graph representation Flowchart class, to_dict() serialization
src/reasoner.py Stage 2 reasoning Reasoner class, graph traversal algorithms
src/vqa.py Pipeline orchestration VQA module coordinating extraction and reasoning
src/models/api_models.py LLM interface API wrappers for language model calls
src/models/model_loader.py Backend management Model loading for local or remote LLMs

Summary

TextFlow's two-stage architecture offers a deterministic alternative to end-to-end vision-language models for flowchart understanding tasks. Key advantages include:

  • Explicit graph representation: The Textualizer converts images to structured Flowchart objects before reasoning begins
  • Modular error handling: Visual parsing errors in Stage 1 can be detected and corrected via to_dict() inspection before Stage 2 executes
  • Efficient reasoning: The Reasoner operates on compact symbolic graphs rather than high-resolution pixel tensors, reducing token consumption and hallucination risks
  • Component independence: Each stage can be upgraded independently (e.g., swapping OCR engines or LLM back-ends) without retraining the entire pipeline

Frequently Asked Questions

What is the main advantage of TextFlow's two-stage architecture over end-to-end VLMs?

The primary advantage is error visibility and correction. TextFlow's Stage 1 produces an explicit graph representation via src/textualizer.py that developers can inspect using Flowchart.to_dict(). If the OCR misses a node or an edge is misidentified, you can apply rule-based corrections before Stage 2 reasoning begins. End-to-end VLMs lack this inspection point; visual errors propagate directly into the reasoning layer, often producing confident but incorrect answers about flowchart structure.

How does the Textualizer convert flowchart images into structured data?

The Textualizer class in src/textualizer.py orchestrates a pipeline of computer vision techniques including OCR and shape detection to identify flowchart elements. It processes the input image and constructs a Flowchart object (defined in src/flowchart.py) containing explicit node and edge representations. This object provides methods like to_dict() to serialize the graph into JSON format, creating a deterministic intermediate representation that Stage 2 components consume for reasoning tasks.

Can TextFlow's reasoning components work with different language models?

Yes, the architecture decouples the reasoning layer from specific LLM implementations. The src/models/api_models.py file provides wrappers for various language model APIs, while src/models/model_loader.py handles the loading of both local and remote back-ends (such as GPT-4 or Claude). The Reasoner class in src/reasoner.py and the VQA module in src/vqa.py interact with these abstracted interfaces, allowing you to upgrade the language model without modifying the graph extraction or reasoning logic.

Is TextFlow more data-efficient than training an end-to-end VLM for flowcharts?

TextFlow is significantly more data-efficient for specialized flowchart understanding tasks. Because Stage 1 extracts a compact symbolic graph representation, Stage 2 language models process concise textual descriptions rather than high-resolution image tokens. This reduces the context window requirements and the amount of training data needed to learn structural reasoning patterns. End-to-end VLMs must simultaneously learn visual feature extraction and logical reasoning from pixel data, requiring substantially larger datasets and compute resources to achieve comparable performance on flowchart-specific queries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →