How TextFlow Improves Explainability Over End-to-End Vision-Language Models
TextFlow improves explainability by replacing monolithic vision-language models with a modular two-stage pipeline that generates human-readable intermediate representations, enabling precise error attribution between visual parsing and textual reasoning.
The junyiye/textflow repository implements a decomposed architecture that explicitly separates visual understanding from language-based reasoning. Unlike end-to-end vision-language models (VLMs) that operate as black boxes, TextFlow produces inspectable text-based graph representations that serve as a transparent bridge between image input and final answers.
The Two-Stage Architecture That Improves Explainability
TextFlow’s explainability stems from its strict separation of concerns into two distinct stages, each with dedicated source files and clear responsibilities.
Stage 1: Vision Textualizer
The Vision Textualizer converts flowchart images into structured text representations using src/textualizer.py. This stage employs a vision-language model to parse visual elements and output human-readable scripts in formats like Mermaid, Graphviz, or PlantUML.
The core generation logic resides in the model interaction loop:
# From src/textualizer.py
response = model.generate(
image=input_image,
prompt=load_textualizer_prompt()
)
representation = extract_representation(response)
Because the visual encoder outputs a human-readable script, developers can directly inspect whether the image was parsed correctly. Any error in the visual stage appears as a malformed or missing graph element, which can be fixed without touching the reasoning component.
Stage 2: Textual Reasoner
The Textual Reasoner operates exclusively on the generated text representation, completely decoupling reasoning from visual processing. Implemented in src/reasoner.py, this stage takes the intermediate script and a user question, then generates answers using a language model.
The reasoning pipeline uses load_reasoner_prompt() to construct the input context and calls model.generate_response() to produce the final output:
# From src/reasoner.py
prompt = load_reasoner_prompt(textual_representation, question)
answer = model.generate_response(prompt)
Since the reasoning stage operates only on text, any mistake can be traced back to language-model inference rather than visual encoding. If the answer is wrong, developers inspect the input script to verify whether the reasoning was given correct premises.
Debugging and Error Attribution
TextFlow’s modular design enables independent debugging of each pipeline stage. When an end-to-end VLM produces an incorrect answer, it is impossible to determine whether the error originated in visual misrecognition or faulty reasoning. TextFlow eliminates this ambiguity.
The project’s README explicitly highlights this advantage: "It improves explainability by helping to attribute errors more clearly to visual or textual processing components"【/cache/repos/github.com/junyiye/textflow/main/README.md#L20-L23】.
Developers can log intermediate representations at the boundary between stages, swap out the reasoning model (e.g., upgrading to a stronger LLM) without retraining the visual model, and validate the textualizer output using standard graph syntax checkers before it ever reaches the reasoning stage.
Running the TextFlow Pipeline
Converting Images to Text Representations
Execute the Vision Textualizer to generate intermediate graph scripts:
python src/textualizer.py \
--dataset flowvqa \
--textualizer gpt-4o \
--output_type mermaid
This produces a JSON file at output/flowvqa/mermaid/gpt-4o.json containing Mermaid scripts for each processed image.
Running Textual Reasoning
Process the generated scripts through the Textual Reasoner:
python src/reasoner.py \
--dataset flowvqa \
--reasoner gpt-4o \
--textualizer gpt-4o \
--input_type mermaid
Output is saved to output/flowvqa/textflow/mermaid_reasoner_gpt-4o_textualizer_gpt-4o.json with questions, model responses, and ground-truth answers.
Enabling Tool Use for Enhanced Debugging
Add the --tool_use flag to allow the reasoner to access external graph execution utilities:
python src/reasoner.py \
--dataset flowvqa \
--reasoner gpt-4o \
--textualizer gpt-4o \
--input_type mermaid \
--tool_use
When enabled, the reasoner receives the raw script and can invoke external utilities for graph validation, providing an additional layer of explainability through executable intermediate representations.
Summary
- TextFlow improves explainability by decomposing vision-language tasks into isolated visual and textual stages rather than using monolithic end-to-end models.
- The Vision Textualizer (
src/textualizer.py) generates human-readable intermediate representations (Mermaid, Graphviz, PlantUML) that can be directly inspected for parsing errors. - The Textual Reasoner (
src/reasoner.py) operates exclusively on text, enabling precise attribution of errors to either visual misrecognition or faulty language model inference. - This modular architecture supports independent debugging, intermediate logging, and component swapping without retraining, addressing the black-box limitations of traditional VLMs.
Frequently Asked Questions
How does TextFlow's two-stage pipeline improve error debugging compared to end-to-end VLMs?
TextFlow isolates visual processing in the Vision Textualizer and reasoning in the Textual Reasoner, allowing developers to inspect the human-readable intermediate representation between stages. When an end-to-end VLM fails, it is impossible to determine whether the error originated in visual misrecognition or reasoning flaws. TextFlow makes this distinction explicit by exposing the textualized graph output, enabling targeted fixes to either src/textualizer.py or src/reasoner.py without affecting the other component.
What intermediate formats does TextFlow use to improve explainability?
TextFlow generates human-readable graph scripts including Mermaid, Graphviz DOT, and PlantUML formats. These textual representations serve as transparent bridges between the image input and final answer. Because these formats are human-readable and syntactically valid, developers can validate the Vision Textualizer's output using standard graph visualization tools before it reaches the reasoning stage, providing an auditable trail of how the visual input was interpreted.
Can the reasoning component be upgraded without retraining the visual model?
Yes, the modular architecture explicitly supports component swapping. Since the Textual Reasoner in src/reasoner.py operates exclusively on the text output from the Vision Textualizer, you can replace the reasoning model (for example, upgrading from GPT-4 to GPT-4o or switching to a different LLM) without modifying or retraining the visual textualizer. This separation ensures that improvements in reasoning capabilities do not require expensive retraining of vision-language models.
How does TextFlow handle tool use for additional validation?
TextFlow supports an optional tool use mode activated by the --tool_use flag when running src/reasoner.py. When enabled, the reasoner receives the raw intermediate script and can invoke external utilities for graph execution and validation. This provides an additional layer of explainability by allowing the system to verify graph syntax or execute the textualized representation to check for logical consistency before generating the final answer, effectively using the intermediate format as an executable audit trail.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →