# How TextFlow Improves Explainability Over End-to-End Vision-Language Models

> TextFlow enhances VLM explainability with a modular pipeline. Generate human-readable intermediate representations for precise error attribution in visual parsing and textual reasoning.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: deep-dive
- Published: 2026-03-05

---

**TextFlow improves explainability by replacing monolithic vision-language models with a modular two-stage pipeline that generates human-readable intermediate representations, enabling precise error attribution between visual parsing and textual reasoning.**

The `junyiye/textflow` repository implements a **decomposed architecture** that explicitly separates visual understanding from language-based reasoning. Unlike end-to-end vision-language models (VLMs) that operate as black boxes, TextFlow produces inspectable text-based graph representations that serve as a transparent bridge between image input and final answers.

## The Two-Stage Architecture That Improves Explainability

TextFlow’s explainability stems from its strict separation of concerns into two distinct stages, each with dedicated source files and clear responsibilities.

### Stage 1: Vision Textualizer

The **Vision Textualizer** converts flowchart images into structured text representations using [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py). This stage employs a vision-language model to parse visual elements and output human-readable scripts in formats like Mermaid, Graphviz, or PlantUML.

The core generation logic resides in the model interaction loop:

```python

# From src/textualizer.py

response = model.generate(
    image=input_image,
    prompt=load_textualizer_prompt()
)
representation = extract_representation(response)

```

Because the visual encoder outputs a **human-readable script**, developers can directly inspect whether the image was parsed correctly. Any error in the visual stage appears as a malformed or missing graph element, which can be fixed without touching the reasoning component.

### Stage 2: Textual Reasoner

The **Textual Reasoner** operates exclusively on the generated text representation, completely decoupling reasoning from visual processing. Implemented in [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), this stage takes the intermediate script and a user question, then generates answers using a language model.

The reasoning pipeline uses `load_reasoner_prompt()` to construct the input context and calls `model.generate_response()` to produce the final output:

```python

# From src/reasoner.py

prompt = load_reasoner_prompt(textual_representation, question)
answer = model.generate_response(prompt)

```

Since the reasoning stage operates only on text, any mistake can be traced back to **language-model inference** rather than visual encoding. If the answer is wrong, developers inspect the input script to verify whether the reasoning was given correct premises.

## Debugging and Error Attribution

TextFlow’s modular design enables **independent debugging** of each pipeline stage. When an end-to-end VLM produces an incorrect answer, it is impossible to determine whether the error originated in visual misrecognition or faulty reasoning. TextFlow eliminates this ambiguity.

The project’s README explicitly highlights this advantage: *"It improves explainability by helping to attribute errors more clearly to visual or textual processing components"*【/cache/repos/github.com/junyiye/textflow/main/README.md#L20-L23】.

Developers can log intermediate representations at the boundary between stages, swap out the reasoning model (e.g., upgrading to a stronger LLM) without retraining the visual model, and validate the textualizer output using standard graph syntax checkers before it ever reaches the reasoning stage.

## Running the TextFlow Pipeline

### Converting Images to Text Representations

Execute the Vision Textualizer to generate intermediate graph scripts:

```bash
python src/textualizer.py \
    --dataset flowvqa \
    --textualizer gpt-4o \
    --output_type mermaid

```

This produces a JSON file at [`output/flowvqa/mermaid/gpt-4o.json`](https://github.com/junyiye/textflow/blob/main/output/flowvqa/mermaid/gpt-4o.json) containing Mermaid scripts for each processed image.

### Running Textual Reasoning

Process the generated scripts through the Textual Reasoner:

```bash
python src/reasoner.py \
    --dataset flowvqa \
    --reasoner gpt-4o \
    --textualizer gpt-4o \
    --input_type mermaid

```

Output is saved to [`output/flowvqa/textflow/mermaid_reasoner_gpt-4o_textualizer_gpt-4o.json`](https://github.com/junyiye/textflow/blob/main/output/flowvqa/textflow/mermaid_reasoner_gpt-4o_textualizer_gpt-4o.json) with questions, model responses, and ground-truth answers.

### Enabling Tool Use for Enhanced Debugging

Add the `--tool_use` flag to allow the reasoner to access external graph execution utilities:

```bash
python src/reasoner.py \
    --dataset flowvqa \
    --reasoner gpt-4o \
    --textualizer gpt-4o \
    --input_type mermaid \
    --tool_use

```

When enabled, the reasoner receives the raw script and can invoke external utilities for graph validation, providing an additional layer of explainability through executable intermediate representations.

## Summary

- **TextFlow improves explainability** by decomposing vision-language tasks into isolated visual and textual stages rather than using monolithic end-to-end models.
- The **Vision Textualizer** ([`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py)) generates human-readable intermediate representations (Mermaid, Graphviz, PlantUML) that can be directly inspected for parsing errors.
- The **Textual Reasoner** ([`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py)) operates exclusively on text, enabling precise attribution of errors to either visual misrecognition or faulty language model inference.
- This modular architecture supports **independent debugging**, intermediate logging, and component swapping without retraining, addressing the black-box limitations of traditional VLMs.

## Frequently Asked Questions

### How does TextFlow's two-stage pipeline improve error debugging compared to end-to-end VLMs?

TextFlow isolates visual processing in the **Vision Textualizer** and reasoning in the **Textual Reasoner**, allowing developers to inspect the human-readable intermediate representation between stages. When an end-to-end VLM fails, it is impossible to determine whether the error originated in visual misrecognition or reasoning flaws. TextFlow makes this distinction explicit by exposing the textualized graph output, enabling targeted fixes to either [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) or [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) without affecting the other component.

### What intermediate formats does TextFlow use to improve explainability?

TextFlow generates **human-readable graph scripts** including Mermaid, Graphviz DOT, and PlantUML formats. These textual representations serve as transparent bridges between the image input and final answer. Because these formats are human-readable and syntactically valid, developers can validate the Vision Textualizer's output using standard graph visualization tools before it reaches the reasoning stage, providing an auditable trail of how the visual input was interpreted.

### Can the reasoning component be upgraded without retraining the visual model?

Yes, the modular architecture explicitly supports **component swapping**. Since the Textual Reasoner in [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) operates exclusively on the text output from the Vision Textualizer, you can replace the reasoning model (for example, upgrading from GPT-4 to GPT-4o or switching to a different LLM) without modifying or retraining the visual textualizer. This separation ensures that improvements in reasoning capabilities do not require expensive retraining of vision-language models.

### How does TextFlow handle tool use for additional validation?

TextFlow supports an optional **tool use mode** activated by the `--tool_use` flag when running [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py). When enabled, the reasoner receives the raw intermediate script and can invoke external utilities for graph execution and validation. This provides an additional layer of explainability by allowing the system to verify graph syntax or execute the textualized representation to check for logical consistency before generating the final answer, effectively using the intermediate format as an executable audit trail.