# Prompt Templates for the Vision Textualizer and Textual Reasoner in TextFlow

> Discover TextFlow's prompt templates for Vision Textualizer and Textual Reasoner. Generate diagrams and perform visual reasoning with these powerful LLM tools.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: api-reference
- Published: 2026-03-05

---

**TextFlow utilizes `load_textualizer_prompt()` to instruct vision-language models on generating diagram code in Mermaid, Graphviz, or PlantUML syntax, while `load_reasoner_prompt()` formats visual reasoning queries that inject the textualized representation into a question-answering context for text-only LLMs.**

The junyiye/textflow repository implements a dual-stage pipeline that converts flowchart images into executable textual representations and subsequently enables visual reasoning over those abstractions. At the core of this architecture lie two specialized prompt-generation utilities defined in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py) that standardize how models interact with visual inputs and textual diagram code. Understanding the prompt templates used for the Vision Textualizer and Textual Reasoner reveals how the system bridges multimodal perception and logical inference.

## Vision Textualizer Prompt Template

The **Vision Textualizer** relies on the `load_textualizer_prompt()` function to generate format-specific instructions that guide vision-language models in converting images into diagram code. Located at lines 5–74 in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py), this function accepts an `output_type` parameter and returns a multi-line instruction string containing concrete syntax examples for the requested format.

### Supported Output Formats and Syntax Examples

The template supports three distinct diagram languages, with each branch embedding specific syntax rules and concrete examples that demonstrate exactly how the model should structure its output:

- **Mermaid** – Returns instructions for generating flowchart code using Mermaid.js syntax, including example graph declarations.
- **Graphviz** – Provides the DOT language structure for directed graphs, with sample node and edge definitions.
- **PlantUML** – Supplies UML diagram specifications using PlantUML syntax, demonstrating proper start/end delimiters.

Each variant includes a complete, runnable example of the target language within the prompt text, ensuring zero-shot generation of syntactically valid code without requiring additional few-shot examples in the API call.

### Pipeline Integration in textualizer.py

In [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), the system invokes `load_textualizer_prompt(output_type)` to retrieve the appropriate template before passing it to the vision-language model. The model receives both the prompt instructions and the image path, generating a response that subsequently undergoes parsing to extract only the diagram code block.

```python

# src/textualizer.py

from prompts import load_textualizer_prompt

prompt = load_textualizer_prompt("mermaid")          # ← template selection

response = model.generate_response(prompt, image_path=image_path)
representation = extract_representation(response)    # isolates code block

```

The resulting `representation` variable contains clean diagram code ready for the reasoning stage.

## Textual Reasoner Prompt Template

The **Textual Reasoner** employs `load_reasoner_prompt()`, defined at lines 77–78 in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py), to construct queries that combine a previously generated textual diagram with a natural language question. This function accepts two parameters: `question` (the visual reasoning query) and `representation` (the diagram code produced by the Vision Textualizer).

### Prompt Construction Logic

The template constructs a strictly formatted string that positions the diagram context before the question using this exact pattern:

```

{representation}

Question: {question}
Answer:

```

This format ensures the text-only LLM receives the full visual context as structured text before attempting to answer questions about node relationships, flow logic, or algorithmic states. The explicit `Answer:` suffix triggers the model to generate the completion immediately without additional formatting.

### Integration in the Reasoning Pipeline

In [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), the system feeds the textual representation into this template, creating a self-contained reasoning context that requires no access to the original image.

```python

# src/reasoner.py

from prompts import load_reasoner_prompt

# representation contains the Mermaid/Graphviz code from the Textualizer

question = "What is the final state of the algorithm if the input array is empty?"
prompt = load_reasoner_prompt(question, representation)   # ← combines diagram + Q

answer = model.generate_response(prompt)

```

This round-trip design enables complex visual reasoning solely through text-based LLM inference, effectively decoupling the reasoning stage from vision processing capabilities.

## End-to-End Workflow Example

These two templates enable TextFlow's complete visual-to-reasoning workflow, allowing specialized models to handle distinct phases of the pipeline. First, `load_textualizer_prompt()` instructs a vision-language model like **Qwen2-VL-7B** to transcribe an image into standardized diagram code. Then, `load_reasoner_prompt()` injects that code into a question-answering context for a text-only model such as **Qwen2-7B**.

```python
from prompts import load_textualizer_prompt, load_reasoner_prompt
from models import ModelWrapper

# Stage 1: Visual to Text

vision_model = ModelWrapper("Qwen2-VL-7B")
textualizer_prompt = load_textualizer_prompt("mermaid")
diagram_code = vision_model.generate_response(
    textualizer_prompt, 
    image_path="flowchart.png"
)

# Stage 2: Textual Reasoning

text_model = ModelWrapper("Qwen2-7B")
reasoner_prompt = load_reasoner_prompt(
    "What is the final state of the algorithm if the input array is empty?",
    diagram_code
)
answer = text_model.generate_response(reasoner_prompt)

```

## Summary

- **`load_textualizer_prompt(output_type)`** in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py) (lines 5–74) generates format-specific instructions for converting images to diagram code, embedding concrete Mermaid, Graphviz, or PlantUML syntax examples for each target language.
- **`load_reasoner_prompt(question, representation)`** in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py) (lines 77–78) constructs a structured prompt using the format `{representation}\n\nQuestion: {question}\nAnswer:` to enable text-only LLMs to reason over visual structures.
- The **Vision Textualizer** operates within [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), utilizing vision-language models to extract clean textual representations from flowchart images.
- The **Textual Reasoner** operates within [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), enabling sophisticated logical analysis of the extracted representations without requiring multimodal capabilities during the reasoning phase.
- Together, these prompt templates facilitate a decoupled architecture where **Qwen2-VL-7B** handles visual perception and **Qwen2-7B** handles logical inference.

## Frequently Asked Questions

### How does the Vision Textualizer prompt template vary between different output formats?

The `load_textualizer_prompt()` function branches based on the `output_type` argument to return specialized instructions for Mermaid, Graphviz, or PlantUML. Each variant includes distinct syntax rules and concrete code examples specific to that diagram language, ensuring the vision-language model generates valid, executable code for the requested format.

### What is the exact string format used by the Textual Reasoner prompt template?

According to the implementation in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py) at lines 77–78, the Textual Reasoner constructs prompts using the strict format: `{representation}\n\nQuestion: {question}\nAnswer:`. This structure places the full diagram code first, followed by the question, with an explicit `Answer:` prefix that cues the LLM to generate the final response.

### Can different language models be used for the Textualizer and Reasoner components?

Yes. The pipeline architecture explicitly supports model specialization. The Vision Textualizer requires a multimodal model like **Qwen2-VL-7B** capable of processing both the prompt template and the input image, while the Textual Reasoner can utilize text-only models like **Qwen2-7B** that receive the pre-extracted diagram representation through the reasoner prompt template.

### Where are the prompt template functions defined in the TextFlow codebase?

Both prompt generation functions are centralized in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py), with `load_textualizer_prompt()` occupying lines 5–74 and `load_reasoner_prompt()` at lines 77–78. These utilities are imported by [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) and [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) respectively to orchestrate the visual-to-text conversion and subsequent reasoning workflow.