Prompt Templates for the Vision Textualizer and Textual Reasoner in TextFlow
TextFlow utilizes load_textualizer_prompt() to instruct vision-language models on generating diagram code in Mermaid, Graphviz, or PlantUML syntax, while load_reasoner_prompt() formats visual reasoning queries that inject the textualized representation into a question-answering context for text-only LLMs.
The junyiye/textflow repository implements a dual-stage pipeline that converts flowchart images into executable textual representations and subsequently enables visual reasoning over those abstractions. At the core of this architecture lie two specialized prompt-generation utilities defined in src/prompts/prompts.py that standardize how models interact with visual inputs and textual diagram code. Understanding the prompt templates used for the Vision Textualizer and Textual Reasoner reveals how the system bridges multimodal perception and logical inference.
Vision Textualizer Prompt Template
The Vision Textualizer relies on the load_textualizer_prompt() function to generate format-specific instructions that guide vision-language models in converting images into diagram code. Located at lines 5–74 in src/prompts/prompts.py, this function accepts an output_type parameter and returns a multi-line instruction string containing concrete syntax examples for the requested format.
Supported Output Formats and Syntax Examples
The template supports three distinct diagram languages, with each branch embedding specific syntax rules and concrete examples that demonstrate exactly how the model should structure its output:
- Mermaid – Returns instructions for generating flowchart code using Mermaid.js syntax, including example graph declarations.
- Graphviz – Provides the DOT language structure for directed graphs, with sample node and edge definitions.
- PlantUML – Supplies UML diagram specifications using PlantUML syntax, demonstrating proper start/end delimiters.
Each variant includes a complete, runnable example of the target language within the prompt text, ensuring zero-shot generation of syntactically valid code without requiring additional few-shot examples in the API call.
Pipeline Integration in textualizer.py
In src/textualizer.py, the system invokes load_textualizer_prompt(output_type) to retrieve the appropriate template before passing it to the vision-language model. The model receives both the prompt instructions and the image path, generating a response that subsequently undergoes parsing to extract only the diagram code block.
# src/textualizer.py
from prompts import load_textualizer_prompt
prompt = load_textualizer_prompt("mermaid") # ← template selection
response = model.generate_response(prompt, image_path=image_path)
representation = extract_representation(response) # isolates code block
The resulting representation variable contains clean diagram code ready for the reasoning stage.
Textual Reasoner Prompt Template
The Textual Reasoner employs load_reasoner_prompt(), defined at lines 77–78 in src/prompts/prompts.py, to construct queries that combine a previously generated textual diagram with a natural language question. This function accepts two parameters: question (the visual reasoning query) and representation (the diagram code produced by the Vision Textualizer).
Prompt Construction Logic
The template constructs a strictly formatted string that positions the diagram context before the question using this exact pattern:
{representation}
Question: {question}
Answer:
This format ensures the text-only LLM receives the full visual context as structured text before attempting to answer questions about node relationships, flow logic, or algorithmic states. The explicit Answer: suffix triggers the model to generate the completion immediately without additional formatting.
Integration in the Reasoning Pipeline
In src/reasoner.py, the system feeds the textual representation into this template, creating a self-contained reasoning context that requires no access to the original image.
# src/reasoner.py
from prompts import load_reasoner_prompt
# representation contains the Mermaid/Graphviz code from the Textualizer
question = "What is the final state of the algorithm if the input array is empty?"
prompt = load_reasoner_prompt(question, representation) # ← combines diagram + Q
answer = model.generate_response(prompt)
This round-trip design enables complex visual reasoning solely through text-based LLM inference, effectively decoupling the reasoning stage from vision processing capabilities.
End-to-End Workflow Example
These two templates enable TextFlow's complete visual-to-reasoning workflow, allowing specialized models to handle distinct phases of the pipeline. First, load_textualizer_prompt() instructs a vision-language model like Qwen2-VL-7B to transcribe an image into standardized diagram code. Then, load_reasoner_prompt() injects that code into a question-answering context for a text-only model such as Qwen2-7B.
from prompts import load_textualizer_prompt, load_reasoner_prompt
from models import ModelWrapper
# Stage 1: Visual to Text
vision_model = ModelWrapper("Qwen2-VL-7B")
textualizer_prompt = load_textualizer_prompt("mermaid")
diagram_code = vision_model.generate_response(
textualizer_prompt,
image_path="flowchart.png"
)
# Stage 2: Textual Reasoning
text_model = ModelWrapper("Qwen2-7B")
reasoner_prompt = load_reasoner_prompt(
"What is the final state of the algorithm if the input array is empty?",
diagram_code
)
answer = text_model.generate_response(reasoner_prompt)
Summary
load_textualizer_prompt(output_type)insrc/prompts/prompts.py(lines 5–74) generates format-specific instructions for converting images to diagram code, embedding concrete Mermaid, Graphviz, or PlantUML syntax examples for each target language.load_reasoner_prompt(question, representation)insrc/prompts/prompts.py(lines 77–78) constructs a structured prompt using the format{representation}\n\nQuestion: {question}\nAnswer:to enable text-only LLMs to reason over visual structures.- The Vision Textualizer operates within
src/textualizer.py, utilizing vision-language models to extract clean textual representations from flowchart images. - The Textual Reasoner operates within
src/reasoner.py, enabling sophisticated logical analysis of the extracted representations without requiring multimodal capabilities during the reasoning phase. - Together, these prompt templates facilitate a decoupled architecture where Qwen2-VL-7B handles visual perception and Qwen2-7B handles logical inference.
Frequently Asked Questions
How does the Vision Textualizer prompt template vary between different output formats?
The load_textualizer_prompt() function branches based on the output_type argument to return specialized instructions for Mermaid, Graphviz, or PlantUML. Each variant includes distinct syntax rules and concrete code examples specific to that diagram language, ensuring the vision-language model generates valid, executable code for the requested format.
What is the exact string format used by the Textual Reasoner prompt template?
According to the implementation in src/prompts/prompts.py at lines 77–78, the Textual Reasoner constructs prompts using the strict format: {representation}\n\nQuestion: {question}\nAnswer:. This structure places the full diagram code first, followed by the question, with an explicit Answer: prefix that cues the LLM to generate the final response.
Can different language models be used for the Textualizer and Reasoner components?
Yes. The pipeline architecture explicitly supports model specialization. The Vision Textualizer requires a multimodal model like Qwen2-VL-7B capable of processing both the prompt template and the input image, while the Textual Reasoner can utilize text-only models like Qwen2-7B that receive the pre-extracted diagram representation through the reasoner prompt template.
Where are the prompt template functions defined in the TextFlow codebase?
Both prompt generation functions are centralized in src/prompts/prompts.py, with load_textualizer_prompt() occupying lines 5–74 and load_reasoner_prompt() at lines 77–78. These utilities are imported by src/textualizer.py and src/reasoner.py respectively to orchestrate the visual-to-text conversion and subsequent reasoning workflow.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →