How to Add Support for a New Text Representation Format in TextFlow: A Complete Guide
To add support for a new text representation format in TextFlow, extend the prompt loader in src/prompts/prompts.py, implement a format-specific extraction function in src/utils.py, and register the new handler in extract_representation.
TextFlow converts flowchart images into textual code representations such as Mermaid, Graphviz, and PlantUML by orchestrating prompts, vision-language models (VLMs), and post-processing utilities. When you need to generate output in a custom format—like SVG, TikZ, or a domain-specific diagram language—you must extend three integration points in the junyiye/textflow codebase. This guide provides the exact implementation steps using the repository's actual source files.
How TextFlow Processes Text Representations
Understanding the data flow helps clarify where to inject new format support. According to the source code in junyiye/textflow, the pipeline executes four distinct stages:
-
Prompt Selection – The
load_textualizer_prompt(output_type)function insrc/prompts/prompts.pyreturns a format-specific system prompt that instructs the VLM to generate code in the requested syntax. -
Model Generation – The
ModelWrapperclass (imported and invoked insrc/textualizer.py) transmits the prompt and image to the LLM, receiving a raw text response that typically includes fenced code blocks. -
Code Extraction – The
extract_representation(response)function insrc/utils.pyscans the LLM output for specific fence identifiers (e.g.,mermaid,dot, ```plantuml) and isolates the inner code using regex patterns. -
Persistence – The extracted string is written to a JSON file in a subdirectory named after the
output_typeparameter.
To add a new format, you must modify stages 1 and 3. The CLI in src/textualizer.py requires no changes because it already forwards the args.output_type value through the pipeline.
Step-by-Step Implementation Guide
Step 1: Create a Format-Specific Prompt in prompts.py
The VLM requires explicit instructions and examples to generate valid code in unfamiliar syntax. In src/prompts/prompts.py, add a new conditional branch to load_textualizer_prompt that returns a tailored prompt for your format.
# src/prompts/prompts.py
def load_textualizer_prompt(output_type):
if output_type == "mermaid":
# ... existing implementation
pass
elif output_type == "graphviz":
# ... existing implementation
pass
elif output_type == "plantuml":
# ... existing implementation
pass
elif output_type == "svg": # ← new branch
return """Generate the SVG code for the provided flowchart.
Here is an example:
```svg
<svg width="200" height="100" xmlns="http://www.w3.org/2000/svg">
<rect x="10" y="10" width="180" height="80" fill="lightblue" stroke="black"/>
<text x="100" y="55" font-family="sans-serif" font-size="14" text-anchor="middle">Start</text>
</svg>
```"""
else:
raise ValueError(f"Unsupported output type: {output_type}")
Include a complete, valid example inside the prompt to minimize hallucination and syntax errors.
Step 2: Implement Custom Extraction Logic in utils.py
LLM responses often contain conversational text surrounding the code block. Create a dedicated extractor function in src/utils.py that uses regex to capture only the content within your new fence type.
# src/utils.py
import re
def extract_svg_code(string):
"""Return only the SVG markup inside a ```svg``` fence."""
svg_pattern = r"```svg\s+([\s\S]*?)```"
match = re.search(svg_pattern, string)
if match:
return match.group(1).strip()
return string
Pattern your implementation after existing extractors like extract_mermaid_code or extract_graphviz_code to maintain consistency across the codebase.
Step 3: Register the Extractor in extract_representation
Update the central dispatcher function extract_representation in src/utils.py to recognize the new format identifier and route to your extractor.
# src/utils.py
def extract_representation(string):
if "```mermaid" in string:
return extract_mermaid_code(string)
elif "```dot" in string:
return extract_graphviz_code(string)
elif "```plantuml" in string:
return extract_plantuml_code(string)
elif "```svg" in string: # ← new condition
return extract_svg_code(string)
else:
return string
This integration ensures that when the LLM returns a response containing ```svg fences, the pipeline automatically isolates the markup.
Step 4: Test via CLI
With the prompt and extraction logic in place, invoke the textualizer using your new output_type value. The existing argument parser in src/textualizer.py accepts arbitrary string values for --output_type and passes them through to the prompt loader and file writer.
python -m src.textualizer \
--dataset flowvqa \
--textualizer Qwen2-VL-7B \
--output_type svg
The command creates an svg/ subdirectory under the configured output path and stores the extracted code in JSON format.
Summary
To successfully add support for a new text representation format in TextFlow, implement these three core changes:
- Extend
load_textualizer_promptinsrc/prompts/prompts.pyto return instructions and examples for the new syntax. - Create a format-specific extractor (e.g.,
extract_svg_code) insrc/utils.pyto parse the LLM's fenced code blocks. - Register the extractor in
extract_representationwithinsrc/utils.pyto enable automatic routing. - Leverage the existing CLI in
src/textualizer.py, which requires no modification to accept newoutput_typevalues.
Frequently Asked Questions
Do I need to modify the CLI code to accept new output formats?
No. The CLI argument parser in src/textualizer.py forwards the --output_type value directly to load_textualizer_prompt and uses it to construct the output directory path. As long as you handle the new value in the prompt loader and extraction logic, the CLI works without modification.
How does TextFlow handle malformed or missing code blocks?
If the LLM response does not contain the expected fence identifier (e.g., ```svg), the extract_representation function returns the raw response string unchanged. You can add validation logic inside your specific extractor to raise warnings or return empty strings when patterns fail to match.
Can I add multiple new formats simultaneously?
Yes. You can extend the conditional chains in both load_textualizer_prompt and extract_representation with additional elif branches for each format. Ensure each format has a unique fence identifier (e.g., tikz, d2) to prevent extraction collisions.
Where should I add unit tests for the new extraction logic?
Add test cases in src/test_models.py following the existing test patterns. Verify that your extractor correctly handles responses with surrounding text, responses without code blocks, and responses with multiple fenced blocks to ensure robustness in production environments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →