# How to Add Support for a New Text Representation Format in TextFlow: A Complete Guide

> Easily add new text representation formats to TextFlow. Follow our guide to extend prompt loaders, implement extraction functions, and register your new handler. Enhance TextFlow today.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: how-to-guide
- Published: 2026-03-05

---

**To add support for a new text representation format in TextFlow, extend the prompt loader in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py), implement a format-specific extraction function in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py), and register the new handler in `extract_representation`.**

TextFlow converts flowchart images into textual code representations such as Mermaid, Graphviz, and PlantUML by orchestrating prompts, vision-language models (VLMs), and post-processing utilities. When you need to generate output in a custom format—like SVG, TikZ, or a domain-specific diagram language—you must extend three integration points in the `junyiye/textflow` codebase. This guide provides the exact implementation steps using the repository's actual source files.

## How TextFlow Processes Text Representations

Understanding the data flow helps clarify where to inject new format support. According to the source code in `junyiye/textflow`, the pipeline executes four distinct stages:

1. **Prompt Selection** – The `load_textualizer_prompt(output_type)` function in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py) returns a format-specific system prompt that instructs the VLM to generate code in the requested syntax.

2. **Model Generation** – The `ModelWrapper` class (imported and invoked in [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py)) transmits the prompt and image to the LLM, receiving a raw text response that typically includes fenced code blocks.

3. **Code Extraction** – The `extract_representation(response)` function in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py) scans the LLM output for specific fence identifiers (e.g., ```mermaid, ```dot, ```plantuml) and isolates the inner code using regex patterns.

4. **Persistence** – The extracted string is written to a JSON file in a subdirectory named after the `output_type` parameter.

To add a new format, you must modify stages 1 and 3. The CLI in `src/textualizer.py` requires no changes because it already forwards the `args.output_type` value through the pipeline.

## Step-by-Step Implementation Guide

### Step 1: Create a Format-Specific Prompt in prompts.py

The VLM requires explicit instructions and examples to generate valid code in unfamiliar syntax. In `src/prompts/prompts.py`, add a new conditional branch to `load_textualizer_prompt` that returns a tailored prompt for your format.

```python

# src/prompts/prompts.py

def load_textualizer_prompt(output_type):
    if output_type == "mermaid":
        # ... existing implementation

        pass
    elif output_type == "graphviz":
        # ... existing implementation

        pass
    elif output_type == "plantuml":
        # ... existing implementation

        pass
    elif output_type == "svg":               # ← new branch

        return """Generate the SVG code for the provided flowchart.

Here is an example:

```svg
<svg width="200" height="100" xmlns="http://www.w3.org/2000/svg">
  <rect x="10" y="10" width="180" height="80" fill="lightblue" stroke="black"/>
  <text x="100" y="55" font-family="sans-serif" font-size="14" text-anchor="middle">Start</text>
</svg>

```"""
    else:
        raise ValueError(f"Unsupported output type: {output_type}")

```

Include a complete, valid example inside the prompt to minimize hallucination and syntax errors.

### Step 2: Implement Custom Extraction Logic in utils.py

LLM responses often contain conversational text surrounding the code block. Create a dedicated extractor function in `src/utils.py` that uses regex to capture only the content within your new fence type.

```python

# src/utils.py

import re

def extract_svg_code(string):
    """Return only the SVG markup inside a ```svg``` fence."""
    svg_pattern = r"```svg\s+([\s\S]*?)```"
    match = re.search(svg_pattern, string)
    if match:
        return match.group(1).strip()
    return string

```

Pattern your implementation after existing extractors like `extract_mermaid_code` or `extract_graphviz_code` to maintain consistency across the codebase.

### Step 3: Register the Extractor in extract_representation

Update the central dispatcher function `extract_representation` in `src/utils.py` to recognize the new format identifier and route to your extractor.

```python

# src/utils.py

def extract_representation(string):
    if "```mermaid" in string:
        return extract_mermaid_code(string)
    elif "```dot" in string:
        return extract_graphviz_code(string)
    elif "```plantuml" in string:
        return extract_plantuml_code(string)
    elif "```svg" in string:               # ← new condition

        return extract_svg_code(string)
    else:
        return string

```

This integration ensures that when the LLM returns a response containing ```svg fences, the pipeline automatically isolates the markup.

### Step 4: Test via CLI

With the prompt and extraction logic in place, invoke the textualizer using your new `output_type` value. The existing argument parser in [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) accepts arbitrary string values for `--output_type` and passes them through to the prompt loader and file writer.

```bash
python -m src.textualizer \
    --dataset flowvqa \
    --textualizer Qwen2-VL-7B \
    --output_type svg

```

The command creates an `svg/` subdirectory under the configured output path and stores the extracted code in JSON format.

## Summary

To successfully add support for a new text representation format in TextFlow, implement these three core changes:

- **Extend `load_textualizer_prompt`** in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py) to return instructions and examples for the new syntax.
- **Create a format-specific extractor** (e.g., `extract_svg_code`) in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py) to parse the LLM's fenced code blocks.
- **Register the extractor** in `extract_representation` within [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py) to enable automatic routing.
- **Leverage the existing CLI** in [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), which requires no modification to accept new `output_type` values.

## Frequently Asked Questions

### Do I need to modify the CLI code to accept new output formats?

No. The CLI argument parser in [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) forwards the `--output_type` value directly to `load_textualizer_prompt` and uses it to construct the output directory path. As long as you handle the new value in the prompt loader and extraction logic, the CLI works without modification.

### How does TextFlow handle malformed or missing code blocks?

If the LLM response does not contain the expected fence identifier (e.g., ```svg), the `extract_representation` function returns the raw response string unchanged. You can add validation logic inside your specific extractor to raise warnings or return empty strings when patterns fail to match.

### Can I add multiple new formats simultaneously?

Yes. You can extend the conditional chains in both `load_textualizer_prompt` and `extract_representation` with additional `elif` branches for each format. Ensure each format has a unique fence identifier (e.g., ```tikz, ```d2) to prevent extraction collisions.

### Where should I add unit tests for the new extraction logic?

Add test cases in [`src/test_models.py`](https://github.com/junyiye/textflow/blob/main/src/test_models.py) following the existing test patterns. Verify that your extractor correctly handles responses with surrounding text, responses without code blocks, and responses with multiple fenced blocks to ensure robustness in production environments.