# How to Integrate TextFlow with Your Own Applications: Complete Developer Guide

> Integrate TextFlow into your Python apps with this developer guide. Programmatically convert flowchart images to text and perform visual question answering using modular pipeline components.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Integrate TextFlow into your Python applications by importing its modular pipeline components—ModelWrapper, Vision Textualizer, and Textual Reasoner—to programmatically convert flowchart images into structured text representations and perform visual question answering.**

The `junyiye/textflow` repository provides a modular framework for flowchart understanding that separates visual perception from textual reasoning. By leveraging its standardized interfaces, you can embed TextFlow's two-stage pipeline into web services, desktop applications, or automated evaluation workflows without modifying the core library source code.

## Understanding the TextFlow Architecture

TextFlow employs a **two-stage pipeline** that decouples visual understanding from logical reasoning. This separation allows you to swap vision-language models (VLMs) and large language models (LLMs) independently while maintaining consistent data flows.

The pipeline consists of these core components:

- **Configuration** ([`src/config.py`](https://github.com/junyiye/textflow/blob/main/src/config.py)): Loads paths, logging settings, and model credentials from [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json).
- **ModelWrapper** ([`src/models/model_wrapper.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_wrapper.py)): Provides a unified interface for both API-based and local models, handling image encoding and tool-use logic.
- **Vision Textualizer** ([`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py)): Converts flowchart images into text representations (Mermaid, Graphviz, or PlantUML) via VLM inference.
- **Textual Reasoner** ([`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py)): Consumes the generated representation to answer visual questions using an LLM, with optional tool-use support.
- **Prompt Templates** ([`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py)): Centralized prompt strings for both pipeline stages.
- **Utility Functions** ([`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py)): Helper methods for image encoding (`encode_image`) and representation extraction (`extract_representation`).

## Programmatic Integration Guide

You can embed TextFlow directly into your codebase by importing the core classes and bypassing the command-line interface. The following examples assume you have installed dependencies via `pip install -r requirements.txt` and configured [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) with your API keys.

### Step 1 — Generate Text Representations from Flowchart Images

To convert an image into a structured text format like Mermaid syntax, instantiate `ModelWrapper` and call the textualizer pipeline:

```python
from src.models.model_wrapper import ModelWrapper
from src.prompts.prompts import load_textualizer_prompt
from src.utils import extract_representation
from src.config import config
import os
import json

def generate_mermaid(image_path, textualizer="gpt-4o"):
    model = ModelWrapper(textualizer)
    prompt = load_textualizer_prompt("mermaid")
    
    # Generate response with image encoding handled internally

    response = model.generate_response(prompt, image_path=image_path)
    representation = extract_representation(response)
    
    # Optional: persist to output directory matching CLI structure

    out_dir = os.path.join(
        config["file_paths"]["output"], 
        "flowvqa", 
        "mermaid"
    )
    os.makedirs(out_dir, exist_ok=True)
    out_file = os.path.join(out_dir, f"{textualizer}.json")
    
    with open(out_file, "w") as f:
        json.dump({os.path.basename(image_path): representation}, f, indent=2)
    
    return representation

```

This function encodes the image, sends it to the specified VLM (e.g., `"gpt-4o"` or `"Qwen2-VL-7B"`), and extracts the diagram code using the utility function `extract_representation`.

### Step 2 — Answer Visual Questions Using Text Representations

Once you have the text representation (Mermaid code), pass it to the Textual Reasoner along with your question:

```python
from src.models.model_wrapper import ModelWrapper
from src.prompts.prompts import load_reasoner_prompt

def answer_question(question, mermaid_code, reasoner="gpt-4o", tool_use=False):
    model = ModelWrapper(reasoner)
    prompt = load_reasoner_prompt(question, mermaid_code)
    
    if tool_use:
        # Enables model-driven tool execution (e.g., Mermaid rendering)

        response = model.generate_response(
            prompt, 
            representation=mermaid_code
        )
    else:
        response = model.generate_response(prompt)
    
    return response

```

When `tool_use=True`, the `ModelWrapper` invokes `generate_api_response_tool_use` (implemented in [`src/models/model_wrapper.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_wrapper.py)), allowing the model to execute the diagram representation for structured validation.

### Step 3 — Build an End-to-End Integration Pipeline

Combine both stages into a single callable function for your application:

```python
def textflow_pipeline(image_path, question, 
                      textualizer="gpt-4o", 
                      reasoner="gpt-4o",
                      tool_use=False):
    """
    End-to-end TextFlow integration for flowchart VQA.
    """
    # Stage 1: Visual understanding

    representation = generate_mermaid(image_path, textualizer)
    
    # Stage 2: Logical reasoning

    answer = answer_question(
        question, 
        representation, 
        reasoner=reasoner, 
        tool_use=tool_use
    )
    
    return answer

# Usage in your application

result = textflow_pipeline(
    "diagrams/process_flow.png", 
    "What is the purpose of step 3?",
    tool_use=True
)

```

## Key Source Files for Integration

When integrating TextFlow, reference these specific files in the `junyiye/textflow` repository:

- **[`src/config.py`](https://github.com/junyiye/textflow/blob/main/src/config.py)**: Loads [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) containing file paths and model credentials.
- **[`src/models/model_wrapper.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_wrapper.py)**: The primary integration point; abstracts API calls (`load_api_model`) and local checkpoint loading (`load_local_model`).
- **[`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py)**: Low-level OpenAI and Anthropic API implementations used by `ModelWrapper`.
- **[`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py)**: Local model initialization for Llama-3.1, Qwen2-VL, and other Hugging Face models.
- **[`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py)**: CLI entry point for the Vision Textualizer stage; useful for understanding the default execution flow.
- **[`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py)**: CLI entry point for the Textual Reasoner stage; demonstrates `--tool_use` flag handling.
- **[`src/flowchart.py`](https://github.com/junyiye/textflow/blob/main/src/flowchart.py)**: In-memory flowchart representation and syntax conversion utilities.
- **[`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py)**: Image encoding and text extraction helper functions.

## Configuration Requirements

Before running the integration code:

1. **Configure [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json)** with your API keys (OpenAI, Anthropic) and adjust `file_paths` for your environment.
2. **Install dependencies**: `pip install -r requirements.txt` includes the required ML and API client libraries.
3. **Select compatible models**: Use any model name recognized by `ModelWrapper`, including API models (`"gpt-4o"`, `"claude-3-5-sonnet"`) or local checkpoints (`"Qwen2-VL-7B"`, `"Llama-3.1-8B"`).

## Summary

- **TextFlow** provides a modular two-stage pipeline for flowchart understanding through the `junyiye/textflow` repository.
- **`ModelWrapper`** ([`src/models/model_wrapper.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_wrapper.py)) offers a unified interface for both cloud and local vision-language models.
- **Vision Textualizer** converts images to Mermaid/Graphviz/PlantUML text via `load_textualizer_prompt` and `extract_representation`.
- **Textual Reasoner** answers questions using the generated representation through `load_reasoner_prompt`.
- Enable **tool-use** by passing `representation` to `generate_response` for model-driven diagram execution.
- You can integrate TextFlow into any Python application by importing these components directly without modifying core library files.

## Frequently Asked Questions

### Can I use local models instead of API-based LLMs with TextFlow?

Yes. The `ModelWrapper` class in [`src/models/model_wrapper.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_wrapper.py) automatically detects whether you provide an API model name (like `"gpt-4o"`) or a local checkpoint name (like `"Qwen2-VL-7B"`). Local models are loaded via `load_local_model` in [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py), while API models route through `load_api_model` in [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py).

### What output formats does the Vision Textualizer support?

According to the source code in [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) and [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py), the Vision Textualizer supports **Mermaid**, **Graphviz**, and **PlantUML** syntax. You specify the desired format via the `output_type` parameter when calling `load_textualizer_prompt(output_type)`.

### How do I enable tool-use for the Textual Reasoner?

Pass `tool_use=True` to your reasoning function, which triggers the `ModelWrapper` to use `generate_api_response_tool_use` instead of the standard generation method. This allows the LLM to invoke execution tools on the Mermaid representation for structured validation, as implemented in [`src/models/model_wrapper.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_wrapper.py).

### Do I need to modify the TextFlow source code to integrate it into my application?

No. You can integrate TextFlow by importing the public classes and functions (such as `ModelWrapper`, `load_textualizer_prompt`, and `load_reasoner_prompt`) directly from the `src` directory. The modular design in `junyiye/textflow` allows you to compose the pipeline stages programmatically without altering the library internals.