How to Integrate TextFlow with Your Own Applications: Complete Developer Guide

Integrate TextFlow into your Python applications by importing its modular pipeline components—ModelWrapper, Vision Textualizer, and Textual Reasoner—to programmatically convert flowchart images into structured text representations and perform visual question answering.

The junyiye/textflow repository provides a modular framework for flowchart understanding that separates visual perception from textual reasoning. By leveraging its standardized interfaces, you can embed TextFlow's two-stage pipeline into web services, desktop applications, or automated evaluation workflows without modifying the core library source code.

Understanding the TextFlow Architecture

TextFlow employs a two-stage pipeline that decouples visual understanding from logical reasoning. This separation allows you to swap vision-language models (VLMs) and large language models (LLMs) independently while maintaining consistent data flows.

The pipeline consists of these core components:

  • Configuration (src/config.py): Loads paths, logging settings, and model credentials from config.json.
  • ModelWrapper (src/models/model_wrapper.py): Provides a unified interface for both API-based and local models, handling image encoding and tool-use logic.
  • Vision Textualizer (src/textualizer.py): Converts flowchart images into text representations (Mermaid, Graphviz, or PlantUML) via VLM inference.
  • Textual Reasoner (src/reasoner.py): Consumes the generated representation to answer visual questions using an LLM, with optional tool-use support.
  • Prompt Templates (src/prompts/prompts.py): Centralized prompt strings for both pipeline stages.
  • Utility Functions (src/utils.py): Helper methods for image encoding (encode_image) and representation extraction (extract_representation).

Programmatic Integration Guide

You can embed TextFlow directly into your codebase by importing the core classes and bypassing the command-line interface. The following examples assume you have installed dependencies via pip install -r requirements.txt and configured config.json with your API keys.

Step 1 — Generate Text Representations from Flowchart Images

To convert an image into a structured text format like Mermaid syntax, instantiate ModelWrapper and call the textualizer pipeline:

from src.models.model_wrapper import ModelWrapper
from src.prompts.prompts import load_textualizer_prompt
from src.utils import extract_representation
from src.config import config
import os
import json

def generate_mermaid(image_path, textualizer="gpt-4o"):
    model = ModelWrapper(textualizer)
    prompt = load_textualizer_prompt("mermaid")
    
    # Generate response with image encoding handled internally

    response = model.generate_response(prompt, image_path=image_path)
    representation = extract_representation(response)
    
    # Optional: persist to output directory matching CLI structure

    out_dir = os.path.join(
        config["file_paths"]["output"], 
        "flowvqa", 
        "mermaid"
    )
    os.makedirs(out_dir, exist_ok=True)
    out_file = os.path.join(out_dir, f"{textualizer}.json")
    
    with open(out_file, "w") as f:
        json.dump({os.path.basename(image_path): representation}, f, indent=2)
    
    return representation

This function encodes the image, sends it to the specified VLM (e.g., "gpt-4o" or "Qwen2-VL-7B"), and extracts the diagram code using the utility function extract_representation.

Step 2 — Answer Visual Questions Using Text Representations

Once you have the text representation (Mermaid code), pass it to the Textual Reasoner along with your question:

from src.models.model_wrapper import ModelWrapper
from src.prompts.prompts import load_reasoner_prompt

def answer_question(question, mermaid_code, reasoner="gpt-4o", tool_use=False):
    model = ModelWrapper(reasoner)
    prompt = load_reasoner_prompt(question, mermaid_code)
    
    if tool_use:
        # Enables model-driven tool execution (e.g., Mermaid rendering)

        response = model.generate_response(
            prompt, 
            representation=mermaid_code
        )
    else:
        response = model.generate_response(prompt)
    
    return response

When tool_use=True, the ModelWrapper invokes generate_api_response_tool_use (implemented in src/models/model_wrapper.py), allowing the model to execute the diagram representation for structured validation.

Step 3 — Build an End-to-End Integration Pipeline

Combine both stages into a single callable function for your application:

def textflow_pipeline(image_path, question, 
                      textualizer="gpt-4o", 
                      reasoner="gpt-4o",
                      tool_use=False):
    """
    End-to-end TextFlow integration for flowchart VQA.
    """
    # Stage 1: Visual understanding

    representation = generate_mermaid(image_path, textualizer)
    
    # Stage 2: Logical reasoning

    answer = answer_question(
        question, 
        representation, 
        reasoner=reasoner, 
        tool_use=tool_use
    )
    
    return answer

# Usage in your application

result = textflow_pipeline(
    "diagrams/process_flow.png", 
    "What is the purpose of step 3?",
    tool_use=True
)

Key Source Files for Integration

When integrating TextFlow, reference these specific files in the junyiye/textflow repository:

  • src/config.py: Loads config.json containing file paths and model credentials.
  • src/models/model_wrapper.py: The primary integration point; abstracts API calls (load_api_model) and local checkpoint loading (load_local_model).
  • src/models/api_models.py: Low-level OpenAI and Anthropic API implementations used by ModelWrapper.
  • src/models/local_models.py: Local model initialization for Llama-3.1, Qwen2-VL, and other Hugging Face models.
  • src/textualizer.py: CLI entry point for the Vision Textualizer stage; useful for understanding the default execution flow.
  • src/reasoner.py: CLI entry point for the Textual Reasoner stage; demonstrates --tool_use flag handling.
  • src/flowchart.py: In-memory flowchart representation and syntax conversion utilities.
  • src/utils.py: Image encoding and text extraction helper functions.

Configuration Requirements

Before running the integration code:

  1. Configure config.json with your API keys (OpenAI, Anthropic) and adjust file_paths for your environment.
  2. Install dependencies: pip install -r requirements.txt includes the required ML and API client libraries.
  3. Select compatible models: Use any model name recognized by ModelWrapper, including API models ("gpt-4o", "claude-3-5-sonnet") or local checkpoints ("Qwen2-VL-7B", "Llama-3.1-8B").

Summary

  • TextFlow provides a modular two-stage pipeline for flowchart understanding through the junyiye/textflow repository.
  • ModelWrapper (src/models/model_wrapper.py) offers a unified interface for both cloud and local vision-language models.
  • Vision Textualizer converts images to Mermaid/Graphviz/PlantUML text via load_textualizer_prompt and extract_representation.
  • Textual Reasoner answers questions using the generated representation through load_reasoner_prompt.
  • Enable tool-use by passing representation to generate_response for model-driven diagram execution.
  • You can integrate TextFlow into any Python application by importing these components directly without modifying core library files.

Frequently Asked Questions

Can I use local models instead of API-based LLMs with TextFlow?

Yes. The ModelWrapper class in src/models/model_wrapper.py automatically detects whether you provide an API model name (like "gpt-4o") or a local checkpoint name (like "Qwen2-VL-7B"). Local models are loaded via load_local_model in src/models/local_models.py, while API models route through load_api_model in src/models/api_models.py.

What output formats does the Vision Textualizer support?

According to the source code in src/textualizer.py and src/prompts/prompts.py, the Vision Textualizer supports Mermaid, Graphviz, and PlantUML syntax. You specify the desired format via the output_type parameter when calling load_textualizer_prompt(output_type).

How do I enable tool-use for the Textual Reasoner?

Pass tool_use=True to your reasoning function, which triggers the ModelWrapper to use generate_api_response_tool_use instead of the standard generation method. This allows the LLM to invoke execution tools on the Mermaid representation for structured validation, as implemented in src/models/model_wrapper.py.

Do I need to modify the TextFlow source code to integrate it into my application?

No. You can integrate TextFlow by importing the public classes and functions (such as ModelWrapper, load_textualizer_prompt, and load_reasoner_prompt) directly from the src directory. The modular design in junyiye/textflow allows you to compose the pipeline stages programmatically without altering the library internals.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →