How to Integrate TextFlow with Your Own Applications: Complete Developer Guide
Integrate TextFlow into your Python applications by importing its modular pipeline components—ModelWrapper, Vision Textualizer, and Textual Reasoner—to programmatically convert flowchart images into structured text representations and perform visual question answering.
The junyiye/textflow repository provides a modular framework for flowchart understanding that separates visual perception from textual reasoning. By leveraging its standardized interfaces, you can embed TextFlow's two-stage pipeline into web services, desktop applications, or automated evaluation workflows without modifying the core library source code.
Understanding the TextFlow Architecture
TextFlow employs a two-stage pipeline that decouples visual understanding from logical reasoning. This separation allows you to swap vision-language models (VLMs) and large language models (LLMs) independently while maintaining consistent data flows.
The pipeline consists of these core components:
- Configuration (
src/config.py): Loads paths, logging settings, and model credentials fromconfig.json. - ModelWrapper (
src/models/model_wrapper.py): Provides a unified interface for both API-based and local models, handling image encoding and tool-use logic. - Vision Textualizer (
src/textualizer.py): Converts flowchart images into text representations (Mermaid, Graphviz, or PlantUML) via VLM inference. - Textual Reasoner (
src/reasoner.py): Consumes the generated representation to answer visual questions using an LLM, with optional tool-use support. - Prompt Templates (
src/prompts/prompts.py): Centralized prompt strings for both pipeline stages. - Utility Functions (
src/utils.py): Helper methods for image encoding (encode_image) and representation extraction (extract_representation).
Programmatic Integration Guide
You can embed TextFlow directly into your codebase by importing the core classes and bypassing the command-line interface. The following examples assume you have installed dependencies via pip install -r requirements.txt and configured config.json with your API keys.
Step 1 — Generate Text Representations from Flowchart Images
To convert an image into a structured text format like Mermaid syntax, instantiate ModelWrapper and call the textualizer pipeline:
from src.models.model_wrapper import ModelWrapper
from src.prompts.prompts import load_textualizer_prompt
from src.utils import extract_representation
from src.config import config
import os
import json
def generate_mermaid(image_path, textualizer="gpt-4o"):
model = ModelWrapper(textualizer)
prompt = load_textualizer_prompt("mermaid")
# Generate response with image encoding handled internally
response = model.generate_response(prompt, image_path=image_path)
representation = extract_representation(response)
# Optional: persist to output directory matching CLI structure
out_dir = os.path.join(
config["file_paths"]["output"],
"flowvqa",
"mermaid"
)
os.makedirs(out_dir, exist_ok=True)
out_file = os.path.join(out_dir, f"{textualizer}.json")
with open(out_file, "w") as f:
json.dump({os.path.basename(image_path): representation}, f, indent=2)
return representation
This function encodes the image, sends it to the specified VLM (e.g., "gpt-4o" or "Qwen2-VL-7B"), and extracts the diagram code using the utility function extract_representation.
Step 2 — Answer Visual Questions Using Text Representations
Once you have the text representation (Mermaid code), pass it to the Textual Reasoner along with your question:
from src.models.model_wrapper import ModelWrapper
from src.prompts.prompts import load_reasoner_prompt
def answer_question(question, mermaid_code, reasoner="gpt-4o", tool_use=False):
model = ModelWrapper(reasoner)
prompt = load_reasoner_prompt(question, mermaid_code)
if tool_use:
# Enables model-driven tool execution (e.g., Mermaid rendering)
response = model.generate_response(
prompt,
representation=mermaid_code
)
else:
response = model.generate_response(prompt)
return response
When tool_use=True, the ModelWrapper invokes generate_api_response_tool_use (implemented in src/models/model_wrapper.py), allowing the model to execute the diagram representation for structured validation.
Step 3 — Build an End-to-End Integration Pipeline
Combine both stages into a single callable function for your application:
def textflow_pipeline(image_path, question,
textualizer="gpt-4o",
reasoner="gpt-4o",
tool_use=False):
"""
End-to-end TextFlow integration for flowchart VQA.
"""
# Stage 1: Visual understanding
representation = generate_mermaid(image_path, textualizer)
# Stage 2: Logical reasoning
answer = answer_question(
question,
representation,
reasoner=reasoner,
tool_use=tool_use
)
return answer
# Usage in your application
result = textflow_pipeline(
"diagrams/process_flow.png",
"What is the purpose of step 3?",
tool_use=True
)
Key Source Files for Integration
When integrating TextFlow, reference these specific files in the junyiye/textflow repository:
src/config.py: Loadsconfig.jsoncontaining file paths and model credentials.src/models/model_wrapper.py: The primary integration point; abstracts API calls (load_api_model) and local checkpoint loading (load_local_model).src/models/api_models.py: Low-level OpenAI and Anthropic API implementations used byModelWrapper.src/models/local_models.py: Local model initialization for Llama-3.1, Qwen2-VL, and other Hugging Face models.src/textualizer.py: CLI entry point for the Vision Textualizer stage; useful for understanding the default execution flow.src/reasoner.py: CLI entry point for the Textual Reasoner stage; demonstrates--tool_useflag handling.src/flowchart.py: In-memory flowchart representation and syntax conversion utilities.src/utils.py: Image encoding and text extraction helper functions.
Configuration Requirements
Before running the integration code:
- Configure
config.jsonwith your API keys (OpenAI, Anthropic) and adjustfile_pathsfor your environment. - Install dependencies:
pip install -r requirements.txtincludes the required ML and API client libraries. - Select compatible models: Use any model name recognized by
ModelWrapper, including API models ("gpt-4o","claude-3-5-sonnet") or local checkpoints ("Qwen2-VL-7B","Llama-3.1-8B").
Summary
- TextFlow provides a modular two-stage pipeline for flowchart understanding through the
junyiye/textflowrepository. ModelWrapper(src/models/model_wrapper.py) offers a unified interface for both cloud and local vision-language models.- Vision Textualizer converts images to Mermaid/Graphviz/PlantUML text via
load_textualizer_promptandextract_representation. - Textual Reasoner answers questions using the generated representation through
load_reasoner_prompt. - Enable tool-use by passing
representationtogenerate_responsefor model-driven diagram execution. - You can integrate TextFlow into any Python application by importing these components directly without modifying core library files.
Frequently Asked Questions
Can I use local models instead of API-based LLMs with TextFlow?
Yes. The ModelWrapper class in src/models/model_wrapper.py automatically detects whether you provide an API model name (like "gpt-4o") or a local checkpoint name (like "Qwen2-VL-7B"). Local models are loaded via load_local_model in src/models/local_models.py, while API models route through load_api_model in src/models/api_models.py.
What output formats does the Vision Textualizer support?
According to the source code in src/textualizer.py and src/prompts/prompts.py, the Vision Textualizer supports Mermaid, Graphviz, and PlantUML syntax. You specify the desired format via the output_type parameter when calling load_textualizer_prompt(output_type).
How do I enable tool-use for the Textual Reasoner?
Pass tool_use=True to your reasoning function, which triggers the ModelWrapper to use generate_api_response_tool_use instead of the standard generation method. This allows the LLM to invoke execution tools on the Mermaid representation for structured validation, as implemented in src/models/model_wrapper.py.
Do I need to modify the TextFlow source code to integrate it into my application?
No. You can integrate TextFlow by importing the public classes and functions (such as ModelWrapper, load_textualizer_prompt, and load_reasoner_prompt) directly from the src directory. The modular design in junyiye/textflow allows you to compose the pipeline stages programmatically without altering the library internals.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →