Performance Difference Between API Models and Local Models in TextFlow: A Complete Analysis

API models in TextFlow achieve state-of-the-art accuracy exceeding 80% on flowchart reasoning benchmarks through built-in tool-use support, while local models peak at 60-70% accuracy and lack graph query capabilities, creating a significant performance gap for complex visual reasoning tasks.

TextFlow is an open-source framework for flowchart understanding and question answering that supports both cloud-based API models and locally-hosted language models. Understanding the performance difference between these two approaches is critical for selecting the right backend for your specific accuracy, latency, and privacy requirements.

How TextFlow Implements API Models

API models in TextFlow are managed through src/models/api_models.py, which provides interfaces to OpenAI and Anthropic SDKs. The system supports high-performance models including gpt-4o, gpt-4o-mini, and claude-3-5-sonnet.

The key performance advantage comes from the generate_api_response_tool_use function (lines 65-124), which implements a graph-tool loop:

  • Automatically converts Mermaid diagram representations into executable Python graph code
  • Calls appropriate graph-query functions during the reasoning process
  • Feeds structured results back to the LLM for final answer generation

This tool-use capability allows API models to achieve state-of-the-art results on the FlowVQA and FlowLearn benchmarks, with exact-match accuracy exceeding 80% when tool-use is enabled.

How TextFlow Implements Local Models

Local models are handled in src/models/local_models.py, which loads models from the Hugging Face hub or local checkpoints. Supported architectures include Llama-3.1-8B, Mixtral-8x22B, Phi-3.5-mini, Qwen2-VL-7B, and Llava-v1.6-110b.

The generate_local_response function (lines 72-185) handles pure text generation, but lacks tool-use support. The repository includes a stub function generate_local_response_tool_use that explicitly raises an error when called, preventing local models from executing graph queries during reasoning.

Local models run on the user's hardware (CPU or GPU), which introduces several performance constraints:

  • Accuracy ceiling: The best-performing local models (such as Qwen2-VL-7B) achieve approximately 60-70% accuracy on FlowVQA, noticeably below API baselines
  • Hardware dependency: Inference quality depends on available GPU memory and may require lower-precision quantization
  • No graph reasoning loop: Without tool-use, local models cannot dynamically query graph structures during generation

Performance Benchmarks: API vs Local Models

The TextFlow evaluation suite, implemented in src/evaluation.py, quantifies the performance gap between deployment modes:

Model Type Best Performer FlowVQA Accuracy Tool-Use Support
API Models gpt-4o with tool-use ~80%+ Yes
Local Models Qwen2-VL-7B ~60-70% No

The 20 percentage point accuracy gap widens significantly on complex graph reasoning tasks that require multi-hop traversal or mathematical operations on flowchart nodes. API models leverage the tool-use loop to execute precise graph algorithms, while local models must rely solely on their parametric knowledge.

Why the Performance Gap Exists

Three architectural factors explain the performance difference between API and local models in TextFlow:

1. Model Scale and Instruction Tuning API models like GPT-4o and Claude 3.5 Sonnet utilize multi-billion parameter architectures trained on massive instruction-following datasets. Local models, while capable, typically lack the same depth of fine-tuning for complex visual reasoning tasks.

2. Tool-Use Architecture The critical differentiator is implemented in src/models/api_models.py: the generate_api_response_tool_use function creates a feedback loop where the LLM can generate Python code to query graph properties, execute that code, and incorporate results into subsequent reasoning steps. Local models cannot access this capability—the generate_local_response_tool_use stub in src/models/local_models.py raises a NotImplementedError.

3. Hardware and Inference Constraints API models run on optimized cloud infrastructure with guaranteed low-latency, high-throughput inference. Local models depend on consumer hardware, often requiring 4-bit or 8-bit quantization that can degrade reasoning performance, particularly for mathematical operations on flowchart nodes.

Code Examples: Calling API and Local Models

Using API Models with Tool-Use

The following example demonstrates how to leverage the high-performance API path with graph reasoning capabilities:

from models.api_models import load_api_model, generate_api_response_tool_use

# Initialize GPT-4o client

client = load_api_model("gpt-4o")

# Define conversation context

messages = [
    {"role": "user", "content": "How many nodes are in the flowchart?"}
]

# Mermaid diagram representation

mermaid_graph = """
graph TD
    A --> B
    B --> C
"""

# Execute with tool-use for accurate graph querying

response = generate_api_response_tool_use(
    model_name="gpt-4o",
    client=client,
    messages=messages,
    representation=mermaid_graph,
)

print("Answer:", response)

Source: src/models/api_models.py, lines 65-124.

Using Local Models for Pure Generation

For offline scenarios, use the local model path (note the absence of tool-use):

from models.local_models import load_local_model, generate_local_response

# Load Qwen2-VL-7B from Hugging Face

model, processor = load_local_model("Qwen2-VL-7B")

# Prepare multimodal input

messages = [
    {"role": "user",
     "content": [
         {"type": "text", "text": "Summarize the flowchart."},
         {"type": "image", "image": "path/to/flowchart.png"},
     ]
    }
]

# Generate response (no tool-use available)

answer = generate_local_response(
    model_name="Qwen2-VL-7B",
    model=model,
    tokenizer=processor,
    messages=messages,
)

print("Answer:", answer)

Source: src/models/local_models.py, lines 72-185.

Summary

  • API models (gpt-4o, claude-3-5-sonnet) deliver state-of-the-art accuracy exceeding 80% on TextFlow benchmarks through the generate_api_response_tool_use function, which enables dynamic graph querying during reasoning.

  • Local models (Qwen2-VL-7B, Llama-3.1-8B, etc.) achieve 60-70% accuracy and rely solely on parametric knowledge, as the generate_local_response_tool_use stub raises a NotImplementedError, preventing graph-tool loops.

  • The 20 percentage point performance gap stems from three factors: model scale and instruction tuning, exclusive tool-use architecture in src/models/api_models.py, and hardware constraints affecting local inference quality.

  • Choose API models for maximum accuracy on complex flowchart reasoning tasks; choose local models only when offline operation or data privacy requirements prohibit cloud API calls.

Frequently Asked Questions

Why do API models outperform local models in TextFlow?

API models outperform local models primarily due to the tool-use architecture implemented in src/models/api_models.py. The generate_api_response_tool_use function allows GPT-4o and Claude 3.5 to execute Python graph queries during reasoning, feeding accurate results back into the generation process. Local models lack this capability—the corresponding function in src/models/local_models.py raises a NotImplementedError. Additionally, API models benefit from larger parameter counts and superior instruction tuning for visual reasoning tasks.

Can I use tool-use features with local models in TextFlow?

No, tool-use is not available for local models in the current TextFlow implementation. While src/models/api_models.py contains the full generate_api_response_tool_use implementation that converts Mermaid diagrams into executable graph code, src/models/local_models.py only provides a stub function that raises an error when called. Local models must rely on pure text generation through generate_local_response, limiting their reasoning capabilities for complex graph traversal tasks.

What accuracy can I expect from local models versus API models?

According to the TextFlow evaluation suite in src/evaluation.py, API models achieve approximately 80%+ exact-match accuracy on the FlowVQA and FlowLearn benchmarks when tool-use is enabled. In contrast, local models typically reach 60-70% accuracy, with the best-performing open models like Qwen2-VL-7B hitting the upper end of that range. The 20-percentage-point gap widens on tasks requiring multi-hop graph reasoning or mathematical operations on flowchart nodes, where API models leverage dynamic graph queries.

Which local models are supported in TextFlow?

TextFlow supports a variety of local vision-language and text-only models through src/models/local_models.py, including Llama-3.1-8B, Mixtral-8x22B, Phi-3.5-mini, Qwen2-VL-7B, and Llava-v1.6-110b. These models are loaded from the Hugging Face hub or local checkpoints using the load_local_model function. However, all local models operate in pure generation mode without tool-use capabilities, and their performance depends heavily on available GPU memory and quantization settings configured in config.json.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →