# Performance Difference Between API Models and Local Models in TextFlow: A Complete Analysis

> Analyze TextFlow performance: API models offer 80%+ accuracy with tool use, while local models reach 60-70% lacking graph queries. Understand the gap for visual reasoning.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: performance
- Published: 2026-03-05

---

**API models in TextFlow achieve state-of-the-art accuracy exceeding 80% on flowchart reasoning benchmarks through built-in tool-use support, while local models peak at 60-70% accuracy and lack graph query capabilities, creating a significant performance gap for complex visual reasoning tasks.**

TextFlow is an open-source framework for flowchart understanding and question answering that supports both cloud-based API models and locally-hosted language models. Understanding the performance difference between these two approaches is critical for selecting the right backend for your specific accuracy, latency, and privacy requirements.

## How TextFlow Implements API Models

API models in TextFlow are managed through **[`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py)**, which provides interfaces to OpenAI and Anthropic SDKs. The system supports high-performance models including `gpt-4o`, `gpt-4o-mini`, and `claude-3-5-sonnet`.

The key performance advantage comes from the **`generate_api_response_tool_use`** function (lines 65-124), which implements a graph-tool loop:

- Automatically converts Mermaid diagram representations into executable Python graph code
- Calls appropriate graph-query functions during the reasoning process
- Feeds structured results back to the LLM for final answer generation

This tool-use capability allows API models to achieve **state-of-the-art results** on the FlowVQA and FlowLearn benchmarks, with exact-match accuracy exceeding **80%** when tool-use is enabled.

## How TextFlow Implements Local Models

Local models are handled in **[`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py)**, which loads models from the Hugging Face hub or local checkpoints. Supported architectures include `Llama-3.1-8B`, `Mixtral-8x22B`, `Phi-3.5-mini`, `Qwen2-VL-7B`, and `Llava-v1.6-110b`.

The **`generate_local_response`** function (lines 72-185) handles pure text generation, but **lacks tool-use support**. The repository includes a stub function `generate_local_response_tool_use` that explicitly raises an error when called, preventing local models from executing graph queries during reasoning.

Local models run on the user's hardware (CPU or GPU), which introduces several performance constraints:

- **Accuracy ceiling**: The best-performing local models (such as `Qwen2-VL-7B`) achieve approximately **60-70%** accuracy on FlowVQA, noticeably below API baselines
- **Hardware dependency**: Inference quality depends on available GPU memory and may require lower-precision quantization
- **No graph reasoning loop**: Without tool-use, local models cannot dynamically query graph structures during generation

## Performance Benchmarks: API vs Local Models

The TextFlow evaluation suite, implemented in **[`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py)**, quantifies the performance gap between deployment modes:

| Model Type | Best Performer | FlowVQA Accuracy | Tool-Use Support |
|------------|---------------|------------------|------------------|
| **API Models** | `gpt-4o` with tool-use | **~80%+** | Yes |
| **Local Models** | `Qwen2-VL-7B` | **~60-70%** | No |

The **20 percentage point accuracy gap** widens significantly on complex graph reasoning tasks that require multi-hop traversal or mathematical operations on flowchart nodes. API models leverage the tool-use loop to execute precise graph algorithms, while local models must rely solely on their parametric knowledge.

## Why the Performance Gap Exists

Three architectural factors explain the performance difference between API and local models in TextFlow:

**1. Model Scale and Instruction Tuning**
API models like GPT-4o and Claude 3.5 Sonnet utilize multi-billion parameter architectures trained on massive instruction-following datasets. Local models, while capable, typically lack the same depth of fine-tuning for complex visual reasoning tasks.

**2. Tool-Use Architecture**
The critical differentiator is implemented in **[`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py)**: the `generate_api_response_tool_use` function creates a feedback loop where the LLM can generate Python code to query graph properties, execute that code, and incorporate results into subsequent reasoning steps. Local models cannot access this capability—the `generate_local_response_tool_use` stub in **[`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py)** raises a `NotImplementedError`.

**3. Hardware and Inference Constraints**
API models run on optimized cloud infrastructure with guaranteed low-latency, high-throughput inference. Local models depend on consumer hardware, often requiring 4-bit or 8-bit quantization that can degrade reasoning performance, particularly for mathematical operations on flowchart nodes.

## Code Examples: Calling API and Local Models

### Using API Models with Tool-Use

The following example demonstrates how to leverage the high-performance API path with graph reasoning capabilities:

```python
from models.api_models import load_api_model, generate_api_response_tool_use

# Initialize GPT-4o client

client = load_api_model("gpt-4o")

# Define conversation context

messages = [
    {"role": "user", "content": "How many nodes are in the flowchart?"}
]

# Mermaid diagram representation

mermaid_graph = """
graph TD
    A --> B
    B --> C
"""

# Execute with tool-use for accurate graph querying

response = generate_api_response_tool_use(
    model_name="gpt-4o",
    client=client,
    messages=messages,
    representation=mermaid_graph,
)

print("Answer:", response)

```

*Source:* **[`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py)**, lines 65-124.

### Using Local Models for Pure Generation

For offline scenarios, use the local model path (note the absence of tool-use):

```python
from models.local_models import load_local_model, generate_local_response

# Load Qwen2-VL-7B from Hugging Face

model, processor = load_local_model("Qwen2-VL-7B")

# Prepare multimodal input

messages = [
    {"role": "user",
     "content": [
         {"type": "text", "text": "Summarize the flowchart."},
         {"type": "image", "image": "path/to/flowchart.png"},
     ]
    }
]

# Generate response (no tool-use available)

answer = generate_local_response(
    model_name="Qwen2-VL-7B",
    model=model,
    tokenizer=processor,
    messages=messages,
)

print("Answer:", answer)

```

*Source:* **[`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py)**, lines 72-185.

## Summary

- **API models** (`gpt-4o`, `claude-3-5-sonnet`) deliver **state-of-the-art accuracy exceeding 80%** on TextFlow benchmarks through the **`generate_api_response_tool_use`** function, which enables dynamic graph querying during reasoning.

- **Local models** (`Qwen2-VL-7B`, `Llama-3.1-8B`, etc.) achieve **60-70% accuracy** and rely solely on parametric knowledge, as the **`generate_local_response_tool_use`** stub raises a `NotImplementedError`, preventing graph-tool loops.

- The **20 percentage point performance gap** stems from three factors: model scale and instruction tuning, exclusive tool-use architecture in **[`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py)**, and hardware constraints affecting local inference quality.

- Choose **API models** for maximum accuracy on complex flowchart reasoning tasks; choose **local models** only when offline operation or data privacy requirements prohibit cloud API calls.

## Frequently Asked Questions

### Why do API models outperform local models in TextFlow?

API models outperform local models primarily due to the **tool-use architecture** implemented in **[`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py)**. The `generate_api_response_tool_use` function allows GPT-4o and Claude 3.5 to execute Python graph queries during reasoning, feeding accurate results back into the generation process. Local models lack this capability—the corresponding function in **[`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py)** raises a `NotImplementedError`. Additionally, API models benefit from larger parameter counts and superior instruction tuning for visual reasoning tasks.

### Can I use tool-use features with local models in TextFlow?

No, **tool-use is not available for local models** in the current TextFlow implementation. While **[`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py)** contains the full `generate_api_response_tool_use` implementation that converts Mermaid diagrams into executable graph code, **[`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py)** only provides a stub function that raises an error when called. Local models must rely on pure text generation through `generate_local_response`, limiting their reasoning capabilities for complex graph traversal tasks.

### What accuracy can I expect from local models versus API models?

According to the TextFlow evaluation suite in **[`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py)**, **API models achieve approximately 80%+ exact-match accuracy** on the FlowVQA and FlowLearn benchmarks when tool-use is enabled. In contrast, **local models typically reach 60-70% accuracy**, with the best-performing open models like `Qwen2-VL-7B` hitting the upper end of that range. The 20-percentage-point gap widens on tasks requiring multi-hop graph reasoning or mathematical operations on flowchart nodes, where API models leverage dynamic graph queries.

### Which local models are supported in TextFlow?

TextFlow supports a variety of local vision-language and text-only models through **[`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py)**, including `Llama-3.1-8B`, `Mixtral-8x22B`, `Phi-3.5-mini`, `Qwen2-VL-7B`, and `Llava-v1.6-110b`. These models are loaded from the Hugging Face hub or local checkpoints using the `load_local_model` function. However, all local models operate in pure generation mode without tool-use capabilities, and their performance depends heavily on available GPU memory and quantization settings configured in [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json).