# Understanding the Differences Between Reasoner Models in TextFlow's Textual Reasoner

> Explore TextFlow's Textual Reasoner: API vs local models. Understand tool-use differences between Claude 3.5-Sonnet, GPT-4o, Llama-3.1, and Mixtral for your NLP needs.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: deep-dive
- Published: 2026-03-05

---

**API-based models (Claude 3.5-Sonnet, GPT-4o) support full tool-use functionality and require external API keys, while local models (Llama-3.1, Mixtral) run entirely offline with lower latency but currently lack functional tool-use support.**

The Textual Reasoner in the junyiye/textflow repository provides a unified interface for interacting with large language models through a modular architecture. When running [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), users select their preferred backend via the `--reasoner` CLI argument, which defaults to `Llama-3.1-8B`. Understanding the architectural differences between API-based and local inference paths is essential for optimizing both capability and operational cost.

## Architecture and Model Selection

The entry point for model selection resides in [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), where the `--reasoner` argument determines which backend powers the reasoning pipeline. This value is passed directly to the `ModelWrapper` class defined in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py), which abstracts away the underlying implementation differences between cloud-hosted and local inference.

At initialization, `ModelWrapper` determines the execution path by checking if the requested model name exists in a predefined list of API models. According to the source code at lines 15-22 of [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py), this boolean flag (`self.is_api_model`) dictates whether to invoke `load_api_model()` or `load_local_model()`:

```python

# Conceptual representation of the initialization logic

if model_name in ["claude-3-5-sonnet-20240620", "gpt-4o", "gpt-4o-mini"]:
    self.is_api_model = True
    self.load_api_model()
else:
    self.is_api_model = False
    self.load_local_model()

```

## API-Based vs. Local Model Categories

### Supported Model portfolios

The available models are enumerated in [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) under the `model_version` key. The system distinguishes between two distinct categories:

- **API-based models**: Claude 3.5-Sonnet, GPT-4o, and GPT-4o-mini. These require active internet connectivity and authentication credentials.
- **Local (offline) models**: Llama-3.1-8B, Llama-3.1-70B, Mixtral-8x22B, Phi-3.5-mini/MoE, Qwen2.5 variants (7B/14B/32B/72B), Llama-3.2-Vision models (11B/90B), Llava-v1.6-110b, and Qwen2-VL (7B/72B).

### Loading Mechanisms

When `is_api_model` evaluates to true, `ModelWrapper` initializes an HTTP client for the provider's API endpoint. For local models, `load_local_model()` utilizes Hugging Face transformers to load model weights and tokenizers into local memory, either from cache or via on-demand download.

## Functional Differences in Inference

### Standard Response Generation

Both execution paths share a common preprocessing stage through `load_messages()`, which constructs the conversation history. However, the generation logic diverges significantly:

- **API path**: Invokes `generate_api_response()`, transmitting the message payload to the provider's endpoint and returning the remote inference result.
- **Local path**: Invokes `generate_local_response()`, executing the forward pass on local hardware via the Hugging Face generation pipeline.

### Critical Difference: Tool-Use Support

The most substantial functional gap between the two categories involves tool-use capabilities. When the `--tool_use` flag is enabled in [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) (lines 99-105), the system passes a `representation` argument to `ModelWrapper.generate_response()`.

According to the implementation at lines 30-51 of [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py), this triggers different execution branches:

- **API models**: Call `generate_api_response_tool_use()`, which sends the textual representation to the remote service and processes tool calls (such as image generation or retrieval) on the provider side.
- **Local models**: Call `generate_local_response_tool_use()`. However, **the current implementation does not return a value**—the function is invoked but its result is discarded, rendering tool-use a no-op for local models.

```python

# From src/models/model_loader.py (conceptual)

if self.is_api_model:
    return self.generate_api_response_tool_use(representation, messages)
else:
    self.generate_local_response_tool_use(representation, messages)  # Result discarded

    # Falls through to standard generation or returns None

```

## Operational Requirements and Constraints

### Authentication and Environment Variables

API-based models require valid credentials stored in environment variables or [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json): specifically `ANTHROPIC_API_KEY` for Claude models or `OPENAI_API_KEY` for GPT variants. The `ModelWrapper` never reads these keys directly; instead, the underlying provider libraries handle authentication transparently.

Local models require no external credentials, eliminating dependency on third-party API access and associated rate limits.

### Performance Characteristics

**API models** introduce network latency and are subject to provider-side throttling, but offer access to massively scaled, instruction-tuned models without local hardware constraints.

**Local models** execute on host CPU/GPU resources with deterministic latency dependent on hardware specifications. While avoiding network delays, performance is constrained by available VRAM and compute capacity, particularly when running larger variants like Llama-3.1-70B or Qwen2.5-72B.

## Configuration Interface

Both model categories utilize the same configuration block in [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) for generation parameters including `max_tokens`, `temperature`, and sampling controls. These parameters are applied either to the API request payload or to the Hugging Face `GenerationConfig`, ensuring consistent behavior across backends regardless of execution location.

## Summary

- **API-based models** (Claude 3.5-Sonnet, GPT-4o) provide full **tool-use functionality** and access to cloud-scale compute, but require **API keys** and incur network latency.
- **Local models** (Llama-3.1, Mixtral, Qwen2.5) run **entirely offline** without authentication, offering lower latency for standard inference, but currently **lack functional tool-use support** according to the source implementation.
- The `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py) handles execution path selection at lines 15-22, while response generation logic diverges at lines 33-51 based on the `is_api_model` boolean flag.
- Tool-use capabilities are available only for API models through `generate_api_response_tool_use()`, whereas the local equivalent `generate_local_response_tool_use()` is currently non-functional in the codebase.

## Frequently Asked Questions

### How do I switch between API and local models when running the Textual Reasoner?

Use the `--reasoner` CLI argument when executing [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py). For example, `--reasoner gpt-4o` selects the cloud-based GPT-4o model, while `--reasoner Llama-3.1-8B` (or omitting the flag, as it is the default) selects local inference. The `ModelWrapper` automatically detects the category and initializes the appropriate loading mechanism.

### Why doesn't tool-use work with local models in TextFlow?

As implemented in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py), the `generate_local_response_tool_use()` function is called when tool-use is enabled for local models, but its return value is discarded (lines 33-51). This makes the tool-use feature effectively a no-op for local backends. In contrast, API models properly implement tool-use through `generate_api_response_tool_use()`, which processes the `representation` argument and handles tool calls on the provider's infrastructure.

### Which models can I use without an internet connection?

All models categorized as "local" in [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) support offline operation, including the Llama family (3.1-8B, 3.1-70B, 3.2-Vision variants), Mixtral-8x22B, Phi-3.5 models, Qwen2.5 series (7B through 72B), Llava-v1.6-110b, and Qwen2-VL models. These are loaded via `load_local_model()` using Hugging Face transformers and require only local hardware resources.

### Do API-based models and local models share the same configuration options?

Yes. Both categories read from the same `model_config` block in [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) for parameters like `max_tokens` and `temperature`. The `ModelWrapper` applies these settings either to the HTTP API request payload (for cloud models) or to the Hugging Face generation pipeline (for local models), ensuring consistent generation behavior across different backends.