# What Local Models Are Supported by TextFlow and How to Load Them

> Explore TextFlow's support for 15+ local models like Llama 3.1 and Mixtral. Learn how to easily load them using the unified load_local_model interface with automatic optimization.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: how-to-guide
- Published: 2026-03-05

---

**TextFlow supports 15+ local models—including Llama 3.1, Mixtral, Phi-3.5, Qwen2.5, LLaVA, and Qwen2-VL—through the unified `load_local_model` interface in [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py), which automatically handles tokenizer initialization, device mapping, and dtype optimization.**

TextFlow (junyiye/textflow) provides a streamlined interface for running large language models (LLMs) and vision-language models (VLMs) locally without external API dependencies. Understanding what local models are supported by TextFlow and how to load them enables you to leverage optimized inference pipelines for both text generation and multimodal reasoning tasks.

## Supported Local Models in TextFlow

TextFlow categorizes supported models into **Large Language Models (LLMs)** for text generation and **Vision-Language Models (VLMs)** for image understanding tasks.

### Large Language Models (LLMs)

The following LLMs are supported through the conditional logic in `load_local_model` (lines 24-34 of [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py)):

- **Llama-3.1-8B** and **Llama-3.1-70B** – Meta's latest Llama 3.1 series models
- **Mixtral-8x22B** – Mistral's sparse mixture-of-experts architecture
- **Phi-3.5-mini** and **Phi-3.5-MoE** – Microsoft's compact and mixture-of-experts variants (require `trust_remote_code=True`)
- **Qwen2.5-7B**, **Qwen2.5-14B**, **Qwen2.5-32B**, and **Qwen2.5-72B** – Alibaba's Qwen 2.5 series across multiple parameter scales

### Vision-Language Models (VLMs)

For multimodal inference combining vision and language, TextFlow supports the following models through the VLM branch of `load_local_model` (lines 46-66 of [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py)):

- **Llava-v1.6-110b** – Large-scale LLaVA Next model for visual question answering
- **Llama-3.2-11B** and **Llama-3.2-90B** – Meta's vision-capable Llama 3.2 variants
- **Qwen2-VL-7B** and **Qwen2-VL-72B** – Qwen's vision-language models in two parameter sizes

## How TextFlow Loads Local Models

The `load_local_model` function in [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py) implements a unified loading mechanism that abstracts Hugging Face Transformers complexity while optimizing for local hardware.

### Configuration Lookup

First, the function queries `config["model_version"].get(model_name)` to map logical names (e.g., `"Llama-3.1-8B"`) to concrete Hugging Face repository identifiers. These mappings are defined in [`src/config.py`](https://github.com/junyiye/textflow/blob/main/src/config.py).

### Tokenizer and Model Initialization

Depending on the model category, TextFlow initializes different components:

**For LLMs:**
- Creates an `AutoTokenizer` instance for text preprocessing
- Loads `AutoModelForCausalLM` with automatic device mapping
- Applies `trust_remote_code=True` for specific architectures (Phi-3.5 series)

**For VLMs:**
- Selects architecture-specific classes: `LlavaNextForConditionalGeneration`, `MllamaForConditionalGeneration`, or `Qwen2VLForConditionalGeneration`
- Initializes corresponding processors (`LlavaNextProcessor` or `AutoProcessor`)
- Handles vision-specific preprocessing requirements

### Optimization Settings

All model loads specify `device_map="auto"` for automatic GPU/CPU distribution and `torch_dtype="auto"` (or explicit `torch.float16`/`torch.bfloat16`) for memory-efficient inference. The function returns a tuple `(model, tokenizer_or_processor)` ready for inference via `generate_local_response`.

## Loading a Local Model: Code Example

To load any supported model, import `load_local_model` from the local models module and specify the logical model name:

```python
from src.models.local_models import load_local_model

# Load an LLM (Llama 3.1 8B)

model_name = "Llama-3.1-8B"
model, tokenizer = load_local_model(model_name)

# Load a VLM (Qwen2-VL-7B)

vl_model_name = "Qwen2-VL-7B"
vl_model, processor = load_local_model(vl_model_name)

print(f"Loaded LLM: {type(model)}")
print(f"Loaded VLM: {type(vl_model)}")

```

The function returns a tuple of `(model, tokenizer_or_processor)` ready for inference. Use `generate_local_response` (defined in the same file) to run text generation or multimodal reasoning.

## Key Source Files

| File | Role |
|------|------|
| **[`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py)** | Implements `load_local_model` and `generate_local_response` for all supported local models |
| **[`src/config.py`](https://github.com/junyiye/textflow/blob/main/src/config.py)** | Contains the `config` dictionary mapping logical model names to Hugging Face repository identifiers |
| **[`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py)** | Provides higher-level utilities that may invoke `load_local_model` for CLI or API interfaces |
| **[`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py)** | Uses loaded models for prompt formatting and text generation tasks |
| **[`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py)** | Implements reasoning pipelines using locally loaded LLMs |
| **[`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py)** | Runs evaluation metrics using the loaded local model infrastructure |

## Summary

- TextFlow supports **10+ LLMs** (Llama 3.1, Mixtral, Phi-3.5, Qwen2.5) and **5+ VLMs** (LLaVA, Llama 3.2 Vision, Qwen2-VL) through a unified interface.
- The **`load_local_model`** function in [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py) handles configuration lookup, tokenizer/processor initialization, and optimized model loading with automatic device mapping.
- Loading requires only the logical model name (e.g., `"Llama-3.1-8B"`) and returns a tuple of `(model, tokenizer_or_processor)` ready for inference.
- Configuration mappings reside in [`src/config.py`](https://github.com/junyiye/textflow/blob/main/src/config.py), while higher-level orchestration uses [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py).

## Frequently Asked Questions

### What hardware requirements are needed for TextFlow local models?

Hardware requirements vary by model size. The 8B parameter models (Llama-3.1-8B, Qwen2.5-7B) typically require 16-24GB VRAM for full precision inference, while 70B+ models need multiple GPUs or significant quantization. TextFlow's `device_map="auto"` setting automatically distributes layers across available CUDA devices and falls back to CPU when GPU memory is insufficient.

### Can I add custom local models to TextFlow?

Yes. To add a custom model, update the `config["model_version"]` dictionary in [`src/config.py`](https://github.com/junyiye/textflow/blob/main/src/config.py) with your model's logical name and Hugging Face repository identifier. Then ensure your model architecture is supported by the conditional logic in `load_local_model` within [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py). If using a novel architecture, you may need to add a new branch to handle the specific model class and processor initialization.

### How does TextFlow handle vision-language model preprocessing?

TextFlow handles VLM preprocessing through architecture-specific processors initialized in `load_local_model`. For LLaVA models, it uses `LlavaNextProcessor`; for Llama 3.2 Vision, it uses `MllamaForConditionalGeneration` with `AutoProcessor`; and for Qwen2-VL, it uses `Qwen2VLForConditionalGeneration`. These processors handle image tokenization and multimodal prompt formatting automatically when you pass images to the `generate_local_response` function.

### What is the difference between load_local_model and generate_local_response?

`load_local_model` is the initialization function that loads the model weights, configures the tokenizer or processor, and returns the model objects. It handles one-time setup operations like device mapping and dtype configuration. `generate_local_response`, defined in the same [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py) file, is the inference function that takes the loaded model, tokenizer/processor, and input prompts (plus optional images for VLMs) to generate text outputs. You call `load_local_model` once at startup, then `generate_local_response` repeatedly during inference.