What Local Models Are Supported by TextFlow and How to Load Them
TextFlow supports 15+ local models—including Llama 3.1, Mixtral, Phi-3.5, Qwen2.5, LLaVA, and Qwen2-VL—through the unified load_local_model interface in src/models/local_models.py, which automatically handles tokenizer initialization, device mapping, and dtype optimization.
TextFlow (junyiye/textflow) provides a streamlined interface for running large language models (LLMs) and vision-language models (VLMs) locally without external API dependencies. Understanding what local models are supported by TextFlow and how to load them enables you to leverage optimized inference pipelines for both text generation and multimodal reasoning tasks.
Supported Local Models in TextFlow
TextFlow categorizes supported models into Large Language Models (LLMs) for text generation and Vision-Language Models (VLMs) for image understanding tasks.
Large Language Models (LLMs)
The following LLMs are supported through the conditional logic in load_local_model (lines 24-34 of src/models/local_models.py):
- Llama-3.1-8B and Llama-3.1-70B – Meta's latest Llama 3.1 series models
- Mixtral-8x22B – Mistral's sparse mixture-of-experts architecture
- Phi-3.5-mini and Phi-3.5-MoE – Microsoft's compact and mixture-of-experts variants (require
trust_remote_code=True) - Qwen2.5-7B, Qwen2.5-14B, Qwen2.5-32B, and Qwen2.5-72B – Alibaba's Qwen 2.5 series across multiple parameter scales
Vision-Language Models (VLMs)
For multimodal inference combining vision and language, TextFlow supports the following models through the VLM branch of load_local_model (lines 46-66 of src/models/local_models.py):
- Llava-v1.6-110b – Large-scale LLaVA Next model for visual question answering
- Llama-3.2-11B and Llama-3.2-90B – Meta's vision-capable Llama 3.2 variants
- Qwen2-VL-7B and Qwen2-VL-72B – Qwen's vision-language models in two parameter sizes
How TextFlow Loads Local Models
The load_local_model function in src/models/local_models.py implements a unified loading mechanism that abstracts Hugging Face Transformers complexity while optimizing for local hardware.
Configuration Lookup
First, the function queries config["model_version"].get(model_name) to map logical names (e.g., "Llama-3.1-8B") to concrete Hugging Face repository identifiers. These mappings are defined in src/config.py.
Tokenizer and Model Initialization
Depending on the model category, TextFlow initializes different components:
For LLMs:
- Creates an
AutoTokenizerinstance for text preprocessing - Loads
AutoModelForCausalLMwith automatic device mapping - Applies
trust_remote_code=Truefor specific architectures (Phi-3.5 series)
For VLMs:
- Selects architecture-specific classes:
LlavaNextForConditionalGeneration,MllamaForConditionalGeneration, orQwen2VLForConditionalGeneration - Initializes corresponding processors (
LlavaNextProcessororAutoProcessor) - Handles vision-specific preprocessing requirements
Optimization Settings
All model loads specify device_map="auto" for automatic GPU/CPU distribution and torch_dtype="auto" (or explicit torch.float16/torch.bfloat16) for memory-efficient inference. The function returns a tuple (model, tokenizer_or_processor) ready for inference via generate_local_response.
Loading a Local Model: Code Example
To load any supported model, import load_local_model from the local models module and specify the logical model name:
from src.models.local_models import load_local_model
# Load an LLM (Llama 3.1 8B)
model_name = "Llama-3.1-8B"
model, tokenizer = load_local_model(model_name)
# Load a VLM (Qwen2-VL-7B)
vl_model_name = "Qwen2-VL-7B"
vl_model, processor = load_local_model(vl_model_name)
print(f"Loaded LLM: {type(model)}")
print(f"Loaded VLM: {type(vl_model)}")
The function returns a tuple of (model, tokenizer_or_processor) ready for inference. Use generate_local_response (defined in the same file) to run text generation or multimodal reasoning.
Key Source Files
| File | Role |
|---|---|
src/models/local_models.py |
Implements load_local_model and generate_local_response for all supported local models |
src/config.py |
Contains the config dictionary mapping logical model names to Hugging Face repository identifiers |
src/models/model_loader.py |
Provides higher-level utilities that may invoke load_local_model for CLI or API interfaces |
src/prompts/prompts.py |
Uses loaded models for prompt formatting and text generation tasks |
src/reasoner.py |
Implements reasoning pipelines using locally loaded LLMs |
src/evaluation.py |
Runs evaluation metrics using the loaded local model infrastructure |
Summary
- TextFlow supports 10+ LLMs (Llama 3.1, Mixtral, Phi-3.5, Qwen2.5) and 5+ VLMs (LLaVA, Llama 3.2 Vision, Qwen2-VL) through a unified interface.
- The
load_local_modelfunction insrc/models/local_models.pyhandles configuration lookup, tokenizer/processor initialization, and optimized model loading with automatic device mapping. - Loading requires only the logical model name (e.g.,
"Llama-3.1-8B") and returns a tuple of(model, tokenizer_or_processor)ready for inference. - Configuration mappings reside in
src/config.py, while higher-level orchestration usessrc/models/model_loader.py.
Frequently Asked Questions
What hardware requirements are needed for TextFlow local models?
Hardware requirements vary by model size. The 8B parameter models (Llama-3.1-8B, Qwen2.5-7B) typically require 16-24GB VRAM for full precision inference, while 70B+ models need multiple GPUs or significant quantization. TextFlow's device_map="auto" setting automatically distributes layers across available CUDA devices and falls back to CPU when GPU memory is insufficient.
Can I add custom local models to TextFlow?
Yes. To add a custom model, update the config["model_version"] dictionary in src/config.py with your model's logical name and Hugging Face repository identifier. Then ensure your model architecture is supported by the conditional logic in load_local_model within src/models/local_models.py. If using a novel architecture, you may need to add a new branch to handle the specific model class and processor initialization.
How does TextFlow handle vision-language model preprocessing?
TextFlow handles VLM preprocessing through architecture-specific processors initialized in load_local_model. For LLaVA models, it uses LlavaNextProcessor; for Llama 3.2 Vision, it uses MllamaForConditionalGeneration with AutoProcessor; and for Qwen2-VL, it uses Qwen2VLForConditionalGeneration. These processors handle image tokenization and multimodal prompt formatting automatically when you pass images to the generate_local_response function.
What is the difference between load_local_model and generate_local_response?
load_local_model is the initialization function that loads the model weights, configures the tokenizer or processor, and returns the model objects. It handles one-time setup operations like device mapping and dtype configuration. generate_local_response, defined in the same src/models/local_models.py file, is the inference function that takes the loaded model, tokenizer/processor, and input prompts (plus optional images for VLMs) to generate text outputs. You call load_local_model once at startup, then generate_local_response repeatedly during inference.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →