How to Switch Between Different VLMs for the Vision Textualizer in TextFlow
You can switch between Vision-Language Models by passing the model name to the --textualizer CLI argument or the ModelWrapper constructor, which automatically routes to API or local inference pipelines.
The Vision Textualizer in the junyiye/textflow repository converts flowchart images into textual representations using Vision-Language Models (VLMs). Switching between different VLMs requires only a single parameter change, as the architecture abstracts model loading and inference behind a unified wrapper interface.
Architecture Overview
The Vision Textualizer acts as a thin orchestration layer that delegates model-specific operations to specialized loaders. When you switch VLMs, the system automatically determines whether to use a cloud API or local GPU inference based on the model identifier you provide.
CLI Argument Parsing
The entry point in src/textualizer.py exposes the --textualizer flag to accept your VLM selection. This argument defaults to Qwen2-VL-7B but accepts any supported model identifier.
parser.add_argument(
"--textualizer",
type=str,
default="Qwen2‑VL‑7B",
help="The VLM to generate the text representation.",
)
Source: src/textualizer.py, lines 27‑31
The parsed value passes directly to the ModelWrapper constructor at line 57:
model = ModelWrapper(textualizer)
ModelWrapper Routing
The ModelWrapper class in src/models/model_loader.py serves as the abstraction layer that routes your VLM choice to the appropriate backend. It distinguishes between API-based services and locally-hosted models by checking the model name against a hard-coded list.
# API models identified by name
api_models = ["claude-3-5-sonnet", "gpt-4o", "gpt-4o-mini"]
Source: src/models/model_loader.py, lines 15‑22
If your selected model appears in this list, ModelWrapper calls load_api_model() from src/models/api_models.py. Otherwise, it invokes load_local_model() from src/models/local_models.py.
Local vs API Model Loading
API Models use cloud endpoints and require authentication via environment variables. The generate_api_response* functions handle image encoding and HTTP requests.
Local VLMs load into GPU memory using HuggingFace transformers. The load_local_model() function in src/models/local_models.py contains conditional branches for each supported architecture, configuring the appropriate processor and model class.
Supported VLMs and How to Select Them
The framework supports both commercial API services and open-source local models. You switch between them by changing the string passed to --textualizer.
API-Based Models
These require valid API keys set in your environment:
- Claude 3.5 Sonnet:
claude-3-5-sonnet - GPT-4o:
gpt-4o - GPT-4o Mini:
gpt-4o-mini
Local VLMs
These download weights from HuggingFace and run on your hardware:
- Qwen2-VL-7B (default)
- Qwen2-VL-72B
- Llava-v1.6-110b
- Llama-3.2-11B
- **Llama-3.2-90B`
Source: src/models/local_models.py, lines 44‑66
Practical Code Examples
Command-Line Usage
Switch VLMs by modifying the --textualizer argument when running src/textualizer.py:
# Use the default Qwen2-VL-7B
python -m src.textualizer --dataset flowvqa --output_type mermaid
# Switch to Llava-v1.6-110b (local VLM)
python -m src.textualizer \
--dataset flowvqa \
--textualizer Llava-v1.6-110b \
--output_type mermaid
# Use GPT-4o API model
python -m src.textualizer \
--dataset flowvqa \
--textualizer gpt-4o \
--output_type mermaid
Programmatic Usage in Python
You can also switch VLMs directly in code by instantiating ModelWrapper with different model names:
from src.models.model_loader import ModelWrapper
from src.prompts import load_textualizer_prompt
from src.utils import extract_representation
# Choose a VLM name – any listed in local_models.load_local_model or an API model
vlm_name = "Qwen2-VL-7B" # or "Llava-v1.6-110b", "gpt-4o", etc.
# Initialise the wrapper
model = ModelWrapper(vlm_name)
# Prepare prompt and image
prompt = load_textualizer_prompt(output_type="mermaid")
image_path = "path/to/flowchart.png"
# Generate textual representation
response = model.generate_response(prompt, image_path=image_path)
representation = extract_representation(response)
print(representation)
Extending to a New VLM
To add support for a new VLM not currently in the supported list:
- Add the model name to the appropriate branch in
load_local_model()insrc/models/local_models.py(or create a new API loader insrc/models/api_models.py). - Implement the loading logic (model, processor/tokenizer) following the existing patterns for similar architectures.
- The
ModelWrapperwill automatically recognize the new name when passed via--textualizerwithout requiring changes to the core orchestration logic.
Key Implementation Files
| File | Role |
|---|---|
src/textualizer.py |
CLI entry point; parses --textualizer and drives the workflow. |
src/models/model_loader.py |
ModelWrapper class – decides API vs. local, loads model, forwards generation calls. |
src/models/local_models.py |
Implements load_local_model() and VLM inference paths for each supported local model. |
src/models/api_models.py |
Handles loading and calling cloud-based VLM APIs (Claude, GPT-4o). |
src/prompts/prompts.py |
Supplies the textualizer prompt used for VLM generation. |
src/utils.py |
Helper extract_representation that parses the VLM's raw response into the requested format. |
Summary
- Switch between VLMs by changing the
--textualizerCLI argument or passing a different model name to theModelWrapperconstructor. - The system automatically routes API models (
gpt-4o,claude-3-5-sonnet) to cloud endpoints and local models (Qwen2-VL,Llava,Llama-3.2) to HuggingFace transformers. - All model-specific logic is encapsulated in
src/models/model_loader.py, making it trivial to add new VLMs without modifying the core textualization pipeline.
Frequently Asked Questions
What is the default VLM used by the Vision Textualizer?
The default VLM is Qwen2-VL-7B, specified in the argument parser within src/textualizer.py at line 29. If you omit the --textualizer flag, the system automatically instantiates this model via the ModelWrapper class.
Can I use commercial API models like GPT-4o instead of local VLMs?
Yes. The ModelWrapper in src/models/model_loader.py recognizes gpt-4o, gpt-4o-mini, and claude-3-5-sonnet as API-based models. When you specify one of these identifiers, the system routes requests to src/models/api_models.py instead of loading weights locally.
How do I add support for a new VLM that is not currently listed?
To add a new VLM, modify the load_local_model() function in src/models/local_models.py (for local models) or create a new loader in src/models/api_models.py (for API services). Implement the model and processor initialization following the existing conditional patterns for similar architectures. Once added, ModelWrapper will automatically recognize the new model name without requiring changes to the CLI or core workflow.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →