How to Switch Between Different VLMs for the Vision Textualizer in TextFlow

You can switch between Vision-Language Models by passing the model name to the --textualizer CLI argument or the ModelWrapper constructor, which automatically routes to API or local inference pipelines.

The Vision Textualizer in the junyiye/textflow repository converts flowchart images into textual representations using Vision-Language Models (VLMs). Switching between different VLMs requires only a single parameter change, as the architecture abstracts model loading and inference behind a unified wrapper interface.

Architecture Overview

The Vision Textualizer acts as a thin orchestration layer that delegates model-specific operations to specialized loaders. When you switch VLMs, the system automatically determines whether to use a cloud API or local GPU inference based on the model identifier you provide.

CLI Argument Parsing

The entry point in src/textualizer.py exposes the --textualizer flag to accept your VLM selection. This argument defaults to Qwen2-VL-7B but accepts any supported model identifier.

parser.add_argument(
    "--textualizer",
    type=str,
    default="Qwen2‑VL‑7B",
    help="The VLM to generate the text representation.",
)

Source: src/textualizer.py, lines 27‑31

The parsed value passes directly to the ModelWrapper constructor at line 57:

model = ModelWrapper(textualizer)

ModelWrapper Routing

The ModelWrapper class in src/models/model_loader.py serves as the abstraction layer that routes your VLM choice to the appropriate backend. It distinguishes between API-based services and locally-hosted models by checking the model name against a hard-coded list.


# API models identified by name

api_models = ["claude-3-5-sonnet", "gpt-4o", "gpt-4o-mini"]

Source: src/models/model_loader.py, lines 15‑22

If your selected model appears in this list, ModelWrapper calls load_api_model() from src/models/api_models.py. Otherwise, it invokes load_local_model() from src/models/local_models.py.

Local vs API Model Loading

API Models use cloud endpoints and require authentication via environment variables. The generate_api_response* functions handle image encoding and HTTP requests.

Local VLMs load into GPU memory using HuggingFace transformers. The load_local_model() function in src/models/local_models.py contains conditional branches for each supported architecture, configuring the appropriate processor and model class.

Supported VLMs and How to Select Them

The framework supports both commercial API services and open-source local models. You switch between them by changing the string passed to --textualizer.

API-Based Models

These require valid API keys set in your environment:

  • Claude 3.5 Sonnet: claude-3-5-sonnet
  • GPT-4o: gpt-4o
  • GPT-4o Mini: gpt-4o-mini

Local VLMs

These download weights from HuggingFace and run on your hardware:

  • Qwen2-VL-7B (default)
  • Qwen2-VL-72B
  • Llava-v1.6-110b
  • Llama-3.2-11B
  • **Llama-3.2-90B`

Source: src/models/local_models.py, lines 44‑66

Practical Code Examples

Command-Line Usage

Switch VLMs by modifying the --textualizer argument when running src/textualizer.py:


# Use the default Qwen2-VL-7B

python -m src.textualizer --dataset flowvqa --output_type mermaid

# Switch to Llava-v1.6-110b (local VLM)

python -m src.textualizer \
    --dataset flowvqa \
    --textualizer Llava-v1.6-110b \
    --output_type mermaid

# Use GPT-4o API model

python -m src.textualizer \
    --dataset flowvqa \
    --textualizer gpt-4o \
    --output_type mermaid

Programmatic Usage in Python

You can also switch VLMs directly in code by instantiating ModelWrapper with different model names:

from src.models.model_loader import ModelWrapper
from src.prompts import load_textualizer_prompt
from src.utils import extract_representation

# Choose a VLM name – any listed in local_models.load_local_model or an API model

vlm_name = "Qwen2-VL-7B"          # or "Llava-v1.6-110b", "gpt-4o", etc.

# Initialise the wrapper

model = ModelWrapper(vlm_name)

# Prepare prompt and image

prompt = load_textualizer_prompt(output_type="mermaid")
image_path = "path/to/flowchart.png"

# Generate textual representation

response = model.generate_response(prompt, image_path=image_path)
representation = extract_representation(response)

print(representation)

Extending to a New VLM

To add support for a new VLM not currently in the supported list:

  1. Add the model name to the appropriate branch in load_local_model() in src/models/local_models.py (or create a new API loader in src/models/api_models.py).
  2. Implement the loading logic (model, processor/tokenizer) following the existing patterns for similar architectures.
  3. The ModelWrapper will automatically recognize the new name when passed via --textualizer without requiring changes to the core orchestration logic.

Key Implementation Files

File Role
src/textualizer.py CLI entry point; parses --textualizer and drives the workflow.
src/models/model_loader.py ModelWrapper class – decides API vs. local, loads model, forwards generation calls.
src/models/local_models.py Implements load_local_model() and VLM inference paths for each supported local model.
src/models/api_models.py Handles loading and calling cloud-based VLM APIs (Claude, GPT-4o).
src/prompts/prompts.py Supplies the textualizer prompt used for VLM generation.
src/utils.py Helper extract_representation that parses the VLM's raw response into the requested format.

Summary

  • Switch between VLMs by changing the --textualizer CLI argument or passing a different model name to the ModelWrapper constructor.
  • The system automatically routes API models (gpt-4o, claude-3-5-sonnet) to cloud endpoints and local models (Qwen2-VL, Llava, Llama-3.2) to HuggingFace transformers.
  • All model-specific logic is encapsulated in src/models/model_loader.py, making it trivial to add new VLMs without modifying the core textualization pipeline.

Frequently Asked Questions

What is the default VLM used by the Vision Textualizer?

The default VLM is Qwen2-VL-7B, specified in the argument parser within src/textualizer.py at line 29. If you omit the --textualizer flag, the system automatically instantiates this model via the ModelWrapper class.

Can I use commercial API models like GPT-4o instead of local VLMs?

Yes. The ModelWrapper in src/models/model_loader.py recognizes gpt-4o, gpt-4o-mini, and claude-3-5-sonnet as API-based models. When you specify one of these identifiers, the system routes requests to src/models/api_models.py instead of loading weights locally.

How do I add support for a new VLM that is not currently listed?

To add a new VLM, modify the load_local_model() function in src/models/local_models.py (for local models) or create a new loader in src/models/api_models.py (for API services). Implement the model and processor initialization following the existing conditional patterns for similar architectures. Once added, ModelWrapper will automatically recognize the new model name without requiring changes to the CLI or core workflow.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →