# How to Switch Between Different VLMs for the Vision Textualizer in TextFlow

> Easily switch Vision-Language Models for TextFlow's Vision Textualizer. Use the --textualizer CLI argument or ModelWrapper constructor for seamless API or local inference.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: how-to-guide
- Published: 2026-03-05

---

**You can switch between Vision-Language Models by passing the model name to the `--textualizer` CLI argument or the `ModelWrapper` constructor, which automatically routes to API or local inference pipelines.**

The Vision Textualizer in the `junyiye/textflow` repository converts flowchart images into textual representations using Vision-Language Models (VLMs). Switching between different VLMs requires only a single parameter change, as the architecture abstracts model loading and inference behind a unified wrapper interface.

## Architecture Overview

The Vision Textualizer acts as a thin orchestration layer that delegates model-specific operations to specialized loaders. When you switch VLMs, the system automatically determines whether to use a cloud API or local GPU inference based on the model identifier you provide.

### CLI Argument Parsing

The entry point in [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) exposes the `--textualizer` flag to accept your VLM selection. This argument defaults to `Qwen2-VL-7B` but accepts any supported model identifier.

```python
parser.add_argument(
    "--textualizer",
    type=str,
    default="Qwen2‑VL‑7B",
    help="The VLM to generate the text representation.",
)

```

*Source: [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), lines 27‑31*

The parsed value passes directly to the `ModelWrapper` constructor at line 57:

```python
model = ModelWrapper(textualizer)

```

### ModelWrapper Routing

The `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py) serves as the abstraction layer that routes your VLM choice to the appropriate backend. It distinguishes between API-based services and locally-hosted models by checking the model name against a hard-coded list.

```python

# API models identified by name

api_models = ["claude-3-5-sonnet", "gpt-4o", "gpt-4o-mini"]

```

*Source: [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py), lines 15‑22*

If your selected model appears in this list, `ModelWrapper` calls `load_api_model()` from [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py). Otherwise, it invokes `load_local_model()` from [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py).

### Local vs API Model Loading

**API Models** use cloud endpoints and require authentication via environment variables. The `generate_api_response*` functions handle image encoding and HTTP requests.

**Local VLMs** load into GPU memory using HuggingFace transformers. The `load_local_model()` function in [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py) contains conditional branches for each supported architecture, configuring the appropriate processor and model class.

## Supported VLMs and How to Select Them

The framework supports both commercial API services and open-source local models. You switch between them by changing the string passed to `--textualizer`.

### API-Based Models

These require valid API keys set in your environment:

- **Claude 3.5 Sonnet**: `claude-3-5-sonnet`
- **GPT-4o**: `gpt-4o`
- **GPT-4o Mini**: `gpt-4o-mini`

### Local VLMs

These download weights from HuggingFace and run on your hardware:

- **Qwen2-VL-7B** (default)
- **Qwen2-VL-72B**
- **Llava-v1.6-110b**
- **Llama-3.2-11B**
- **Llama-3.2-90B`

*Source: [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py), lines 44‑66*

## Practical Code Examples

### Command-Line Usage

Switch VLMs by modifying the `--textualizer` argument when running [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py):

```bash

# Use the default Qwen2-VL-7B

python -m src.textualizer --dataset flowvqa --output_type mermaid

# Switch to Llava-v1.6-110b (local VLM)

python -m src.textualizer \
    --dataset flowvqa \
    --textualizer Llava-v1.6-110b \
    --output_type mermaid

# Use GPT-4o API model

python -m src.textualizer \
    --dataset flowvqa \
    --textualizer gpt-4o \
    --output_type mermaid

```

### Programmatic Usage in Python

You can also switch VLMs directly in code by instantiating `ModelWrapper` with different model names:

```python
from src.models.model_loader import ModelWrapper
from src.prompts import load_textualizer_prompt
from src.utils import extract_representation

# Choose a VLM name – any listed in local_models.load_local_model or an API model

vlm_name = "Qwen2-VL-7B"          # or "Llava-v1.6-110b", "gpt-4o", etc.

# Initialise the wrapper

model = ModelWrapper(vlm_name)

# Prepare prompt and image

prompt = load_textualizer_prompt(output_type="mermaid")
image_path = "path/to/flowchart.png"

# Generate textual representation

response = model.generate_response(prompt, image_path=image_path)
representation = extract_representation(response)

print(representation)

```

### Extending to a New VLM

To add support for a new VLM not currently in the supported list:

1. Add the model name to the appropriate branch in `load_local_model()` in [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py) (or create a new API loader in [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py)).
2. Implement the loading logic (model, processor/tokenizer) following the existing patterns for similar architectures.
3. The `ModelWrapper` will automatically recognize the new name when passed via `--textualizer` without requiring changes to the core orchestration logic.

## Key Implementation Files

| File | Role |
|------|------|
| [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) | CLI entry point; parses `--textualizer` and drives the workflow. |
| [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py) | `ModelWrapper` class – decides API vs. local, loads model, forwards generation calls. |
| [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py) | Implements `load_local_model()` and VLM inference paths for each supported local model. |
| [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py) | Handles loading and calling cloud-based VLM APIs (Claude, GPT-4o). |
| [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py) | Supplies the textualizer prompt used for VLM generation. |
| [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py) | Helper `extract_representation` that parses the VLM's raw response into the requested format. |

## Summary

- Switch between VLMs by changing the `--textualizer` CLI argument or passing a different model name to the `ModelWrapper` constructor.
- The system automatically routes API models (`gpt-4o`, `claude-3-5-sonnet`) to cloud endpoints and local models (`Qwen2-VL`, `Llava`, `Llama-3.2`) to HuggingFace transformers.
- All model-specific logic is encapsulated in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py), making it trivial to add new VLMs without modifying the core textualization pipeline.

## Frequently Asked Questions

### What is the default VLM used by the Vision Textualizer?

The default VLM is **Qwen2-VL-7B**, specified in the argument parser within [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) at line 29. If you omit the `--textualizer` flag, the system automatically instantiates this model via the `ModelWrapper` class.

### Can I use commercial API models like GPT-4o instead of local VLMs?

Yes. The `ModelWrapper` in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py) recognizes `gpt-4o`, `gpt-4o-mini`, and `claude-3-5-sonnet` as API-based models. When you specify one of these identifiers, the system routes requests to [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py) instead of loading weights locally.

### How do I add support for a new VLM that is not currently listed?

To add a new VLM, modify the `load_local_model()` function in [`src/models/local_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/local_models.py) (for local models) or create a new loader in [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py) (for API services). Implement the model and processor initialization following the existing conditional patterns for similar architectures. Once added, `ModelWrapper` will automatically recognize the new model name without requiring changes to the CLI or core workflow.