# How TextFlow's Modular Architecture Enables Swapping Components

> Discover how TextFlow's modular architecture lets you swap components like vision-language models and encoders easily. Edit one file, no refactoring needed, to customize your pipeline.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: internals
- Published: 2026-03-05

---

**TextFlow decouples model inference, image preprocessing, and prompt generation into isolated modules, allowing you to substitute any component—vision-language models, encoders, or templates—by editing a single file without refactoring the core pipeline.**

The `junyiye/textflow` repository implements a **modular architecture** that splits the visual question answering workflow into distinct, interchangeable layers. By isolating responsibilities—model loading, image encoding, and prompt management—into dedicated modules under `src/`, the system lets developers swap individual parts without cascading changes throughout the codebase.

## Core Abstraction Layers

### Model Wrapper Interface

The `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py) serves as the central abstraction for all vision-language model interactions. It maintains a dictionary `self.is_api_model` that maps model names to boolean flags, distinguishing between API-based services (OpenAI, Anthropic) and local deployments (LLaVA, Qwen2-VL). When `generate_response` is called, the wrapper delegates to either `generate_api_response` or `generate_local_response` based on this mapping. Adding a new provider requires only extending the dictionary and implementing a thin adapter function.

### Image Encoding Utilities

Visual input handling is isolated in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py) through the `encode_image` function. This utility inspects the `model_name` parameter to route images through model-specific encoders: `encode_image_anthropic` for Claude models, `encode_image_openai` for GPT-4o variants, and a default PIL loader for local models. The function also applies conditional resizing when `model_name` contains "claude" and the image exceeds 2000 pixels, ensuring provider-specific constraints are enforced without cluttering the main pipeline.

### Prompt Template System

Natural language instructions are externalized into the `src/prompts/` package. The [`vqa_prompt.py`](https://github.com/junyiye/textflow/blob/main/vqa_prompt.py) module exposes `load_vqa_prompt(question)`, while [`reasoner_prompt.py`](https://github.com/junyiye/textflow/blob/main/reasoner_prompt.py) provides `load_reasoner_prompt(question, textualized_caption, textualized_choices)`. These functions return formatted strings, allowing the VQA and reasoning stages to request prompts without embedding template logic. To adopt a new prompting strategy—such as chain-of-thought or few-shot examples—developers simply modify the return value in the relevant prompt file.

## Swapping Components in Practice

### Switching Vision-Language Models

Changing the underlying VLM requires only two steps: updating [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) with the new model identifier and, if the model is not already supported, adding a loader function in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py). For example, to migrate from GPT-4o to Claude-3-5-sonnet:

```python

# In src/models/model_loader.py

self.is_api_model = {
    "gpt-4o": True,
    "gpt-4o-mini": True,
    "claude-3-5-sonnet": True,  # newly added

    "llava-v1.6-34b": False,
    "Qwen2-VL-7B-Instruct": False,
}

```

The `ModelWrapper` automatically routes calls to the Anthropic client because `is_api_model["claude-3-5-sonnet"]` evaluates to `True`.

### Customizing Image Preprocessing

When integrating a vision model that requires unique image preparation—such as specific normalization or tensor conversion—you extend `encode_image` in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py). Suppose you add a custom local model named "MyVision-7B" that expects base64 PNG strings:

```python

# In src/utils.py

def encode_image_myvision(image_path):
    with open(image_path, "rb") as img_file:
        return base64.b64encode(img_file.read()).decode("utf-8")

def encode_image(image_path, model_name=None):
    if model_name == "claude-3-5-sonnet":
        return encode_image_anthropic(image_path)
    elif model_name in ["gpt-4o", "gpt-4o-mini"]:
        return encode_image_openai(image_path)
    elif model_name == "my-vision-7b":
        return encode_image_myvision(image_path)
    else:
        return Image.open(image_path)

```

The VQA pipeline in [`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py) continues to call `encode_image(image_path, model_name)` without modification.

### Injecting New Prompt Templates

To experiment with a chain-of-thought prompting style for the reasoning stage, you edit [`src/prompts/reasoner_prompt.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/reasoner_prompt.py):

```python

# In src/prompts/reasoner_prompt.py

def load_reasoner_prompt(question, textualized_caption, textualized_choices):
    return (
        "You are an expert visual reasoning assistant.\n"
        "Caption: {caption}\n"
        "Choices: {choices}\n"
        "Question: {question}\n"
        "Let's think step by step before answering.\nAnswer:"
    ).format(
        caption=textualized_caption,
        choices=textualized_choices,
        question=question
    )

```

Because [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) imports `load_reasoner_prompt` from the prompts package, the new template is automatically picked up on the next run.

## Summary

- **Model abstraction** in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py) isolates API and local inference behind a unified `ModelWrapper`, making model swaps a matter of updating a dictionary.
- **Image encoding** in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py) routes preprocessing through model-specific branches, allowing new vision encoders to be plugged in without touching the VQA loop.
- **Prompt templates** in `src/prompts/` externalize language instructions, enabling rapid experimentation with different prompting strategies.
- **Pipeline scripts** ([`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py), [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py)) remain agnostic to these details, ensuring that component changes do not cascade into breaking changes elsewhere.

## Frequently Asked Questions

### Can I use a local vision model instead of an API service?

Yes. Add the model name to `self.is_api_model` with a value of `False` in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py), then implement a `generate_local_response` function that loads your model and returns text. The `ModelWrapper` will automatically route inference to your local implementation.

### How do I add support for a new image format or preprocessing step?

Extend [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py) by creating a new encoder function (e.g., `encode_image_custom`) and add a conditional branch in `encode_image` that checks for your model name. The VQA and reasoning scripts call `encode_image` with the current model identifier, so your new logic will be invoked automatically.

### Is it possible to switch prompt templates at runtime?

The prompt loaders are simple Python functions imported at the top of [`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py) and [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py). While the current implementation loads templates statically, you can modify the loader functions to accept an additional `template_name` parameter and return different strings based on runtime configuration, enabling dynamic prompt selection without altering the core pipeline logic.

### Will swapping components break the evaluation metrics?

No. The evaluation module in [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py) operates on the final JSON output produced by [`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py) or [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py). As long as the new component adheres to the existing return format (a dictionary with keys such as `answer` or `reasoning`), the accuracy and consistency calculations remain valid.