How TextFlow's Modular Architecture Enables Swapping Components

TextFlow decouples model inference, image preprocessing, and prompt generation into isolated modules, allowing you to substitute any component—vision-language models, encoders, or templates—by editing a single file without refactoring the core pipeline.

The junyiye/textflow repository implements a modular architecture that splits the visual question answering workflow into distinct, interchangeable layers. By isolating responsibilities—model loading, image encoding, and prompt management—into dedicated modules under src/, the system lets developers swap individual parts without cascading changes throughout the codebase.

Core Abstraction Layers

Model Wrapper Interface

The ModelWrapper class in src/models/model_loader.py serves as the central abstraction for all vision-language model interactions. It maintains a dictionary self.is_api_model that maps model names to boolean flags, distinguishing between API-based services (OpenAI, Anthropic) and local deployments (LLaVA, Qwen2-VL). When generate_response is called, the wrapper delegates to either generate_api_response or generate_local_response based on this mapping. Adding a new provider requires only extending the dictionary and implementing a thin adapter function.

Image Encoding Utilities

Visual input handling is isolated in src/utils.py through the encode_image function. This utility inspects the model_name parameter to route images through model-specific encoders: encode_image_anthropic for Claude models, encode_image_openai for GPT-4o variants, and a default PIL loader for local models. The function also applies conditional resizing when model_name contains "claude" and the image exceeds 2000 pixels, ensuring provider-specific constraints are enforced without cluttering the main pipeline.

Prompt Template System

Natural language instructions are externalized into the src/prompts/ package. The vqa_prompt.py module exposes load_vqa_prompt(question), while reasoner_prompt.py provides load_reasoner_prompt(question, textualized_caption, textualized_choices). These functions return formatted strings, allowing the VQA and reasoning stages to request prompts without embedding template logic. To adopt a new prompting strategy—such as chain-of-thought or few-shot examples—developers simply modify the return value in the relevant prompt file.

Swapping Components in Practice

Switching Vision-Language Models

Changing the underlying VLM requires only two steps: updating config.json with the new model identifier and, if the model is not already supported, adding a loader function in src/models/model_loader.py. For example, to migrate from GPT-4o to Claude-3-5-sonnet:


# In src/models/model_loader.py

self.is_api_model = {
    "gpt-4o": True,
    "gpt-4o-mini": True,
    "claude-3-5-sonnet": True,  # newly added

    "llava-v1.6-34b": False,
    "Qwen2-VL-7B-Instruct": False,
}

The ModelWrapper automatically routes calls to the Anthropic client because is_api_model["claude-3-5-sonnet"] evaluates to True.

Customizing Image Preprocessing

When integrating a vision model that requires unique image preparation—such as specific normalization or tensor conversion—you extend encode_image in src/utils.py. Suppose you add a custom local model named "MyVision-7B" that expects base64 PNG strings:


# In src/utils.py

def encode_image_myvision(image_path):
    with open(image_path, "rb") as img_file:
        return base64.b64encode(img_file.read()).decode("utf-8")

def encode_image(image_path, model_name=None):
    if model_name == "claude-3-5-sonnet":
        return encode_image_anthropic(image_path)
    elif model_name in ["gpt-4o", "gpt-4o-mini"]:
        return encode_image_openai(image_path)
    elif model_name == "my-vision-7b":
        return encode_image_myvision(image_path)
    else:
        return Image.open(image_path)

The VQA pipeline in src/vqa.py continues to call encode_image(image_path, model_name) without modification.

Injecting New Prompt Templates

To experiment with a chain-of-thought prompting style for the reasoning stage, you edit src/prompts/reasoner_prompt.py:


# In src/prompts/reasoner_prompt.py

def load_reasoner_prompt(question, textualized_caption, textualized_choices):
    return (
        "You are an expert visual reasoning assistant.\n"
        "Caption: {caption}\n"
        "Choices: {choices}\n"
        "Question: {question}\n"
        "Let's think step by step before answering.\nAnswer:"
    ).format(
        caption=textualized_caption,
        choices=textualized_choices,
        question=question
    )

Because src/reasoner.py imports load_reasoner_prompt from the prompts package, the new template is automatically picked up on the next run.

Summary

  • Model abstraction in src/models/model_loader.py isolates API and local inference behind a unified ModelWrapper, making model swaps a matter of updating a dictionary.
  • Image encoding in src/utils.py routes preprocessing through model-specific branches, allowing new vision encoders to be plugged in without touching the VQA loop.
  • Prompt templates in src/prompts/ externalize language instructions, enabling rapid experimentation with different prompting strategies.
  • Pipeline scripts (src/vqa.py, src/reasoner.py) remain agnostic to these details, ensuring that component changes do not cascade into breaking changes elsewhere.

Frequently Asked Questions

Can I use a local vision model instead of an API service?

Yes. Add the model name to self.is_api_model with a value of False in src/models/model_loader.py, then implement a generate_local_response function that loads your model and returns text. The ModelWrapper will automatically route inference to your local implementation.

How do I add support for a new image format or preprocessing step?

Extend src/utils.py by creating a new encoder function (e.g., encode_image_custom) and add a conditional branch in encode_image that checks for your model name. The VQA and reasoning scripts call encode_image with the current model identifier, so your new logic will be invoked automatically.

Is it possible to switch prompt templates at runtime?

The prompt loaders are simple Python functions imported at the top of src/vqa.py and src/reasoner.py. While the current implementation loads templates statically, you can modify the loader functions to accept an additional template_name parameter and return different strings based on runtime configuration, enabling dynamic prompt selection without altering the core pipeline logic.

Will swapping components break the evaluation metrics?

No. The evaluation module in src/evaluation.py operates on the final JSON output produced by src/vqa.py or src/reasoner.py. As long as the new component adheres to the existing return format (a dictionary with keys such as answer or reasoning), the accuracy and consistency calculations remain valid.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →