How to Build Multimodal Vision-Language Models: A Complete Technical Guide

To build multimodal vision-language models, combine a vision encoder that converts images into patch tokens with a language model via a projection layer and cross-attention fusion, then train the system with contrastive losses on image-text pairs.

Multimodal vision-language models (VLMs) integrate computer vision and natural language processing to enable AI systems that can interpret visual content while generating or processing text. According to the rohitg00/ai-engineering-from-scratch repository, these architectures rely on four fundamental layers that bridge visual and textual modalities through careful engineering of token representations and attention mechanisms. This guide breaks down each component with reference implementations and production-ready design patterns.

The Four-Layer Architecture of Vision-Language Models

Modern VLMs consist of four critical layers that transform raw pixels into semantic understanding aligned with language.

The Patch-Token Primitive

Every vision-language model begins by converting images into sequences of tokens. An image is split into fixed-size patches (typically 14×14 or 16×16 pixels), each flattened and projected into a vector representation. This approach, detailed in phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md, creates a flexible sequence whose length equals (H/P)×(W/P), where H and W are image dimensions and P is patch size.

The Projection Layer

Vision tokens exist in a different embedding space than text tokens. A small two-layer MLP called the projector aligns these modalities by mapping visual vectors into the language model's embedding dimension. As implemented in phases/19-capstone-projects/60-projection-layer-modality-align/code/main.py, this layer is trained with a cosine-alignment loss against paired captions to ensure semantic compatibility.

Cross-Attention Fusion

To ground language generation in visual content, models insert cross-attention blocks every K transformer layers (typically every 4-6 blocks). These blocks allow the language model's queries to attend to vision keys and values, enabling the model to "look at" specific image regions while generating text. The implementation in phases/19-capstone-projects/61-cross-attention-fusion/code/main.py demonstrates this fusion mechanism.

Joint Pre-training

The complete system undergoes contrastive pre-training on billions of image-text pairs using InfoNCE or sigmoid-pairwise losses (SigLIP). This phase, covered in phases/12-multimodal-ai/02-clip-contrastive-pretraining/docs/en.md, teaches the model a shared multimodal representation space where matching images and text are pulled together while non-matching pairs are pushed apart.

Key Design Choices for Production VLMs

When building multimodal vision-language models, several architectural decisions significantly impact performance:

Design Choice Impact Typical Defaults (2026)
Patch Size Controls token count vs. fidelity. Smaller patches improve OCR but increase compute. 14px (SigLIP 2) or 16px (ViT-B/16)
Positional Encoding 1D learnable embeddings fix resolution; 2D-RoPE supports arbitrary aspect ratios. 2D-RoPE (native-resolution)
Pooling Strategy CLS token for classification; mean-pooling for dense features; no pooling for VLMs. No pooling (full patch sequence)
Pre-training Objective Contrastive (InfoNCE) vs. sigmoid pairwise (SigLIP). Sigmoid scales to larger batches. SigLIP 2 (Sigmoid + NaFlex)
Fusion Frequency Cross-attention every K blocks. More frequent fusion improves grounding but adds compute. Every 4-6 blocks (Flamingo/IDEFICS)

Implementation: Building Each Component

Follow these steps to implement a complete VLM pipeline using the reference code from ai-engineering-from-scratch.

Step 1: Create the Patch Tokenizer

Start by implementing the patch tokenizer using a convolutional projection. In phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py, the geometry calculator determines token sequence length:

def patch_tokens(image_h: int, image_w: int, patch_p: int, hidden_dim: int):
    """Return the number of tokens and the projection weight shape."""
    grid_h, grid_w = image_h // patch_p, image_w // patch_p
    seq_len = grid_h * grid_w          # no CLS or registers

    # Linear projection from (3*P^2) → hidden_dim

    proj_shape = (3 * patch_p * patch_p, hidden_dim)
    return seq_len, proj_shape

This function returns the sequence length and projection shape needed to convert image patches into the hidden dimension expected by your vision transformer.

Step 2: Implement the Projection MLP

The projection layer aligns vision and language embedding spaces. From phases/19-capstone-projects/60-projection-layer-modality-align/code/main.py:

def projection_mlp(vision_vec, hidden_dim):
    """Two-layer MLP that maps vision vectors to LLM space."""
    import math
    w1 = [[0.01] * hidden_dim for _ in range(len(vision_vec))]
    w2 = [[0.01] * hidden_dim for _ in range(hidden_dim)]
    # Linear → ReLU → Linear

    hidden = [math.tanh(sum(v * w for v, w in zip(vision_vec, row))) for row in w1]
    out = [sum(h * w for h, w in zip(hidden, col)) for col in zip(*w2)]
    return out   # same dimensionality as LLM embeddings

This two-layer network with ReLU activation transforms visual features into the language model's input space, enabling direct concatenation with text tokens.

Step 3: Add Cross-Attention Layers

Implement the fusion mechanism that allows text queries to attend to vision keys and values. From phases/19-capstone-projects/61-cross-attention-fusion/code/main.py:

def cross_attention(text_tokens, vision_tokens):
    """Simple scaled dot-product cross-attention (no masking)."""
    import math
    # Queries from text, keys/values from vision

    q = [t * 0.1 for t in text_tokens]
    k = [v * 0.1 for v in vision_tokens]
    v = vision_tokens
    # Scaled dot-product

    dk = math.sqrt(len(k[0]))
    scores = [[sum(q_i * k_j) / dk for k_j in k] for q_i in q]
    # Softmax across vision tokens

    attn = [[math.exp(s) / sum(math.exp(x) for x in row) for s in row] for row in scores]
    # Weighted sum

    new_text = [sum(a * v_j for a, v_j in zip(row, v)) for row in attn]
    return new_text

Insert this block every K layers in your language model decoder to maintain visual grounding throughout text generation.

Step 4: Train with Contrastive Loss

Use symmetric InfoNCE loss to align image and text representations. From phases/12-multimodal-ai/02-clip-contrastive-pretraining/code/main.py:

def info_nce_loss(image_embs, text_embs, temperature=0.07):
    """Symmetric InfoNCE loss for a batch of image-text pairs."""
    import math
    # Normalize

    i_norm = [vec / math.sqrt(sum(x*x for x in vec)) for vec in image_embs]
    t_norm = [vec / math.sqrt(sum(x*x for x in vec)) for vec in text_embs]
    # Similarity matrix

    sims = [[sum(i * t for i, t in zip(i_vec, t_vec)) / temperature
             for t_vec in t_norm] for i_vec in i_norm]
    # Softmax cross-entropy on rows and columns

    def cross_entropy(row):
        max_logit = max(row)
        exp_row = [math.exp(r - max_logit) for r in row]
        sum_exp = sum(exp_row)
        return -math.log(exp_row[row.index(max(row))] / sum_exp)
    loss_i2t = sum(cross_entropy(r) for r in sims) / len(sims)
    loss_t2i = sum(cross_entropy([r[i] for r in sims]) for i in range(len(sims))) / len(sims)
    return (loss_i2t + loss_t2i) / 2

Train on large batches of image-text pairs to maximize the similarity between matching pairs while minimizing non-matching pairs.

End-to-End Inference Pipeline

Combine all components for caption generation or visual question answering. From phases/19-capstone-projects/62-vision-language-pretraining/code/main.py:

def generate_caption(image, tokenizer, llm):
    """Run vision encoder → projector → cross-attention → LLM decoder."""
    # 1. Tokenize image patches

    seq_len, _ = patch_tokens(*image.shape[:2], patch_p=14, hidden_dim=768)
    vision_tokens = [0.0] * seq_len   # placeholder embeddings

    # 2. Project to LLM space

    proj_tokens = projection_mlp(vision_tokens, hidden_dim=1024)
    # 3. Fuse with text prompt

    fused = cross_attention([], proj_tokens)
    # 4. Decode with LLM

    caption = llm.decode(fused)
    return caption

This pipeline demonstrates the complete data flow: image patches are tokenized, projected into language space, fused via cross-attention, and decoded into natural language.

Summary

  • Patch-token primitives convert images into sequences suitable for transformer processing by splitting them into fixed-size patches and projecting them into embedding vectors.
  • Projection layers align visual and textual embedding spaces through learned two-layer MLPs trained with cosine-alignment losses.
  • Cross-attention fusion grounds language generation in specific image regions by allowing text queries to attend to vision keys and values every K transformer blocks.
  • Contrastive pre-training creates a shared multimodal representation space using InfoNCE or SigLIP losses on billions of image-text pairs.

Frequently Asked Questions

What is the optimal patch size for vision-language models?

Smaller patches (14px) provide higher fidelity for OCR and dense prediction tasks but increase computational cost and sequence length, while 16px patches offer a balanced trade-off for general purpose VLMs. According to the ai-engineering-from-scratch implementation, 14px is now standard for high-resolution models like SigLIP 2, though 16px remains common for ViT-B/16 architectures.

How does the projection layer prevent modality misalignment?

The projection layer acts as a learnable adapter that maps visual tokens from the vision encoder's output space into the language model's input embedding space. By training this two-layer MLP with a cosine-alignment loss against paired caption embeddings, the model learns to place semantically similar images and text in nearby regions of the embedding space, enabling seamless concatenation during inference.

Why is cross-attention inserted every K layers rather than every layer?

Inserting cross-attention every 4-6 transformer blocks (as seen in Flamingo and IDEFICS architectures) provides a balance between computational efficiency and visual grounding quality. Less frequent fusion reduces memory usage and training time, while still allowing the language model to access visual context at multiple depths of the network. Full fusion (every layer) improves grounding but typically increases training costs by 3-4x.

Can I use a pre-trained vision encoder instead of training from scratch?

Yes, most production VLMs utilize pre-trained vision encoders like CLIP, SigLIP, or DINOv2 and freeze their weights during initial training, only updating the projection layers and cross-attention parameters. This approach leverages the rich visual representations learned from billions of images and significantly reduces the compute requirements for building multimodal vision-language models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →