Building Vision Transformers (ViT) from Scratch Without Pre-Built Libraries

Vision Transformers (ViTs) replace convolutional front-ends with a patch-token pipeline that converts images into sequences of vectors using only patch extraction, linear projection, and positional encoding, allowing standard transformer blocks to process visual data.

The rohitg00/ai-engineering-from-scratch repository demonstrates how to build Vision Transformers without relying on high-level libraries like PyTorch or TensorFlow. This approach implements the complete patch-to-sequence pipeline using only NumPy and Python's standard library, revealing the fundamental operations that frameworks typically abstract away.

The Patch-to-Token Pipeline

Vision Transformers process images by first converting spatial pixel data into a sequence of token vectors. According to phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md (lines 31-36), this transformation occurs through two primary operations: patch extraction and linear projection.

Non-Overlapping Patch Extraction

The first step splits an input image of height H and width W into a grid of non-overlapping P×P pixel patches. This yields a total of (H/P)×(W/P) patches, each containing 3P² values (for RGB channels). As implemented in phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py, the extract_patches function reshapes the image tensor from (H, W, 3) to (grid_h, patch_h, grid_w, patch_w, 3) before flattening each patch into a vector.

Linear Projection as Convolution

Each flattened patch vector undergoes linear projection to a D-dimensional embedding space via a shared weight matrix W_E ∈ ℝ^(3P²×D). In practice, this operation is equivalent to a 2D convolution with kernel_size=P, stride=P, and out_channels=D (docs line 38). The linear_proj function in the reference implementation performs this matrix multiplication, transforming the (N, 3P²) patch tensor into (N, D) token embeddings, where N represents the number of patches.

Position Encoding and Global Representations

Once patches become tokens, the architecture must inject spatial information and aggregate global context.

2D Rotary Position Embeddings (2D-RoPE)

Early ViT architectures used learnable 1D positional vectors added directly to token embeddings. Modern implementations prefer 2D Rotary Position Embeddings (2D-RoPE), which rotate query and key vectors based on each patch's row and column coordinates (docs lines 42-45). This approach eliminates the need for fixed position tables and generalizes better to variable image resolutions, as demonstrated in the SigLIP-2 SO400m/14 architecture.

[CLS] Tokens and Register Tokens

To obtain a single vector representing the entire image, ViTs prepend a learnable [CLS] token to the patch sequence. Alternatively, architectures may use register tokens or mean-pooling strategies. The choice impacts downstream tasks: classification favors the [CLS] token output, while vision-language models (VLMs) often feed all patch tokens directly to the language model and discard registers (docs lines 46-53).

Implementing the Transformer Stack

The token sequence—including any [CLS] or register tokens—flows through standard transformer blocks comprising multi-head self-attention and feed-forward networks.

Token Sequence Construction

The build_tokens function in phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py orchestrates the pipeline:

  1. Extract patches using extract_patches
  2. Project patches via linear_proj with learned weights W_E
  3. Prepend a randomly initialized [CLS] token
  4. Apply rotary positional encoding via apply_rotary

This produces a sequence of shape (N+1, D), where the first row represents the global image token.

Multi-Head Self-Attention and FFN

The transformer_block function implements the ViT-B/16 architecture specified in phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md (lines 78-88):

  • 12 attention heads operating in parallel
  • LayerNorm applied before attention (pre-normalization)
  • GELU activation in the feed-forward network
  • 4× expansion ratio in the FFN hidden layer (D → 4D → D)

Each block computes scaled dot-product attention, applies residual connections, and processes tokens through the expanded feed-forward network.

The Complete Encoder

The vit_encoder function chains 12 identical transformer_block layers to match the ViT-B/16 depth. After processing, the function returns the first token (the [CLS] embedding) as the final image representation of shape (D,).

From Scratch to Production: Scaling ViTs

By 2026, the dominant vision architecture is the SigLIP-2 SO400m/14 model, which adds several enhancements to the basic ViT structure:

  • Native 2D-RoPE for coordinate-aware attention
  • NaFlex (native flexible resolution) support for variable input sizes
  • Register tokens for improved global aggregation
  • Optimized geometry for 400 million parameters and 14×14 patch sizing

For arbitrary configurations, the repository includes a geometry calculator in phases/12-multimodal-ai/01-vision-transformer-patch-tokens/outputs/skill-patch-geometry-reader.md. This utility estimates token counts, parameter counts, and FLOPs based on image resolution, patch size, and model depth.

Complete NumPy Implementation

Below is the minimal, self-contained implementation from phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py that constructs a ViT-B/16 encoder without external deep learning dependencies:

import numpy as np

def extract_patches(img, patch_size):
    """Extract non-overlapping patches from image.
    img: (H, W, 3) uint8 array"""
    H, W, C = img.shape
    assert H % patch_size == 0 and W % patch_size == 0
    img = img.reshape(H // patch_size, patch_size,
                      W // patch_size, patch_size, C)
    patches = img.transpose(0, 2, 1, 3, 4).reshape(-1, patch_size * patch_size * C)
    return patches

def linear_proj(patches, proj_weight, proj_bias):
    """Project patches to D dimensions.
    patches: (N, 3*P^2), proj_weight: (3*P^2, D)"""
    return patches @ proj_weight + proj_bias

def apply_rotary(x, seq_len, dim):
    """Apply 2D-RoPE positional encoding."""
    x1 = x[..., ::2]
    x2 = x[..., 1::2]
    theta = np.arange(seq_len)[:, None] / (10000 ** (np.arange(dim // 2) / (dim // 2)))
    sin, cos = np.sin(theta), np.cos(theta)
    return np.concatenate([x1 * cos - x2 * sin,
                           x1 * sin + x2 * cos], axis=-1)

def build_tokens(img, patch_size=16, hidden_dim=768):
    patches = extract_patches(img, patch_size)
    D = hidden_dim
    W_E = np.random.randn(3 * patch_size * patch_size, D) * 0.02
    b_E = np.zeros(D)
    patch_tokens = linear_proj(patches, W_E, b_E)
    cls_token = np.random.randn(1, D) * 0.02
    tokens = np.concatenate([cls_token, patch_tokens], axis=0)
    tokens = apply_rotary(tokens, tokens.shape[0], D)
    return tokens

def transformer_block(x, n_head=12):
    B, D = x.shape
    Q = x @ np.random.randn(D, D)
    K = x @ np.random.randn(D, D)
    V = x @ np.random.randn(D, D)
    
    attn_scores = Q @ K.T / np.sqrt(D)
    attn_weights = np.exp(attn_scores) / np.exp(attn_scores).sum(axis=-1, keepdims=True)
    attn_out = attn_weights @ V
    
    x = x + attn_out
    x = x / np.linalg.norm(x, axis=-1, keepdims=True)
    
    w1, b1 = np.random.randn(D, 4 * D), np.zeros(4 * D)
    w2, b2 = np.random.randn(4 * D, D), np.zeros(D)
    ff = np.maximum(0, x @ w1 + b1)  # GELU approximated

    ff = ff @ w2 + b2
    x = x + ff
    x = x / np.linalg.norm(x, axis=-1, keepdims=True)
    return x

def vit_encoder(img):
    tokens = build_tokens(img, patch_size=16, hidden_dim=768)
    for _ in range(12):  # ViT-B/16 depth

        tokens = transformer_block(tokens)
    return tokens[0]  # Return [CLS] embedding

if __name__ == "__main__":
    dummy_img = np.random.randint(0, 255, (224, 224, 3), dtype=np.uint8)
    embedding = vit_encoder(dummy_img)
    print("Image embedding shape:", embedding.shape)

Summary

  • Patch extraction converts spatial images into (N, 3P²) tensors by splitting inputs into non-overlapping grids.
  • Linear projection maps patches to embedding space via matrix multiplication equivalent to a strided convolution.
  • 2D-RoPE provides coordinate-aware positional encoding without fixed position tables.
  • [CLS] tokens aggregate global image context for classification tasks.
  • Transformer blocks follow the ViT-B/16 specification: 12 layers, 12 heads, GELU activation, and 4× FFN expansion.
  • The reference implementation in phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py demonstrates the complete architecture using only NumPy.

Frequently Asked Questions

Why implement a ViT without deep learning frameworks?

Implementing Vision Transformers from scratch using only NumPy and the standard library reveals the underlying mechanics of patch tokenization, positional encoding, and self-attention that high-level frameworks abstract away. This approach, as demonstrated in rohitg00/ai-engineering-from-scratch, provides foundational understanding of how image data transforms into sequence representations before introducing framework-specific optimizations.

How does patch extraction convert images to sequences?

Patch extraction splits an input image of size H×W into non-overlapping P×P pixel grids, yielding (H/P)×(W/P) patches. Each patch flattens into a vector of length 3P² (for RGB), creating a sequence of N vectors. According to phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md (lines 31-36), this operation uses reshape and transpose operations to maintain spatial coherence while converting 2D spatial data into a 1D sequence format suitable for transformer processing.

What are the advantages of 2D-RoPE over 1D positional embeddings?

2D Rotary Position Embeddings (2D-RoPE) rotate query and key vectors based on each patch's row and column coordinates, while 1D embeddings use learnable vectors added to token representations. 2D-RoPE eliminates the need for fixed position tables, reduces parameter count, and generalizes better to variable image resolutions. Modern architectures like SigLIP-2 SO400m/14 adopt 2D-RoPE for native flexible resolution (NaFlex) support, as documented in phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md (lines 42-45).

How do you calculate parameters and FLOPs for a custom ViT?

The geometry calculator in phases/12-multimodal-ai/01-vision-transformer-patch-tokens/outputs/skill-patch-geometry-reader.md computes token counts, parameter counts, and FLOPs based on image resolution, patch size P, hidden dimension D, and transformer depth. For a ViT-B/16 configuration, the calculator accounts for 12 blocks with 12 heads each, 4× feed-forward expansion, and embedding matrices to estimate computational requirements for arbitrary input configurations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →