# Building Vision Transformers (ViT) from Scratch Without Pre-Built Libraries

> Learn to build Vision Transformers ViT from scratch. Understand patch extraction, linear projection, and positional encoding for visual data processing. No pre-built libraries needed.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-26

---

**Vision Transformers (ViTs) replace convolutional front-ends with a patch-token pipeline that converts images into sequences of vectors using only patch extraction, linear projection, and positional encoding, allowing standard transformer blocks to process visual data.**

The `rohitg00/ai-engineering-from-scratch` repository demonstrates how to build Vision Transformers without relying on high-level libraries like PyTorch or TensorFlow. This approach implements the complete patch-to-sequence pipeline using only NumPy and Python's standard library, revealing the fundamental operations that frameworks typically abstract away.

## The Patch-to-Token Pipeline

Vision Transformers process images by first converting spatial pixel data into a sequence of token vectors. According to [`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md) (lines 31-36), this transformation occurs through two primary operations: patch extraction and linear projection.

### Non-Overlapping Patch Extraction

The first step splits an input image of height **H** and width **W** into a grid of non-overlapping **P×P** pixel patches. This yields a total of **(H/P)×(W/P)** patches, each containing **3P²** values (for RGB channels). As implemented in [`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py), the `extract_patches` function reshapes the image tensor from `(H, W, 3)` to `(grid_h, patch_h, grid_w, patch_w, 3)` before flattening each patch into a vector.

### Linear Projection as Convolution

Each flattened patch vector undergoes linear projection to a **D**-dimensional embedding space via a shared weight matrix **W_E ∈ ℝ^(3P²×D)**. In practice, this operation is equivalent to a 2D convolution with `kernel_size=P`, `stride=P`, and `out_channels=D` (docs line 38). The `linear_proj` function in the reference implementation performs this matrix multiplication, transforming the `(N, 3P²)` patch tensor into `(N, D)` token embeddings, where **N** represents the number of patches.

## Position Encoding and Global Representations

Once patches become tokens, the architecture must inject spatial information and aggregate global context.

### 2D Rotary Position Embeddings (2D-RoPE)

Early ViT architectures used learnable 1D positional vectors added directly to token embeddings. Modern implementations prefer **2D Rotary Position Embeddings (2D-RoPE)**, which rotate query and key vectors based on each patch's row and column coordinates (docs lines 42-45). This approach eliminates the need for fixed position tables and generalizes better to variable image resolutions, as demonstrated in the SigLIP-2 SO400m/14 architecture.

### [CLS] Tokens and Register Tokens

To obtain a single vector representing the entire image, ViTs prepend a learnable `[CLS]` token to the patch sequence. Alternatively, architectures may use register tokens or mean-pooling strategies. The choice impacts downstream tasks: classification favors the `[CLS]` token output, while vision-language models (VLMs) often feed all patch tokens directly to the language model and discard registers (docs lines 46-53).

## Implementing the Transformer Stack

The token sequence—including any `[CLS]` or register tokens—flows through standard transformer blocks comprising multi-head self-attention and feed-forward networks.

### Token Sequence Construction

The `build_tokens` function in [`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py) orchestrates the pipeline:

1. Extract patches using `extract_patches`
2. Project patches via `linear_proj` with learned weights **W_E**
3. Prepend a randomly initialized `[CLS]` token
4. Apply rotary positional encoding via `apply_rotary`

This produces a sequence of shape `(N+1, D)`, where the first row represents the global image token.

### Multi-Head Self-Attention and FFN

The `transformer_block` function implements the ViT-B/16 architecture specified in [`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md) (lines 78-88):

- **12 attention heads** operating in parallel
- **LayerNorm** applied before attention (pre-normalization)
- **GELU activation** in the feed-forward network
- **4× expansion** ratio in the FFN hidden layer (D → 4D → D)

Each block computes scaled dot-product attention, applies residual connections, and processes tokens through the expanded feed-forward network.

### The Complete Encoder

The `vit_encoder` function chains 12 identical `transformer_block` layers to match the ViT-B/16 depth. After processing, the function returns the first token (the `[CLS]` embedding) as the final image representation of shape `(D,)`.

## From Scratch to Production: Scaling ViTs

By 2026, the dominant vision architecture is the **SigLIP-2 SO400m/14** model, which adds several enhancements to the basic ViT structure:

- Native **2D-RoPE** for coordinate-aware attention
- **NaFlex** (native flexible resolution) support for variable input sizes
- **Register tokens** for improved global aggregation
- Optimized geometry for 400 million parameters and 14×14 patch sizing

For arbitrary configurations, the repository includes a geometry calculator in [`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/outputs/skill-patch-geometry-reader.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/outputs/skill-patch-geometry-reader.md). This utility estimates token counts, parameter counts, and FLOPs based on image resolution, patch size, and model depth.

## Complete NumPy Implementation

Below is the minimal, self-contained implementation from [`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py) that constructs a ViT-B/16 encoder without external deep learning dependencies:

```python
import numpy as np

def extract_patches(img, patch_size):
    """Extract non-overlapping patches from image.
    img: (H, W, 3) uint8 array"""
    H, W, C = img.shape
    assert H % patch_size == 0 and W % patch_size == 0
    img = img.reshape(H // patch_size, patch_size,
                      W // patch_size, patch_size, C)
    patches = img.transpose(0, 2, 1, 3, 4).reshape(-1, patch_size * patch_size * C)
    return patches

def linear_proj(patches, proj_weight, proj_bias):
    """Project patches to D dimensions.
    patches: (N, 3*P^2), proj_weight: (3*P^2, D)"""
    return patches @ proj_weight + proj_bias

def apply_rotary(x, seq_len, dim):
    """Apply 2D-RoPE positional encoding."""
    x1 = x[..., ::2]
    x2 = x[..., 1::2]
    theta = np.arange(seq_len)[:, None] / (10000 ** (np.arange(dim // 2) / (dim // 2)))
    sin, cos = np.sin(theta), np.cos(theta)
    return np.concatenate([x1 * cos - x2 * sin,
                           x1 * sin + x2 * cos], axis=-1)

def build_tokens(img, patch_size=16, hidden_dim=768):
    patches = extract_patches(img, patch_size)
    D = hidden_dim
    W_E = np.random.randn(3 * patch_size * patch_size, D) * 0.02
    b_E = np.zeros(D)
    patch_tokens = linear_proj(patches, W_E, b_E)
    cls_token = np.random.randn(1, D) * 0.02
    tokens = np.concatenate([cls_token, patch_tokens], axis=0)
    tokens = apply_rotary(tokens, tokens.shape[0], D)
    return tokens

def transformer_block(x, n_head=12):
    B, D = x.shape
    Q = x @ np.random.randn(D, D)
    K = x @ np.random.randn(D, D)
    V = x @ np.random.randn(D, D)
    
    attn_scores = Q @ K.T / np.sqrt(D)
    attn_weights = np.exp(attn_scores) / np.exp(attn_scores).sum(axis=-1, keepdims=True)
    attn_out = attn_weights @ V
    
    x = x + attn_out
    x = x / np.linalg.norm(x, axis=-1, keepdims=True)
    
    w1, b1 = np.random.randn(D, 4 * D), np.zeros(4 * D)
    w2, b2 = np.random.randn(4 * D, D), np.zeros(D)
    ff = np.maximum(0, x @ w1 + b1)  # GELU approximated

    ff = ff @ w2 + b2
    x = x + ff
    x = x / np.linalg.norm(x, axis=-1, keepdims=True)
    return x

def vit_encoder(img):
    tokens = build_tokens(img, patch_size=16, hidden_dim=768)
    for _ in range(12):  # ViT-B/16 depth

        tokens = transformer_block(tokens)
    return tokens[0]  # Return [CLS] embedding

if __name__ == "__main__":
    dummy_img = np.random.randint(0, 255, (224, 224, 3), dtype=np.uint8)
    embedding = vit_encoder(dummy_img)
    print("Image embedding shape:", embedding.shape)

```

## Summary

- **Patch extraction** converts spatial images into `(N, 3P²)` tensors by splitting inputs into non-overlapping grids.
- **Linear projection** maps patches to embedding space via matrix multiplication equivalent to a strided convolution.
- **2D-RoPE** provides coordinate-aware positional encoding without fixed position tables.
- **[CLS] tokens** aggregate global image context for classification tasks.
- **Transformer blocks** follow the ViT-B/16 specification: 12 layers, 12 heads, GELU activation, and 4× FFN expansion.
- The reference implementation in [`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py) demonstrates the complete architecture using only NumPy.

## Frequently Asked Questions

### Why implement a ViT without deep learning frameworks?

Implementing Vision Transformers from scratch using only NumPy and the standard library reveals the underlying mechanics of patch tokenization, positional encoding, and self-attention that high-level frameworks abstract away. This approach, as demonstrated in `rohitg00/ai-engineering-from-scratch`, provides foundational understanding of how image data transforms into sequence representations before introducing framework-specific optimizations.

### How does patch extraction convert images to sequences?

Patch extraction splits an input image of size **H×W** into non-overlapping **P×P** pixel grids, yielding **(H/P)×(W/P)** patches. Each patch flattens into a vector of length **3P²** (for RGB), creating a sequence of **N** vectors. According to [`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md) (lines 31-36), this operation uses reshape and transpose operations to maintain spatial coherence while converting 2D spatial data into a 1D sequence format suitable for transformer processing.

### What are the advantages of 2D-RoPE over 1D positional embeddings?

**2D Rotary Position Embeddings (2D-RoPE)** rotate query and key vectors based on each patch's row and column coordinates, while 1D embeddings use learnable vectors added to token representations. 2D-RoPE eliminates the need for fixed position tables, reduces parameter count, and generalizes better to variable image resolutions. Modern architectures like SigLIP-2 SO400m/14 adopt 2D-RoPE for native flexible resolution (NaFlex) support, as documented in [`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/docs/en.md) (lines 42-45).

### How do you calculate parameters and FLOPs for a custom ViT?

The geometry calculator in [`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/outputs/skill-patch-geometry-reader.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/outputs/skill-patch-geometry-reader.md) computes token counts, parameter counts, and FLOPs based on image resolution, patch size **P**, hidden dimension **D**, and transformer depth. For a ViT-B/16 configuration, the calculator accounts for 12 blocks with 12 heads each, 4× feed-forward expansion, and embedding matrices to estimate computational requirements for arbitrary input configurations.