# Building Multimodal AI Systems: A Complete Guide to the AI Engineering from Scratch Curriculum

> Learn to build multimodal AI systems from scratch with this comprehensive guide covering vision transformers, cross-modal embeddings, and agent architectures. Explore the AI Engineering from Scratch curriculum.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: tutorial
- Published: 2026-07-19

---

**Building multimodal AI systems requires implementing vision transformers, cross-modal embeddings, and agent architectures from mathematical first principles, all of which are covered in the open-source AI Engineering from Scratch repository.**

The [**AI Engineering from Scratch**](https://github.com/rohitg00/ai-engineering-from-scratch) curriculum provides a structured 20-phase pathway for building multimodal AI systems incrementally. Each phase contains self-contained lessons with runnable code, narrative documentation, and reusable artifacts that form a personal AI toolkit. The repository implements everything from scratch before introducing production frameworks, ensuring deep understanding of the underlying mathematics.

## Curriculum Architecture and Phase Structure

The repository organizes content into **20 phases** (numbered 0–19), with each phase focusing on a specific domain such as mathematics, machine learning fundamentals, or multimodal AI. Every lesson follows a consistent three-part structure:

- `code/` – Implementations in Python, TypeScript, Rust, or Julia
- `docs/` – Narrative explanations (e.g., [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md))
- `outputs/` – Reusable artifacts including prompts, skills, agents, and MCP servers

According to the [README.md](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/README.md), this scaffolded approach ensures that Phase 12 (Multimodal AI) builds upon a solid foundation of transformer architectures and deep learning principles established in preceding phases.

## The Build-It / Use-It Methodology

Each lesson employs a **dual-track approach** that separates theoretical construction from practical application. First, you **build** algorithms using raw mathematical operations and minimal dependencies. Then, you **use** optimized production libraries like PyTorch or JAX to achieve the same results efficiently.

This methodology appears in [[`phases/14-agent-engineering/01-the-agent-loop/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/01-the-agent-loop/code/main.py)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/01-the-agent-loop/code/main.py), where agent loops are implemented from scratch before framework-based alternatives are introduced. In Phase 12, this pattern ensures you understand the matrix operations behind vision transformers before using high-level implementations.

## Phase 12: Vision Transformers and Multimodal Integration

Phase 12 specifically targets **multimodal AI systems**, covering vision-language models, audio-language integration, and cross-modal retrieval mechanisms. The phase demonstrates how to construct systems that can see, hear, and act through token-level vision patches and CLIP contrastive training.

The entry lesson for this phase is located at [[`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py). This file implements Vision Transformer patch tokenization from first principles, providing the foundation for subsequent lessons on cross-modal attention and multimodal agent engineering.

## Installing Reusable AI Skills

The repository includes a skill installation system that exports lesson artifacts into reusable components. The [[`scripts/install_skills.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py) script installs prompts, agents, and MCP servers from lesson `outputs/` directories into your local environment.

To install all available skills:

```bash
python3 scripts/install_skills.py

```

After installation, you can invoke these skills in compatible AI assistants like Claude or Cursor. For example, the [`find-your-level`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/.claude/skills/find-your-level/SKILL.md) skill assesses your knowledge and recommends appropriate starting phases, while other skills provide specialized prompts for multimodal agent construction.

## Hands-On Implementation Guide

To begin building multimodal AI systems with this curriculum, execute the following steps:

1. Clone the repository and navigate to the multimodal phase:

```bash
git clone https://github.com/rohitg00/ai-engineering-from-scratch.git
cd ai-engineering-from-scratch

```

2. Run the Vision Transformer implementation directly:

```bash
python phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py

```

3. Generate the documentation website locally using the build system:

```bash
node site/build.js

```

The [[`site/build.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js) script processes the markdown sources and generates [`site/data.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/data.js) and [`site/index.html`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/index.html). Opening [`site/index.html`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/index.html) in a browser provides an interactive interface with phase tables and lesson diagrams.

## Key Files for Multimodal Development

Understanding the repository structure is essential for navigating the multimodal content:

- **[[`README.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/README.md)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/README.md)** – High-level curriculum overview with phase dependencies and getting-started instructions
- **[[`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py)** – First principles implementation of Vision Transformer patch tokenization
- **[[`scripts/install_skills.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py)** – CLI tool for installing reusable prompts and agent skills
- **[[`site/build.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js)** – Static site generator that transforms markdown lessons into the interactive website
- **[[`AGENTS.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/AGENTS.md)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/AGENTS.md)** – Operating manual for AI assistants and contributors interacting with the codebase

## Summary

- The **AI Engineering from Scratch** repository offers a comprehensive 20-phase curriculum for building multimodal AI systems from mathematical foundations to production deployment.
- **Phase 12** focuses specifically on multimodal AI, implementing vision transformers in [[`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py) before introducing framework optimizations.
- The **Build-It / Use-It** methodology ensures you understand raw matrix operations governing cross-modal attention before using PyTorch abstractions.
- **Reusable artifacts** from lesson `outputs/` directories can be installed via [[`scripts/install_skills.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py) for use in Claude, Cursor, and other AI environments.
- The curriculum website is generated by [[`site/build.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js), providing an interactive reference at `aiengineeringfromscratch.com`.

## Frequently Asked Questions

### What prerequisites are required before studying Phase 12 on multimodal AI?

You should complete Phases 0 through 11, which cover linear algebra, calculus, probability, machine learning fundamentals, deep learning, and transformer architectures. The curriculum is designed as a dependency graph where each phase builds upon the mathematical and computational concepts established in previous phases, ensuring you have the necessary foundation for understanding vision transformers and cross-modal embeddings.

### How does the Vision Transformer implementation in Phase 12 differ from standard tutorials?

The implementation in [[`phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py) constructs patch tokenization and multi-head attention mechanisms using raw NumPy operations before introducing PyTorch optimizations. This contrasts with standard tutorials that begin with high-level `nn.Module` classes, ensuring you understand the underlying matrix multiplication and positional encoding schemes that enable multimodal systems to process visual and textual information simultaneously.

### Can I reuse the agents and skills from this repository in production workflows?

Yes. Running [[`scripts/install_skills.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py) exports artifacts from lesson `outputs/` directories, including MCP servers and prompt templates defined in files like [[`.claude/skills/find-your-level/SKILL.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/.claude/skills/find-your-level/SKILL.md)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/.claude/skills/find-your-level/SKILL.md). These skills integrate with Claude, Cursor, and other AI assistants, allowing you to deploy multimodal agent architectures and specialized prompts directly into your development environment.

### What is the role of the site/build.js file in the curriculum?

The [[`site/build.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js)](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js) script compiles the markdown documentation, code examples, and interactive diagrams into a static website. It generates [`site/data.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/data.js) and [`site/index.html`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/index.html), transforming the repository's raw content into the browseable learning platform at `aiengineeringfromscratch.com`. This build system ensures that mathematical notation renders correctly and phase dependencies are displayed as interactive tables.