Building Multimodal AI Systems: A Complete Guide to the AI Engineering from Scratch Curriculum

Building multimodal AI systems requires implementing vision transformers, cross-modal embeddings, and agent architectures from mathematical first principles, all of which are covered in the open-source AI Engineering from Scratch repository.

The AI Engineering from Scratch curriculum provides a structured 20-phase pathway for building multimodal AI systems incrementally. Each phase contains self-contained lessons with runnable code, narrative documentation, and reusable artifacts that form a personal AI toolkit. The repository implements everything from scratch before introducing production frameworks, ensuring deep understanding of the underlying mathematics.

Curriculum Architecture and Phase Structure

The repository organizes content into 20 phases (numbered 0–19), with each phase focusing on a specific domain such as mathematics, machine learning fundamentals, or multimodal AI. Every lesson follows a consistent three-part structure:

  • code/ – Implementations in Python, TypeScript, Rust, or Julia
  • docs/ – Narrative explanations (e.g., docs/en.md)
  • outputs/ – Reusable artifacts including prompts, skills, agents, and MCP servers

According to the README.md, this scaffolded approach ensures that Phase 12 (Multimodal AI) builds upon a solid foundation of transformer architectures and deep learning principles established in preceding phases.

The Build-It / Use-It Methodology

Each lesson employs a dual-track approach that separates theoretical construction from practical application. First, you build algorithms using raw mathematical operations and minimal dependencies. Then, you use optimized production libraries like PyTorch or JAX to achieve the same results efficiently.

This methodology appears in [phases/14-agent-engineering/01-the-agent-loop/code/main.py](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/01-the-agent-loop/code/main.py), where agent loops are implemented from scratch before framework-based alternatives are introduced. In Phase 12, this pattern ensures you understand the matrix operations behind vision transformers before using high-level implementations.

Phase 12: Vision Transformers and Multimodal Integration

Phase 12 specifically targets multimodal AI systems, covering vision-language models, audio-language integration, and cross-modal retrieval mechanisms. The phase demonstrates how to construct systems that can see, hear, and act through token-level vision patches and CLIP contrastive training.

The entry lesson for this phase is located at [phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py). This file implements Vision Transformer patch tokenization from first principles, providing the foundation for subsequent lessons on cross-modal attention and multimodal agent engineering.

Installing Reusable AI Skills

The repository includes a skill installation system that exports lesson artifacts into reusable components. The [scripts/install_skills.py](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py) script installs prompts, agents, and MCP servers from lesson outputs/ directories into your local environment.

To install all available skills:

python3 scripts/install_skills.py

After installation, you can invoke these skills in compatible AI assistants like Claude or Cursor. For example, the find-your-level skill assesses your knowledge and recommends appropriate starting phases, while other skills provide specialized prompts for multimodal agent construction.

Hands-On Implementation Guide

To begin building multimodal AI systems with this curriculum, execute the following steps:

  1. Clone the repository and navigate to the multimodal phase:
git clone https://github.com/rohitg00/ai-engineering-from-scratch.git
cd ai-engineering-from-scratch
  1. Run the Vision Transformer implementation directly:
python phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py
  1. Generate the documentation website locally using the build system:
node site/build.js

The [site/build.js](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js) script processes the markdown sources and generates site/data.js and site/index.html. Opening site/index.html in a browser provides an interactive interface with phase tables and lesson diagrams.

Key Files for Multimodal Development

Understanding the repository structure is essential for navigating the multimodal content:

Summary

Frequently Asked Questions

What prerequisites are required before studying Phase 12 on multimodal AI?

You should complete Phases 0 through 11, which cover linear algebra, calculus, probability, machine learning fundamentals, deep learning, and transformer architectures. The curriculum is designed as a dependency graph where each phase builds upon the mathematical and computational concepts established in previous phases, ensuring you have the necessary foundation for understanding vision transformers and cross-modal embeddings.

How does the Vision Transformer implementation in Phase 12 differ from standard tutorials?

The implementation in [phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py) constructs patch tokenization and multi-head attention mechanisms using raw NumPy operations before introducing PyTorch optimizations. This contrasts with standard tutorials that begin with high-level nn.Module classes, ensuring you understand the underlying matrix multiplication and positional encoding schemes that enable multimodal systems to process visual and textual information simultaneously.

Can I reuse the agents and skills from this repository in production workflows?

Yes. Running [scripts/install_skills.py](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py) exports artifacts from lesson outputs/ directories, including MCP servers and prompt templates defined in files like [.claude/skills/find-your-level/SKILL.md](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/.claude/skills/find-your-level/SKILL.md). These skills integrate with Claude, Cursor, and other AI assistants, allowing you to deploy multimodal agent architectures and specialized prompts directly into your development environment.

What is the role of the site/build.js file in the curriculum?

The [site/build.js](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js) script compiles the markdown documentation, code examples, and interactive diagrams into a static website. It generates site/data.js and site/index.html, transforming the repository's raw content into the browseable learning platform at aiengineeringfromscratch.com. This build system ensures that mathematical notation renders correctly and phase dependencies are displayed as interactive tables.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →