# How the AI Agent Book Structures Its Curriculum Around the Core Formula Agent = LLM + Context + Tools

> Discover how the AI Agent Book structures its curriculum around the core formula Agent = LLM + Context + Tools transforming theory into production ready AI agents.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: architecture
- Published: 2026-08-23

---

**The AI Agent Book uses the equation `Agent = LLM + Context + Tools` as its architectural spine, organizing ten chapters into a systematic exploration of the LLM "brain," context "eyes," and tool "hands" to transform theoretical concepts into production-ready implementations.**

The open-source educational repository `bojieli/ai-agent-book` demonstrates that robust agent engineering requires more than isolated prompting techniques. By anchoring every concept to the core formula **Agent = LLM + Context + Tools**, the curriculum provides a layered roadmap that progresses from configuring language model providers to orchestrating multi-agent societies. This structure ensures that each component—defined in [`book/introduction.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/introduction.md) at line 29 as the foundational metaphor of brain, eyes, and hands—receives dedicated depth through specific chapters and runnable code experiments.

## Deconstructing the Three Pillars

The book treats the formula not as abstract theory but as a concrete architectural blueprint. Each variable in the equation maps to specific cognitive and functional responsibilities within an agent system.

### LLM (The Brain): Chapters 1 and 5

The **LLM** serves as the reasoning engine that interprets instructions and generates outputs. According to [`book/chapter1.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/chapter1.md) at line 13, the text establishes that the modern agent fundamentally requires this computational brain, but emphasizes that the brain alone remains insufficient without the supporting pillars. Chapter 1 introduces basic provider configuration (SiliconFlow, Kimi, OpenRouter) while Chapter 5 culminates in a full-stack *Coding Agent* that leverages the LLM to generate and execute code, demonstrating how the brain directs the other components.

### Context (The Eyes): Chapters 2 and 3

**Context** encompasses every piece of information the model can observe during inference—prompt history, retrieved knowledge, tool definitions, and auxiliary data. Chapter 2 (Context Engineering) delves into prompt design, KV-Cache optimization, and context compression techniques found in [`book/chapter2.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/chapter2.md). Chapter 3 extends this into user-specific memory and knowledge base integration, illustrating how richer, well-structured context directly raises the agent's capability ceiling by expanding what the "eyes" can perceive.

### Tools (The Hands): Chapter 4

**Tools** represent the external capabilities the agent can invoke, including search APIs, file I/O, code execution, and computer-use interfaces. As stated in [`chapter4/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter4/README.md) at line 3, "Tools are the Agent's hands," and this chapter provides concrete implementations using the Model Context Protocol (MCP). The content covers perception tools, execution environments, and active tool discovery mechanisms, giving the LLM brain the physical capability to manipulate external systems.

## Architectural Roadmap: Mapping the Formula to 10 Chapters

The book's progression follows a logical construction sequence: build the components, combine them, measure their performance, refine them, and scale them. The following table illustrates how each chapter expands upon the core formula:

| Formula Component | Chapters | Engineering Focus |
|---|---|---|
| **LLM (Brain)** | 1, 5 | Provider configuration and full-stack coding agent implementation |
| **Context (Eyes)** | 2, 3 | Prompt engineering, retrieval-augmented generation, and memory systems |
| **Tools (Hands)** | 4 | MCP protocol, tool categories (perception/execution/collaboration) |
| **Interaction & Expanded Action Space** | 6 | Asynchronous, multimodal, and robotic interaction extending both context and tools |
| **Evaluation** | 7 | Metrics and statistical significance for measuring component effectiveness |
| **Post-Training** | 8 | SFT/RL fine-tuning to improve how the LLM utilizes context and tools |
| **Continual Evolution** | 9 | Using execution traces to update knowledge, policies, and tool definitions |
| **Multi-Agent Collaboration** | 10 | Societies of agents sharing context, delegating tools, and co-evolving |

Chapters 1 through 6 focus on **constructing** the three pillars and demonstrating their integration. Chapters 7 through 9 establish the **feedback loops** necessary to measure, refine, and evolve the assembled agent. Chapter 10 **scales** the identical formula from single agents to collaborative networks, where each agent still operates under the `LLM + Context + Tools` rule but shares resources across the collective.

## Practical Implementation: The Formula in Code

The repository materializes the abstract formula into executable Python. The following example from [`chapter1/context/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter1/context/main.py) demonstrates how the three components instantiate a working agent:

```python
from agent import ContextAwareAgent, ContextMode

# 1. Initialize the three pillars: LLM credential, context mode, and provider

agent = ContextAwareAgent(
    api_key="YOUR_API_KEY",          # LLM (Brain) credential

    context_mode=ContextMode.FULL,  # Context (Eyes) configuration

    provider="siliconflow",          # LLM provider selection

    model=None                       # Use provider default model

)

# 2. Define a task requiring rich context and specific tools

task = """
Analyze this PDF (https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf):
- Extract all monetary amounts.
- Convert them to USD, EUR and JPY.
- Summarize the total in each currency.
"""

# 3. Execute: the agent loop stitches LLM reasoning, context retrieval, and tool invocation

result = agent.execute_task(task)

# 4. Inspect the integration results

print("✅ Completed:", result.get("completed"))
print("🛠️ Tool calls:", len(result["trajectory"].tool_calls))
print("🔎 Final answer:", result.get("final_answer"))

```

This implementation explicitly demonstrates **Agent = LLM + Context + Tools** in action:
- **Agent creation** configures the LLM provider (brain) and selects `ContextMode.FULL` (eyes)
- **Task definition** provides the contextual input (PDF URL) and implicitly requires tool capabilities (PDF parsing, currency conversion)
- **Execution** runs the internal loop that binds reasoning with retrieval and external action
- **Result inspection** yields metrics discussed in Chapters 7-9, including completion status and tool call frequency

## Summary

- The **AI Agent Book** anchors its entire curriculum to the formula `Agent = LLM + Context + Tools`, treating it as the architectural spine rather than a marketing tagline.
- **Chapter 1** and **Chapter 5** establish the LLM as the "brain," while **Chapters 2-3** develop the "eyes" through context engineering and **Chapter 4** builds the "hands" via tool implementations.
- The progression moves from component construction (Chapters 1-6) through evaluation and refinement (Chapters 7-9) to multi-agent scaling (Chapter 10), with each stage explicitly referencing the core formula.
- Runnable code in [`chapter1/context/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter1/context/main.py) provides a concrete instantiation, showing how `ContextAwareAgent` unifies provider configuration, context modes, and tool invocation into a single execution loop.
- Configuration logic resides in [`config.py`](https://github.com/bojieli/ai-agent-book/blob/main/config.py), which serves as the central registry for LLM providers and model defaults referenced throughout the experiments.

## Frequently Asked Questions

### What does the Agent = LLM + Context + Tools formula represent in the book?

According to [`book/introduction.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/introduction.md) at line 29, this equation represents the minimal complete architecture for an autonomous agent. The **LLM** acts as the reasoning brain, **Context** provides the sensory input (everything the model sees at inference time), and **Tools** serve as the effectors that allow the agent to modify external states. The book uses this formula to determine topic ordering, ensuring students understand each pillar before attempting integration.

### Which chapters focus specifically on context engineering?

**Chapter 2** provides the deep technical dive into context engineering, covering prompt design, KV-Cache management, and retrieval-augmented generation as documented in [`book/chapter2.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/chapter2.md). **Chapter 3** extends this into persistent user memory and knowledge base architectures. Together, these chapters establish how to expand and compress the "eyes" of the agent to optimize perception without exceeding token limits.

### How does the repository implement the Tools component?

The "hands" of the agent receive detailed treatment in **Chapter 4**, which explains the Model Context Protocol (MCP) and categorizes tools into perception, execution, and collaboration types. The implementation includes active tool discovery mechanisms and concrete API integrations, allowing the LLM to invoke external capabilities like file parsing, code execution, and search operations through standardized interfaces.

### How does the core formula scale to multi-agent systems?

**Chapter 10** applies the identical `Agent = LLM + Context + Tools` structure to societies of agents, where individual agents retain the three-pillar architecture but share context windows and delegate tool usage across the network. This scaling demonstrates that the formula remains valid at higher levels of abstraction, whether coordinating two agents or twenty, with each participant maintaining its own LLM brain, context eyes, and tool hands while collaborating on collective objectives.