How the AI Agent Book Structures Its Curriculum Around the Core Formula Agent = LLM + Context + Tools

The AI Agent Book uses the equation Agent = LLM + Context + Tools as its architectural spine, organizing ten chapters into a systematic exploration of the LLM "brain," context "eyes," and tool "hands" to transform theoretical concepts into production-ready implementations.

The open-source educational repository bojieli/ai-agent-book demonstrates that robust agent engineering requires more than isolated prompting techniques. By anchoring every concept to the core formula Agent = LLM + Context + Tools, the curriculum provides a layered roadmap that progresses from configuring language model providers to orchestrating multi-agent societies. This structure ensures that each component—defined in book/introduction.md at line 29 as the foundational metaphor of brain, eyes, and hands—receives dedicated depth through specific chapters and runnable code experiments.

Deconstructing the Three Pillars

The book treats the formula not as abstract theory but as a concrete architectural blueprint. Each variable in the equation maps to specific cognitive and functional responsibilities within an agent system.

LLM (The Brain): Chapters 1 and 5

The LLM serves as the reasoning engine that interprets instructions and generates outputs. According to book/chapter1.md at line 13, the text establishes that the modern agent fundamentally requires this computational brain, but emphasizes that the brain alone remains insufficient without the supporting pillars. Chapter 1 introduces basic provider configuration (SiliconFlow, Kimi, OpenRouter) while Chapter 5 culminates in a full-stack Coding Agent that leverages the LLM to generate and execute code, demonstrating how the brain directs the other components.

Context (The Eyes): Chapters 2 and 3

Context encompasses every piece of information the model can observe during inference—prompt history, retrieved knowledge, tool definitions, and auxiliary data. Chapter 2 (Context Engineering) delves into prompt design, KV-Cache optimization, and context compression techniques found in book/chapter2.md. Chapter 3 extends this into user-specific memory and knowledge base integration, illustrating how richer, well-structured context directly raises the agent's capability ceiling by expanding what the "eyes" can perceive.

Tools (The Hands): Chapter 4

Tools represent the external capabilities the agent can invoke, including search APIs, file I/O, code execution, and computer-use interfaces. As stated in chapter4/README.md at line 3, "Tools are the Agent's hands," and this chapter provides concrete implementations using the Model Context Protocol (MCP). The content covers perception tools, execution environments, and active tool discovery mechanisms, giving the LLM brain the physical capability to manipulate external systems.

Architectural Roadmap: Mapping the Formula to 10 Chapters

The book's progression follows a logical construction sequence: build the components, combine them, measure their performance, refine them, and scale them. The following table illustrates how each chapter expands upon the core formula:

Formula Component Chapters Engineering Focus
LLM (Brain) 1, 5 Provider configuration and full-stack coding agent implementation
Context (Eyes) 2, 3 Prompt engineering, retrieval-augmented generation, and memory systems
Tools (Hands) 4 MCP protocol, tool categories (perception/execution/collaboration)
Interaction & Expanded Action Space 6 Asynchronous, multimodal, and robotic interaction extending both context and tools
Evaluation 7 Metrics and statistical significance for measuring component effectiveness
Post-Training 8 SFT/RL fine-tuning to improve how the LLM utilizes context and tools
Continual Evolution 9 Using execution traces to update knowledge, policies, and tool definitions
Multi-Agent Collaboration 10 Societies of agents sharing context, delegating tools, and co-evolving

Chapters 1 through 6 focus on constructing the three pillars and demonstrating their integration. Chapters 7 through 9 establish the feedback loops necessary to measure, refine, and evolve the assembled agent. Chapter 10 scales the identical formula from single agents to collaborative networks, where each agent still operates under the LLM + Context + Tools rule but shares resources across the collective.

Practical Implementation: The Formula in Code

The repository materializes the abstract formula into executable Python. The following example from chapter1/context/main.py demonstrates how the three components instantiate a working agent:

from agent import ContextAwareAgent, ContextMode

# 1. Initialize the three pillars: LLM credential, context mode, and provider

agent = ContextAwareAgent(
    api_key="YOUR_API_KEY",          # LLM (Brain) credential

    context_mode=ContextMode.FULL,  # Context (Eyes) configuration

    provider="siliconflow",          # LLM provider selection

    model=None                       # Use provider default model

)

# 2. Define a task requiring rich context and specific tools

task = """
Analyze this PDF (https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf):
- Extract all monetary amounts.
- Convert them to USD, EUR and JPY.
- Summarize the total in each currency.
"""

# 3. Execute: the agent loop stitches LLM reasoning, context retrieval, and tool invocation

result = agent.execute_task(task)

# 4. Inspect the integration results

print("✅ Completed:", result.get("completed"))
print("🛠️ Tool calls:", len(result["trajectory"].tool_calls))
print("🔎 Final answer:", result.get("final_answer"))

This implementation explicitly demonstrates Agent = LLM + Context + Tools in action:

  • Agent creation configures the LLM provider (brain) and selects ContextMode.FULL (eyes)
  • Task definition provides the contextual input (PDF URL) and implicitly requires tool capabilities (PDF parsing, currency conversion)
  • Execution runs the internal loop that binds reasoning with retrieval and external action
  • Result inspection yields metrics discussed in Chapters 7-9, including completion status and tool call frequency

Summary

  • The AI Agent Book anchors its entire curriculum to the formula Agent = LLM + Context + Tools, treating it as the architectural spine rather than a marketing tagline.
  • Chapter 1 and Chapter 5 establish the LLM as the "brain," while Chapters 2-3 develop the "eyes" through context engineering and Chapter 4 builds the "hands" via tool implementations.
  • The progression moves from component construction (Chapters 1-6) through evaluation and refinement (Chapters 7-9) to multi-agent scaling (Chapter 10), with each stage explicitly referencing the core formula.
  • Runnable code in chapter1/context/main.py provides a concrete instantiation, showing how ContextAwareAgent unifies provider configuration, context modes, and tool invocation into a single execution loop.
  • Configuration logic resides in config.py, which serves as the central registry for LLM providers and model defaults referenced throughout the experiments.

Frequently Asked Questions

What does the Agent = LLM + Context + Tools formula represent in the book?

According to book/introduction.md at line 29, this equation represents the minimal complete architecture for an autonomous agent. The LLM acts as the reasoning brain, Context provides the sensory input (everything the model sees at inference time), and Tools serve as the effectors that allow the agent to modify external states. The book uses this formula to determine topic ordering, ensuring students understand each pillar before attempting integration.

Which chapters focus specifically on context engineering?

Chapter 2 provides the deep technical dive into context engineering, covering prompt design, KV-Cache management, and retrieval-augmented generation as documented in book/chapter2.md. Chapter 3 extends this into persistent user memory and knowledge base architectures. Together, these chapters establish how to expand and compress the "eyes" of the agent to optimize perception without exceeding token limits.

How does the repository implement the Tools component?

The "hands" of the agent receive detailed treatment in Chapter 4, which explains the Model Context Protocol (MCP) and categorizes tools into perception, execution, and collaboration types. The implementation includes active tool discovery mechanisms and concrete API integrations, allowing the LLM to invoke external capabilities like file parsing, code execution, and search operations through standardized interfaces.

How does the core formula scale to multi-agent systems?

Chapter 10 applies the identical Agent = LLM + Context + Tools structure to societies of agents, where individual agents retain the three-pillar architecture but share context windows and delegate tool usage across the network. This scaling demonstrates that the formula remains valid at higher levels of abstraction, whether coordinating two agents or twenty, with each participant maintaining its own LLM brain, context eyes, and tool hands while collaborating on collective objectives.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →