Context Engineering for AI Agents: Shaping the Informational Environment That Drives Decisions
Context engineering is the systematic design and management of the information an AI agent perceives at every decision point, defined in bojieli/ai-agent-book as "the art of shaping an AI Agent's informational environment." This practice directly determines how much of a language model's capability can be harnessed in agentic workflows.
According to the source code in book-en/chapter2.md, agents lack permanent memory. Every model call receives two distinct components: a static prefix (system prompt plus tool definitions) and a dynamic trajectory (conversation history, tool results, and status messages). How these elements are assembled separates high-performing agents from inefficient ones.
What Context Engineering Controls
Context engineering implements the "Context & Tools" layer of an agent harness, deciding what the agent sees and how that information is structured. A well-engineered context supplies precise background knowledge—code layout, process rules, environment configuration—enabling smaller models to outperform larger, poorly-contextualized alternatives.
The Five Architectural Layers of Context
The repository defines five interconnected layers in book-en/chapter2.md that govern agent perception:
| Layer | Function | Implementation Detail |
|---|---|---|
| System Prompt | Encodes behavior rules, identity, and constraints | First message in the list; kept stable to preserve KV-Cache hits (book-en/chapter2.md#L49-L53) |
| Tool Definitions | Declares available functions with JSON schemas | Sent in top-level tools field, not embedded in messages (book-en/chapter2.md#L54-L56) |
| Message List | Holds evolving conversation state | Roles: system, user, assistant, tool; appended each turn (book-en/chapter2.md#L45-L53) |
| KV-Cache-Friendly Design | Minimizes redundant computation | Static prefix reuse eliminates re-encoding of early tokens (book-en/chapter2.md#L44-L46) |
| Dynamic End-Appending | Adds variable information without cache invalidation | Timestamps and status bars appended as new messages, never injected into system prompt (book-en/chapter2.md#L44-L46) |
Implementing Context Engineering: The ReAct Loop
The book provides a complete Python implementation demonstrating context engineering in practice. The pattern follows four strict steps:
- Build the static prefix — combine
systemmessage withtoolsschema - Send message list to the LLM API
- Execute and append tool results when
tool_callsare returned - Repeat until a plain
assistantmessage signals completion
# Minimal ReAct loop – demonstrates context engineering in action
from openai import OpenAI
client = OpenAI()
# 1️⃣ Static prefix: system prompt + tool schemas
tools = [
{
"type": "function",
"function": {
"name": "get_current_time",
"description": "Get the current date and time in a specific timezone",
"parameters": {
"type": "object",
"properties": {
"timezone": {"type": "string", "description": "Timezone name, e.g. America/Vancouver"}
},
},
},
},
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a specific city",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
},
},
},
]
# 2️⃣ Initial message list (static prefix + user query)
messages = [
{"role": "system", "content": "You are a helpful assistant. Use tools when needed."},
{"role": "user", "content": "What's the current time and weather in Vancouver?"},
]
# 3️⃣ Core ReAct loop – context grows each iteration
while True:
resp = client.chat.completions.create(
model="Qwen3-0.6B", messages=messages, tools=tools
)
assistant_msg = resp.choices[0].message
messages.append(assistant_msg) # Append model output
# If no tool calls → final answer
if not getattr(assistant_msg, "tool_calls", None):
print(assistant_msg.content)
break
# Execute each requested tool and append its result
for call in assistant_msg.tool_calls:
result = execute_tool(call.function.name, call.function.arguments)
messages.append(
{"role": "tool", "tool_call_id": call.id, "content": result}
)
Three critical principles emerge from this implementation:
- Static prefix stability — The
systemmessage andtoolsarray never change, enabling the model to reuse KV-Cache states and reducing latency dramatically - Trajectory preservation — Each iteration appends rather than replaces, maintaining complete history for multi-turn reasoning
- Termination guarantee — The loop only exits on a plain
assistantmessage, ensuring the final answer incorporates all accumulated tool results
The KV-Cache Optimization Strategy
Performance optimization is central to effective context engineering. By keeping the static prefix unchanged across calls, agents exploit the key-value cache mechanism in modern transformers. As noted in book-en/chapter2.md#L44-L46, this design avoids "cache-miss" slowdowns that occur when variable data contaminates the system prompt.
Dynamic information—timestamps, progress indicators, status updates—must always be appended as new messages at the trajectory's end. This discipline preserves computational efficiency without sacrificing information richness.
Source Files Reference
| File | Location | Relevance |
|---|---|---|
book-en/chapter2.md |
main/book-en/chapter2.md#L1-L70 |
Core definition, message-role taxonomy, KV-Cache discussion, complete Python example |
book-en/chapter1.md |
main/book-en/chapter1.md#L7-L14 |
Harness perspective: context as the "eyes" of an agent |
README.en.md |
main/README.en.md#L1-L10 |
High-level overview with chapter pointers |
Summary
- Context engineering shapes every decision an AI agent makes by controlling its informational environment
- Static prefix (system + tools) and dynamic trajectory (message history) form the complete context structure
- KV-Cache-friendly design requires keeping the static prefix stable while appending variable data at the end
- The ReAct loop pattern in
book-en/chapter2.mddemonstrates practical implementation with tool calling and trajectory management - Well-engineered context enables smaller models to outperform larger ones through precise information structuring
Frequently Asked Questions
What is the difference between context engineering and prompt engineering?
Prompt engineering optimizes individual inputs; context engineering designs the entire informational system an agent operates within. Prompt engineering might craft a single effective question, while context engineering—as defined in book-en/chapter2.md—manages the static prefix, tool definitions, message trajectory, and KV-Cache optimization across all agent interactions.
Why is KV-Cache preservation important for AI agents?
KV-Cache preservation eliminates redundant computation and reduces latency. When the static prefix remains unchanged, the model reuses pre-computed key-value states for early tokens. According to book-en/chapter2.md#L44-L46, injecting dynamic data into the system prompt invalidates this cache, forcing expensive re-encoding on every call.
How do tool definitions fit into context engineering?
Tool definitions are part of the static prefix but transmitted separately from the message list. In the API implementation shown in book-en/chapter2.md#L54-L56, tools are passed in the top-level tools field rather than embedded as messages. This separation maintains clean message history while ensuring the model understands available capabilities.
Can effective context engineering compensate for smaller model size?
Yes—precise context engineering often outperforms larger, poorly-contextualized models. As stated in book-en/chapter2.md#L19-L22, supplying an agent with exact background knowledge (code layout, process rules, configuration) enables modest-size models to exceed the capabilities of larger alternatives lacking structured context.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →