# Context Engineering for AI Agents: Shaping the Informational Environment That Drives Decisions

> Discover context engineering for AI agents and learn how to shape their informational environment to enhance decision making. Unlock full AI potential.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-25

---

**Context engineering is the systematic design and management of the information an AI agent perceives at every decision point, defined in `bojieli/ai-agent-book` as "the art of shaping an AI Agent's informational environment."** This practice directly determines how much of a language model's capability can be harnessed in agentic workflows.

According to the source code in [`book-en/chapter2.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter2.md), agents lack permanent memory. Every model call receives two distinct components: a **static prefix** (system prompt plus tool definitions) and a **dynamic trajectory** (conversation history, tool results, and status messages). How these elements are assembled separates high-performing agents from inefficient ones.

## What Context Engineering Controls

Context engineering implements the "Context & Tools" layer of an agent harness, deciding **what** the agent sees and **how** that information is structured. A well-engineered context supplies precise background knowledge—code layout, process rules, environment configuration—enabling smaller models to outperform larger, poorly-contextualized alternatives.

## The Five Architectural Layers of Context

The repository defines five interconnected layers in [`book-en/chapter2.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter2.md) that govern agent perception:

| Layer | Function | Implementation Detail |
|-------|----------|----------------------|
| **System Prompt** | Encodes behavior rules, identity, and constraints | First message in the list; kept stable to preserve KV-Cache hits (`book-en/chapter2.md#L49-L53`) |
| **Tool Definitions** | Declares available functions with JSON schemas | Sent in top-level `tools` field, not embedded in messages (`book-en/chapter2.md#L54-L56`) |
| **Message List** | Holds evolving conversation state | Roles: `system`, `user`, `assistant`, `tool`; appended each turn (`book-en/chapter2.md#L45-L53`) |
| **KV-Cache-Friendly Design** | Minimizes redundant computation | Static prefix reuse eliminates re-encoding of early tokens (`book-en/chapter2.md#L44-L46`) |
| **Dynamic End-Appending** | Adds variable information without cache invalidation | Timestamps and status bars appended as new messages, never injected into system prompt (`book-en/chapter2.md#L44-L46`) |

## Implementing Context Engineering: The ReAct Loop

The book provides a complete Python implementation demonstrating context engineering in practice. The pattern follows four strict steps:

1. **Build the static prefix** — combine `system` message with `tools` schema
2. **Send message list** to the LLM API
3. **Execute and append tool results** when `tool_calls` are returned
4. **Repeat** until a plain `assistant` message signals completion

``` python

# Minimal ReAct loop – demonstrates context engineering in action

from openai import OpenAI

client = OpenAI()

# 1️⃣ Static prefix: system prompt + tool schemas

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_current_time",
            "description": "Get the current date and time in a specific timezone",
            "parameters": {
                "type": "object",
                "properties": {
                    "timezone": {"type": "string", "description": "Timezone name, e.g. America/Vancouver"}
                },
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather for a specific city",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {"type": "string", "description": "City name"},
                    "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
                },
            },
        },
    },
]

# 2️⃣ Initial message list (static prefix + user query)

messages = [
    {"role": "system", "content": "You are a helpful assistant. Use tools when needed."},
    {"role": "user", "content": "What's the current time and weather in Vancouver?"},
]

# 3️⃣ Core ReAct loop – context grows each iteration

while True:
    resp = client.chat.completions.create(
        model="Qwen3-0.6B", messages=messages, tools=tools
    )
    assistant_msg = resp.choices[0].message
    messages.append(assistant_msg)                       # Append model output

    # If no tool calls → final answer

    if not getattr(assistant_msg, "tool_calls", None):
        print(assistant_msg.content)
        break

    # Execute each requested tool and append its result

    for call in assistant_msg.tool_calls:
        result = execute_tool(call.function.name, call.function.arguments)
        messages.append(
            {"role": "tool", "tool_call_id": call.id, "content": result}
        )

```

Three critical principles emerge from this implementation:

- **Static prefix stability** — The `system` message and `tools` array never change, enabling the model to reuse KV-Cache states and reducing latency dramatically
- **Trajectory preservation** — Each iteration appends rather than replaces, maintaining complete history for multi-turn reasoning
- **Termination guarantee** — The loop only exits on a plain `assistant` message, ensuring the final answer incorporates all accumulated tool results

## The KV-Cache Optimization Strategy

Performance optimization is central to effective context engineering. By keeping the static prefix unchanged across calls, agents exploit the key-value cache mechanism in modern transformers. As noted in `book-en/chapter2.md#L44-L46`, this design avoids "cache-miss" slowdowns that occur when variable data contaminates the system prompt.

Dynamic information—timestamps, progress indicators, status updates—must always be appended as new messages at the trajectory's end. This discipline preserves computational efficiency without sacrificing information richness.

## Source Files Reference

| File | Location | Relevance |
|------|----------|-----------|
| [`book-en/chapter2.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter2.md) | `main/book-en/chapter2.md#L1-L70` | Core definition, message-role taxonomy, KV-Cache discussion, complete Python example |
| [`book-en/chapter1.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter1.md) | `main/book-en/chapter1.md#L7-L14` | Harness perspective: context as the "eyes" of an agent |
| [`README.en.md`](https://github.com/bojieli/ai-agent-book/blob/main/README.en.md) | `main/README.en.md#L1-L10` | High-level overview with chapter pointers |

## Summary

- **Context engineering** shapes every decision an AI agent makes by controlling its informational environment
- **Static prefix** (system + tools) and **dynamic trajectory** (message history) form the complete context structure
- **KV-Cache-friendly design** requires keeping the static prefix stable while appending variable data at the end
- The **ReAct loop pattern** in [`book-en/chapter2.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter2.md) demonstrates practical implementation with tool calling and trajectory management
- Well-engineered context enables **smaller models to outperform larger ones** through precise information structuring

## Frequently Asked Questions

### What is the difference between context engineering and prompt engineering?

**Prompt engineering optimizes individual inputs; context engineering designs the entire informational system an agent operates within.** Prompt engineering might craft a single effective question, while context engineering—as defined in [`book-en/chapter2.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter2.md)—manages the static prefix, tool definitions, message trajectory, and KV-Cache optimization across all agent interactions.

### Why is KV-Cache preservation important for AI agents?

**KV-Cache preservation eliminates redundant computation and reduces latency.** When the static prefix remains unchanged, the model reuses pre-computed key-value states for early tokens. According to `book-en/chapter2.md#L44-L46`, injecting dynamic data into the system prompt invalidates this cache, forcing expensive re-encoding on every call.

### How do tool definitions fit into context engineering?

**Tool definitions are part of the static prefix but transmitted separately from the message list.** In the API implementation shown in `book-en/chapter2.md#L54-L56`, tools are passed in the top-level `tools` field rather than embedded as messages. This separation maintains clean message history while ensuring the model understands available capabilities.

### Can effective context engineering compensate for smaller model size?

**Yes—precise context engineering often outperforms larger, poorly-contextualized models.** As stated in `book-en/chapter2.md#L19-L22`, supplying an agent with exact background knowledge (code layout, process rules, configuration) enables modest-size models to exceed the capabilities of larger alternatives lacking structured context.