# How to Implement RAG in Agent Context Windows: A 12-Factor Agents Guide

> Learn to implement RAG in agent context windows by encoding documents as structured events for efficient LLM reasoning over current information. Explore the 12-Factor Agents guide.

- Repository: [HumanLayer/12-factor-agents](https://github.com/humanlayer/12-factor-agents)
- Tags: how-to-guide
- Published: 2026-05-19

---

**Retrieval-Augmented Generation (RAG) integrates external knowledge into an agent's context window by encoding retrieved documents as structured events, allowing the LLM to reason over up-to-date information while maintaining token efficiency.**

In the `humanlayer/12-factor-agents` framework, implementing RAG requires treating retrieved documents as first-class components of the context window alongside prompts, tool calls, and conversation history. This approach ensures agents can access external knowledge without breaking the deterministic loops that govern their behavior.

## Understanding RAG in the 12-Factor Agents Framework

According to the source code in [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md) (lines 16-18), **RAG** is treated as one of the core pieces of *context* that an agent must own and manage. Unlike traditional implementations that may inject retrieved text directly into prompts, the 12-Factor methodology requires explicit ownership of the entire context window.

This means retrieval operations become explicit events in the agent's event loop. When the agent determines it needs external knowledge, it generates a retrieval event, fetches relevant documents, and inserts them back into the context as structured data. The framework stores these as distinct event types—`retrieve_document` and `retrieve_document_result`—enabling precise tracking and debugging of what information enters the LLM's view.

## The Context Window Architecture for RAG

The architecture relies on five interconnected components that transform raw retrieval operations into LLM-ready context.

### The Context Engine

The **Context Engine** aggregates all information into a unified window before serialization. As implemented in [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md) (lines 33-41), this includes:

- System instructions and base prompts
- Retrieved RAG documents (raw text, embeddings, or structured extracts)
- Prior tool calls and their execution results
- Persisted memory of past conversations

### The Retrieval Layer

When the agent determines external knowledge is required, it executes a retrieval step. This involves querying a **vector store** (such as FAISS or Pinecone) using a search phrase extracted from the current thread context. The framework supports dedicated "search-intent" events that trigger this workflow without ambiguity.

### RAG Document Integration

Retrieved chunks enter the context window through a custom formatting system. The framework encourages **XML-style tags** (e.g., `<retrieved_doc>`) inserted between event markers, enabling the LLM to readily identify and reason over each piece of external data according to [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md) (lines 71-84).

### The Decision Loop

After the LLM processes the enriched context, it produces the next intent—either answering the user, calling a tool, or requesting additional retrieval. If the initial RAG results prove insufficient, the loop repeats, fetching supplementary documents and rebuilding the context window dynamically.

### Token Efficiency Controls

Because context windows face hard limits, the agent must **filter** retrieved documents for relevance and compress them through summarization or selective field extraction before insertion, as noted in [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md) (lines 29-34).

## Implementing RAG Retrieval Events

The core implementation relies on event-driven architecture where retrieval operations are explicit events in a thread. Here is the minimal Python structure extracted from the repository:

```python
from typing import List, Literal, Union
import yaml

class Thread:
    events: List[Event]

class Event:
    # type can be a tool name, retrieval tag, etc.

    type: Literal[
        "list_git_tags", "deploy_backend", "retrieve_document",
        "retrieve_document_result", "error", "human_response"
    ]
    data: Union[str, dict]   # flexible payload

def event_to_prompt(event: Event) -> str:
    payload = event.data if isinstance(event.data, str) else yaml.safe_dump(event.data)
    return f"<{event.type}>\n{payload}\n</{event.type}>"

def thread_to_prompt(thread: Thread) -> str:
    # Join all events into a single user-message context window

    return "\n\n".join(event_to_prompt(e) for e in thread.events)

```

Key implementation details:

- **`retrieve_document`** and **`retrieve_document_result`** are distinct event types that bookend the retrieval operation
- **`event_to_prompt`** serializes each event into XML-style tags that preserve structure
- **`thread_to_prompt`** collapses the entire thread into a single user message, minimizing token overhead

## Encoding Retrieved Documents for the LLM

Standard message-based APIs (system/user/assistant) prove insufficient for complex RAG scenarios. The framework advocates for **custom context formats** that pack dense information into single messages with clear delimiters.

Example usage showing how RAG fits into the event loop:

```python

# Example usage:

thread = Thread(events=[
    Event(type="user_message", data="What is the company's refund policy?"),
    Event(type="retrieve_document", data={"query": "refund policy 2024"}),
    Event(type="retrieve_document_result", data={"content": "Our refund policy ..."})
])
prompt = thread_to_prompt(thread)
next_step = await determine_next_step(prompt)

```

As documented in [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md) (lines 69-77), this custom formatting reduces token overhead while giving the model explicit boundaries between retrieved knowledge and conversation history.

## Token Efficiency and Safety Considerations

When implementing RAG in agent context windows, two constraints dominate: **token limits** and **data privacy**.

Always filter retrieved documents for relevance before encoding them into the context window. The framework emphasizes *controlling* exactly what reaches the LLM, requiring preprocessing steps that remove sensitive content from external documents according to [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md) (lines 30-33).

Additionally, implement compression strategies such as:

- Summarizing long documents before insertion
- Selecting only specific fields from structured retrieval results
- Deduplicating redundant information across multiple retrieved chunks

## Summary

- **RAG is context**: In the 12-Factor Agents framework, retrieved documents are first-class context components, not prompt injections.
- **Event-driven architecture**: Use explicit `retrieve_document` and `retrieve_document_result` events to track knowledge retrieval in the agent's event loop.
- **Custom formatting**: Encode retrieved data using XML-style tags via `event_to_prompt()` and `thread_to_prompt()` to maintain structure while minimizing tokens.
- **Token efficiency**: Filter and compress retrieved documents before insertion to respect context window limits.
- **Safety first**: Sanitize external data before adding it to the context window to prevent sensitive information leakage.

## Frequently Asked Questions

### How does the 12-Factor Agents framework differ from standard RAG implementations?

Standard RAG often injects retrieved text directly into system prompts, obscuring the retrieval step from the agent's logic. The 12-Factor approach treats retrieval as an explicit event type within the `Thread` class, enabling the agent to decide when to retrieve, handle retrieval failures, and track exactly what knowledge influenced each decision according to [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md).

### What file contains the core logic for context window management?

The primary documentation and implementation guidance resides in [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md), which details how RAG documents integrate with other context components like tool calls and memory. Supporting context appears in [`content/factor-02-own-your-prompts.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-02-own-your-prompts.md) and [`content/factor-04-tools-are-structured-outputs.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-04-tools-are-structured-outputs.md).

### Why use XML-style tags instead of standard JSON for retrieved documents?

XML-style tags created by `event_to_prompt()` provide clear visual delimiters that help LLMs distinguish between different context components (user messages, tool results, retrieved documents). This format reduces token overhead compared to repeated JSON schemas while maintaining human readability for debugging, as shown in the framework's code examples.

### Can the agent perform multiple retrieval steps in one conversation turn?

Yes. The decision loop allows the agent to issue multiple `retrieve_document` events if the initial results prove insufficient. The `thread_to_prompt()` function aggregates all events chronologically, enabling iterative retrieval where each subsequent query can reference previous results before calling `determine_next_step()`.