How to Implement RAG in Agent Context Windows: A 12-Factor Agents Guide

Retrieval-Augmented Generation (RAG) integrates external knowledge into an agent's context window by encoding retrieved documents as structured events, allowing the LLM to reason over up-to-date information while maintaining token efficiency.

In the humanlayer/12-factor-agents framework, implementing RAG requires treating retrieved documents as first-class components of the context window alongside prompts, tool calls, and conversation history. This approach ensures agents can access external knowledge without breaking the deterministic loops that govern their behavior.

Understanding RAG in the 12-Factor Agents Framework

According to the source code in content/factor-03-own-your-context-window.md (lines 16-18), RAG is treated as one of the core pieces of context that an agent must own and manage. Unlike traditional implementations that may inject retrieved text directly into prompts, the 12-Factor methodology requires explicit ownership of the entire context window.

This means retrieval operations become explicit events in the agent's event loop. When the agent determines it needs external knowledge, it generates a retrieval event, fetches relevant documents, and inserts them back into the context as structured data. The framework stores these as distinct event types—retrieve_document and retrieve_document_result—enabling precise tracking and debugging of what information enters the LLM's view.

The Context Window Architecture for RAG

The architecture relies on five interconnected components that transform raw retrieval operations into LLM-ready context.

The Context Engine

The Context Engine aggregates all information into a unified window before serialization. As implemented in content/factor-03-own-your-context-window.md (lines 33-41), this includes:

  • System instructions and base prompts
  • Retrieved RAG documents (raw text, embeddings, or structured extracts)
  • Prior tool calls and their execution results
  • Persisted memory of past conversations

The Retrieval Layer

When the agent determines external knowledge is required, it executes a retrieval step. This involves querying a vector store (such as FAISS or Pinecone) using a search phrase extracted from the current thread context. The framework supports dedicated "search-intent" events that trigger this workflow without ambiguity.

RAG Document Integration

Retrieved chunks enter the context window through a custom formatting system. The framework encourages XML-style tags (e.g., <retrieved_doc>) inserted between event markers, enabling the LLM to readily identify and reason over each piece of external data according to content/factor-03-own-your-context-window.md (lines 71-84).

The Decision Loop

After the LLM processes the enriched context, it produces the next intent—either answering the user, calling a tool, or requesting additional retrieval. If the initial RAG results prove insufficient, the loop repeats, fetching supplementary documents and rebuilding the context window dynamically.

Token Efficiency Controls

Because context windows face hard limits, the agent must filter retrieved documents for relevance and compress them through summarization or selective field extraction before insertion, as noted in content/factor-03-own-your-context-window.md (lines 29-34).

Implementing RAG Retrieval Events

The core implementation relies on event-driven architecture where retrieval operations are explicit events in a thread. Here is the minimal Python structure extracted from the repository:

from typing import List, Literal, Union
import yaml

class Thread:
    events: List[Event]

class Event:
    # type can be a tool name, retrieval tag, etc.

    type: Literal[
        "list_git_tags", "deploy_backend", "retrieve_document",
        "retrieve_document_result", "error", "human_response"
    ]
    data: Union[str, dict]   # flexible payload

def event_to_prompt(event: Event) -> str:
    payload = event.data if isinstance(event.data, str) else yaml.safe_dump(event.data)
    return f"<{event.type}>\n{payload}\n</{event.type}>"

def thread_to_prompt(thread: Thread) -> str:
    # Join all events into a single user-message context window

    return "\n\n".join(event_to_prompt(e) for e in thread.events)

Key implementation details:

  • retrieve_document and retrieve_document_result are distinct event types that bookend the retrieval operation
  • event_to_prompt serializes each event into XML-style tags that preserve structure
  • thread_to_prompt collapses the entire thread into a single user message, minimizing token overhead

Encoding Retrieved Documents for the LLM

Standard message-based APIs (system/user/assistant) prove insufficient for complex RAG scenarios. The framework advocates for custom context formats that pack dense information into single messages with clear delimiters.

Example usage showing how RAG fits into the event loop:


# Example usage:

thread = Thread(events=[
    Event(type="user_message", data="What is the company's refund policy?"),
    Event(type="retrieve_document", data={"query": "refund policy 2024"}),
    Event(type="retrieve_document_result", data={"content": "Our refund policy ..."})
])
prompt = thread_to_prompt(thread)
next_step = await determine_next_step(prompt)

As documented in content/factor-03-own-your-context-window.md (lines 69-77), this custom formatting reduces token overhead while giving the model explicit boundaries between retrieved knowledge and conversation history.

Token Efficiency and Safety Considerations

When implementing RAG in agent context windows, two constraints dominate: token limits and data privacy.

Always filter retrieved documents for relevance before encoding them into the context window. The framework emphasizes controlling exactly what reaches the LLM, requiring preprocessing steps that remove sensitive content from external documents according to content/factor-03-own-your-context-window.md (lines 30-33).

Additionally, implement compression strategies such as:

  • Summarizing long documents before insertion
  • Selecting only specific fields from structured retrieval results
  • Deduplicating redundant information across multiple retrieved chunks

Summary

  • RAG is context: In the 12-Factor Agents framework, retrieved documents are first-class context components, not prompt injections.
  • Event-driven architecture: Use explicit retrieve_document and retrieve_document_result events to track knowledge retrieval in the agent's event loop.
  • Custom formatting: Encode retrieved data using XML-style tags via event_to_prompt() and thread_to_prompt() to maintain structure while minimizing tokens.
  • Token efficiency: Filter and compress retrieved documents before insertion to respect context window limits.
  • Safety first: Sanitize external data before adding it to the context window to prevent sensitive information leakage.

Frequently Asked Questions

How does the 12-Factor Agents framework differ from standard RAG implementations?

Standard RAG often injects retrieved text directly into system prompts, obscuring the retrieval step from the agent's logic. The 12-Factor approach treats retrieval as an explicit event type within the Thread class, enabling the agent to decide when to retrieve, handle retrieval failures, and track exactly what knowledge influenced each decision according to content/factor-03-own-your-context-window.md.

What file contains the core logic for context window management?

The primary documentation and implementation guidance resides in content/factor-03-own-your-context-window.md, which details how RAG documents integrate with other context components like tool calls and memory. Supporting context appears in content/factor-02-own-your-prompts.md and content/factor-04-tools-are-structured-outputs.md.

Why use XML-style tags instead of standard JSON for retrieved documents?

XML-style tags created by event_to_prompt() provide clear visual delimiters that help LLMs distinguish between different context components (user messages, tool results, retrieved documents). This format reduces token overhead compared to repeated JSON schemas while maintaining human readability for debugging, as shown in the framework's code examples.

Can the agent perform multiple retrieval steps in one conversation turn?

Yes. The decision loop allows the agent to issue multiple retrieve_document events if the initial results prove insufficient. The thread_to_prompt() function aggregates all events chronologically, enabling iterative retrieval where each subsequent query can reference previous results before calling determine_next_step().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →