# How to Design Agent Memory and Context Persistence in 12-Factor Agents

> Learn to design agent memory and context persistence by treating threads as the source of truth and persisting events to a ThreadStore. Ensure seamless state retention for your stateless LLMs.

- Repository: [HumanLayer/12-factor-agents](https://github.com/humanlayer/12-factor-agents)
- Tags: how-to-guide
- Published: 2026-05-19

---

**TLDR:** You design agent memory by treating the **Thread** as the canonical source of truth—persisting historical events to a **ThreadStore** (such as **FileSystemThreadStore**) and re-injecting the serialized context into the LLM's context window on every turn, since LLMs are stateless functions that cannot retain state across invocations.

Designing robust agent memory and context persistence in the `humanlayer/12-factor-agents` framework requires treating memory as explicit historical data rather than implicit state. Because LLMs are stateless functions, the only way an agent retains context across turns is by persisting **Thread** objects containing past prompts, tool calls, and results, then deliberately shaping what gets injected into the next request's context window.

## Core Architectural Concepts for Agent Memory

### Thread as the Canonical Source of Truth

A **Thread** is an ordered list of events—including messages, tool calls, results, and errors—that represents a single conversation or workflow. Defined in [`packages/create-12-factor-agent/template/src/agent.ts`](https://github.com/humanlayer/12-factor-agents/blob/main/packages/create-12-factor-agent/template/src/agent.ts), the Thread is the canonical source of truth for everything the LLM sees. It can be serialized to a string for the LLM and deserialized back into objects for the runtime.

### ThreadStore Interface

The **ThreadStore** is an abstract interface defined in [`packages/create-12-factor-agent/template/src/state.ts`](https://github.com/humanlayer/12-factor-agents/blob/main/packages/create-12-factor-agent/template/src/state.ts) that decouples the agent from specific storage backends. It specifies three core methods:

- `create(thread)` – Initializes and persists a new thread.
- `get(id)` – Retrieves a thread by its unique identifier.
- `update(id, thread)` – Writes the latest state after each step.

This abstraction allows you to swap between simple in-memory stores, filesystems, Redis, SQLite, or Postgres without changing agent logic.

### FileSystemThreadStore Implementation

**FileSystemThreadStore** is a concrete implementation that demonstrates minimal persistence requirements for production-ready agents. When you call `create(thread)`, it generates a UUID and writes two files under `.threads/`:

- [`thread.json`](https://github.com/humanlayer/12-factor-agents/blob/main/thread.json) – Contains full event objects for runtime deserialization.
- [`thread.txt`](https://github.com/humanlayer/12-factor-agents/blob/main/thread.txt) – Contains a human-readable, plain-text version for LLM consumption and debugging.

The `update(id, thread)` method rewrites both files after each LLM step, ensuring the latest context survives process restarts or crashes.

### Context Window Ownership

The **Context Window** comprises all information sent to the LLM on each turn, including system instructions, historic events, and memory. According to [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md), properly shaping this context directly determines token efficiency, response quality, and safety. The `Thread.serializeForLLM()` method converts the internal event list into a custom XML-style format, giving you full control over token density and ordering before injection.

### Memory Management Strategies

As processes grow longer, Threads risk exceeding model token limits. [`content/factor-08-own-your-control-flow.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-08-own-your-control-flow.md) outlines **Memory Management** strategies for pruning, compressing, or splitting context, including techniques like compaction, summarization, or selective retrieval. These strategies enable agents to run for many steps without running out of tokens.

## How Thread Persistence Works

1. **Create a thread** – When a conversation starts, `FileSystemThreadStore.create(thread)` generates a UUID, writes [`thread.json`](https://github.com/humanlayer/12-factor-agents/blob/main/thread.json) (full event objects), and writes [`thread.txt`](https://github.com/humanlayer/12-factor-agents/blob/main/thread.txt) (human-readable format) under `.threads/`.

2. **Read a thread** – Subsequent calls use `get(id)` to retrieve the JSON file, deserialize it back into a Thread object, and resume processing exactly where the agent left off.

3. **Update after each step** – After the LLM responds and the agent adds new events (tool calls or results), `update(id, thread)` rewrites both files. This guarantees that the latest context is always persisted to disk.

## Implementing Context Persistence in Node.js

Below is a minimal TypeScript implementation showing an agent creating a thread, adding an event, persisting it, and feeding the stored context back to the LLM.

```typescript
import { Thread } from './src/agent';
import { FileSystemThreadStore } from './src/state';
import { llm } from './src/llm'; // your LLM wrapper

async function run() {
  const store = new FileSystemThreadStore();

  // 1️⃣ Start a new thread with the initial user request
  const thread = new Thread([
    { type: 'user_message', data: 'Deploy the latest backend?' },
  ]);
  const id = await store.create(thread);

  // 2️⃣ Ask the LLM what to do next
  const prompt = thread.serializeForLLM();
  const nextStep = await llm.determineNextStep(prompt); // custom call
  thread.events.push(nextStep); // event is a tool call or action

  // 3️⃣ Persist the updated thread
  await store.update(id, thread);

  // 4️⃣ Later – recover the thread and continue
  const resumed = await store.get(id);
  if (resumed) {
    const nextPrompt = resumed.serializeForLLM();
    const followUp = await llm.determineNextStep(nextPrompt);
    console.log('Next step:', followUp);
  }
}
run();

```

Key implementation details from the source code:

- The thread is the *single source of truth* for everything the LLM sees.
- `FileSystemThreadStore` guarantees durability across process restarts by writing to the filesystem.
- The custom XML-style format produced by `serializeForLLM()` lets you pack dense, token-efficient context.

## Summary

- **Threads are immutable history**: In [`packages/create-12-factor-agent/template/src/agent.ts`](https://github.com/humanlayer/12-factor-agents/blob/main/packages/create-12-factor-agent/template/src/agent.ts), the Thread class maintains an ordered list of events that constitutes the agent's memory.
- **Explicit persistence required**: Because LLMs are stateless, you must use a ThreadStore implementation like `FileSystemThreadStore` (defined in [`packages/create-12-factor-agent/template/src/state.ts`](https://github.com/humanlayer/12-factor-agents/blob/main/packages/create-12-factor-agent/template/src/state.ts)) to write threads to disk via `create()` and `update()`.
- **Control the context window**: Use `serializeForLLM()` to convert threads into token-efficient formats before sending to the model, as emphasized in Factor 3 ([`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md)).
- **Plan for scale**: Implement memory management strategies such as compaction or summarization (Factor 8) to handle long-running conversations without exceeding token limits.

## Frequently Asked Questions

### What constitutes agent memory in 12-Factor Agents?

Agent memory includes any historical data the LLM can draw upon when making decisions, such as past prompts, tool calls, RAG documents, and results from earlier steps. In the 12-Factor framework, this data is explicitly stored as structured events within a Thread object, not as implicit state in the model.

### How does FileSystemThreadStore persist thread data?

`FileSystemThreadStore` writes each thread to two files under the `.threads/` directory: a [`thread.json`](https://github.com/humanlayer/12-factor-agents/blob/main/thread.json) containing serialized event objects for the runtime to deserialize, and a [`thread.txt`](https://github.com/humanlayer/12-factor-agents/blob/main/thread.txt) containing a human-readable format suitable for debugging or direct LLM consumption. This dual-format approach separates machine-readable state from human-readable logs.

### Why must agents re-inject memory on every LLM turn?

LLMs are stateless functions; they do not maintain internal state between requests. The only mechanism for retaining context is to persist the Thread to a store after each turn, then retrieve it and call `serializeForLLM()` to inject the serialized historical context into the next request's context window.

### How do you prevent token overflow with long-running threads?

As outlined in [`content/factor-08-own-your-control-flow.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-08-own-your-control-flow.md), you implement **memory management** strategies such as context compaction, summarization of older events, or selective retrieval of relevant history. These techniques prune the Thread before serialization, ensuring the context window stays within the specific model's token limits.