# How to Implement Session Memory with Agents SDK for Short-Term Memory Management

> Learn to implement session memory with the Agents SDK. Use custom TrimmingSession or SummarizingSession for effective short-term memory management and token limit control.

- Repository: [OpenAI/openai-cookbook](https://github.com/openai/openai-cookbook)
- Tags: how-to-guide
- Published: 2026-03-02

---

**Replace the default `Session` with a custom implementation of `SessionABC`—such as `TrimmingSession` to retain only the last N user turns or `SummarizingSession` to compress older context via LLM—to manage short-term memory while staying within token limits.**

The OpenAI Agents SDK automatically stores conversation history through its `Session` abstraction, but long-running chats quickly exceed model context windows. For robust short-term memory management, the openai/openai-cookbook repository demonstrates custom session classes that selectively retain or summarize historical interactions. These implementations conform to the `SessionABC` interface, allowing seamless integration with `Runner.run` and the Responses API.

## Understanding Session Memory Architecture

The Agents SDK defines a **turn** as a conversational unit starting with a real user message and including all subsequent items—assistant replies, tool calls, and tool results—until the next user message. The default session keeps every turn indefinitely, which becomes problematic in extended interactions.

Custom session implementations solve this by implementing `SessionABC` from `agents.memory.session` and overriding three core methods:
- `get_items()` – Retrieve the current conversation history
- `add_items()` – Append new messages to the session
- `clear_session()` – Reset the conversation state

Both strategies maintain **async-compatible** architectures using `asyncio.Lock()` to prevent race conditions during concurrent modifications.

## Implementing a Trimming Session

The **trimming session** provides deterministic, low-latency short-term memory by retaining only the most recent `max_turns` user turns. This approach requires no additional model calls, minimizing both cost and response time.

### The TrimmingSession Class

As implemented in `examples/agents_sdk/session_memory.ipynb`, the `TrimmingSession` class uses a `deque` to store items and trims history based on user message boundaries:

```python
from agents.memory.session import SessionABC
from agents.items import TResponseInputItem
from collections import deque
import asyncio

ROLE_USER = "user"

def _is_user_msg(item: TResponseInputItem) -> bool:
    """Detect a user-role message in the generic SDK item."""
    if isinstance(item, dict):
        if "role" in item:
            return item["role"] == ROLE_USER
        if item.get("type") == "message":
            return item.get("role") == ROLE_USER
    return getattr(item, "role", None) == ROLE_USER


class TrimmingSession(SessionABC):
    """Keep only the last `max_turns` user turns."""
    def __init__(self, session_id: str, max_turns: int = 8):
        self.session_id = session_id
        self.max_turns = max(1, int(max_turns))
        self._items: deque[TResponseInputItem] = deque()
        self._lock = asyncio.Lock()

    async def get_items(self, limit: int | None = None):
        async with self._lock:
            trimmed = self._trim_to_last_turns(list(self._items))
            return trimmed[-limit:] if limit is not None else trimmed

    async def add_items(self, items: list[TResponseInputItem]):
        if not items:
            return
        async with self._lock:
            self._items.extend(items)
            self._items = deque(self._trim_to_last_turns(list(self._items)))

    async def pop_item(self):
        async with self._lock:
            return self._items.pop() if self._items else None

    async def clear_session(self):
        async with self._lock:
            self._items.clear()

    def _trim_to_last_turns(self, items: list[TResponseInputItem]) -> list[TResponseInputItem]:
        """Return the suffix that contains the last `max_turns` user messages."""
        if not items:
            return items
        count = 0
        start_idx = 0
        for i in range(len(items) - 1, -1, -1):
            if _is_user_msg(items[i]):
                count += 1
                if count == self.max_turns:
                    start_idx = i
                    break
        return items[start_idx:]

```

The `_trim_to_last_turns` method iterates backward through the deque, counting user messages until it reaches `max_turns`, then returns the slice from that index forward. This preserves tool calls and assistant responses associated with the retained turns.

### Usage with Runner

To use the trimming session, instantiate it and pass it to `Runner.run`:

```python
session = TrimmingSession("my_session", max_turns=3)

# Run the agent with the custom session

result = await Runner.run(support_agent, "My laptop shows a red blinking light.", session=session)
history = await session.get_items()   # → only the last 3 user turns are kept

```

## Implementing a Summarizing Session

When applications require long-range context beyond the immediate turns, the **summarizing session** compresses older conversation into synthetic summaries. This approach triggers only when the user turn count exceeds a configurable `context_limit`.

### Core Architecture

The `SummarizingSession` class (cells 48-118 of `session_memory.ipynb`) maintains two categories of records:
1. **Verbatim turns** – The most recent `keep_last_n_turns` retained in full
2. **Synthetic summary** – A compressed representation of earlier conversation stored as a shadow user message and assistant response pair

Key parameters include:
- `context_limit` – Maximum real user turns before summarization triggers
- `keep_last_n_turns` – How many recent turns remain uncompressed (must be ≤ `context_limit`)

### Implementation Details

The class tracks metadata separately from messages to identify synthetic entries:

```python
class SummarizingSession:
    _ALLOWED_MSG_KEYS = {"role", "content", "name"}
    
    def __init__(
        self,
        keep_last_n_turns: int = 2,
        context_limit: int = 4,
        summarizer: Optional["Summarizer"] = None,
        session_id: Optional[str] = None,
    ):
        assert context_limit >= 1
        assert 0 <= keep_last_n_turns <= context_limit
        self.keep_last_n_turns = keep_last_n_turns
        self.context_limit = context_limit
        self.summarizer = summarizer
        self.session_id = session_id or "default"
        self._records: deque[Record] = deque()
        self._lock = asyncio.Lock()

```

The `add_items` method checks `self._summarize_decision_locked()` to determine if the conversation exceeds `context_limit`. When true, it:
1. Extracts the prefix (turns to be summarized)
2. Calls the summarizer outside the lock to prevent blocking
3. Atomically replaces the prefix with synthetic summary messages

### LLM Summarizer Integration

The `LLMSummarizer` class handles the compression logic:

```python
class LLMSummarizer:
    def __init__(self, client, model="gpt-4o", max_tokens=400, tool_trim_limit=600):
        self.client = client
        self.model = model
        self.max_tokens = max_tokens
        self.tool_trim_limit = tool_trim_limit

    async def summarize(self, messages: List[Dict[str, Any]]) -> Tuple[str, str]:
        user_shadow = "Summarize the conversation we had so far."
        TOOL_ROLES = {"tool", "tool_result"}

        def to_snippet(m: Dict[str, Any]) -> Optional[str]:
            role = (m.get("role") or "assistant").lower()
            content = (m.get("content") or "").strip()
            if not content:
                return None
            if role in TOOL_ROLES and len(content) > self.tool_trim_limit:
                content = content[: self.tool_trim_limit] + " …"
            return f"{role.upper()}: {content}"

        snippets = [s for m in messages if (s := to_snippet(m))]
        prompt = [
            {"role": "system", "content": SUMMARY_PROMPT},
            {"role": "user", "content": "\n".join(snippets)},
        ]

        resp = await asyncio.to_thread(
            self.client.responses.create,
            model=self.model,
            input=prompt,
            max_output_tokens=self.max_tokens,
        )
        return user_shadow, resp.output_text

```

This implementation truncates tool outputs exceeding `tool_trim_limit` characters to prevent token bloat from verbose API responses.

### Complete Usage Pattern

```python
summariser = LLMSummarizer(client)

session = SummarizingSession(
    keep_last_n_turns=2,
    context_limit=4,
    summarizer=summariser,
)

await session.add_items([{"role": "user", "content": "My router disconnects constantly."}])
await session.add_items([{"role": "assistant", "content": "Let’s check the firmware…"}])

# … keep adding items …

history = await session.get_items()   # model-ready list with synthetic summary + last 2 turns

```

## Selecting the Right Memory Strategy

Choose between these approaches based on your application's requirements:

- **TrimmingSession** – Use when you need deterministic latency, zero additional API costs, and only recent context matters (e.g., command-line tools, simple Q&A bots).
- **SummarizingSession** – Use when conversations require long-range context awareness and you can tolerate occasional summary latency for reduced token usage over time (e.g., customer support agents, long-form writing assistants).

## Summary

- The OpenAI Agents SDK uses `SessionABC` as the contract for conversation history management, allowing custom implementations to replace the default unlimited session.
- **TrimmingSession** provides cost-free short-term memory by keeping only the last N user turns, implemented in `examples/agents_sdk/session_memory.ipynb`.
- **SummarizingSession** compresses older conversation into synthetic summaries when user turns exceed `context_limit`, preserving long-range context without exceeding token windows.
- Both implementations use `asyncio.Lock()` for thread safety and integrate seamlessly with `Runner.run` and the Responses API.
- The `LLMSummarizer` class handles intelligent compression, truncating verbose tool outputs to maintain efficiency.

## Frequently Asked Questions

### How does the Agents SDK define a conversation turn?

A **turn** begins with a real user message and includes every subsequent item—assistant responses, tool calls, and tool results—until the next user message appears. Both `TrimmingSession` and `SummarizingSession` use this boundary definition to determine which content to retain or compress, checking for the `role` attribute or `synthetic` metadata flag to identify true user-initiated turns.

### Can I use these custom sessions with the Responses API directly?

Yes. Because both `TrimmingSession` and `SummarizingSession` implement `SessionABC`, they work with any SDK component accepting session objects. You can pass them to `Runner.run()` for agent execution or use `session.get_items()` to retrieve formatted message lists for direct insertion into `client.responses.create()` calls.

### What happens if I set `keep_last_n_turns` equal to `context_limit` in SummarizingSession?

When `keep_last_n_turns` equals `context_limit`, the session never triggers summarization because the number of allowed turns equals the retention threshold. The `_summarize_decision_locked()` method returns `False` when `len(user_starts) <= self.context_limit`, effectively making the session behave like a standard unlimited session until you adjust the parameters or implement additional logic.

### Do these session implementations maintain thread safety in multi-agent workflows?

Yes. Both implementations use `asyncio.Lock()` to protect the internal `deque` storage during concurrent access. This is essential for multi-agent collaboration scenarios where multiple agents might call `add_items()` or `get_items()` simultaneously, as demonstrated in `examples/agents_sdk/multi-agent-portfolio-collaboration/multi_agent_portfolio_collaboration.ipynb`.