How to Implement Session Memory with Agents SDK for Short-Term Memory Management

Replace the default Session with a custom implementation of SessionABC—such as TrimmingSession to retain only the last N user turns or SummarizingSession to compress older context via LLM—to manage short-term memory while staying within token limits.

The OpenAI Agents SDK automatically stores conversation history through its Session abstraction, but long-running chats quickly exceed model context windows. For robust short-term memory management, the openai/openai-cookbook repository demonstrates custom session classes that selectively retain or summarize historical interactions. These implementations conform to the SessionABC interface, allowing seamless integration with Runner.run and the Responses API.

Understanding Session Memory Architecture

The Agents SDK defines a turn as a conversational unit starting with a real user message and including all subsequent items—assistant replies, tool calls, and tool results—until the next user message. The default session keeps every turn indefinitely, which becomes problematic in extended interactions.

Custom session implementations solve this by implementing SessionABC from agents.memory.session and overriding three core methods:

  • get_items() – Retrieve the current conversation history
  • add_items() – Append new messages to the session
  • clear_session() – Reset the conversation state

Both strategies maintain async-compatible architectures using asyncio.Lock() to prevent race conditions during concurrent modifications.

Implementing a Trimming Session

The trimming session provides deterministic, low-latency short-term memory by retaining only the most recent max_turns user turns. This approach requires no additional model calls, minimizing both cost and response time.

The TrimmingSession Class

As implemented in examples/agents_sdk/session_memory.ipynb, the TrimmingSession class uses a deque to store items and trims history based on user message boundaries:

from agents.memory.session import SessionABC
from agents.items import TResponseInputItem
from collections import deque
import asyncio

ROLE_USER = "user"

def _is_user_msg(item: TResponseInputItem) -> bool:
    """Detect a user-role message in the generic SDK item."""
    if isinstance(item, dict):
        if "role" in item:
            return item["role"] == ROLE_USER
        if item.get("type") == "message":
            return item.get("role") == ROLE_USER
    return getattr(item, "role", None) == ROLE_USER


class TrimmingSession(SessionABC):
    """Keep only the last `max_turns` user turns."""
    def __init__(self, session_id: str, max_turns: int = 8):
        self.session_id = session_id
        self.max_turns = max(1, int(max_turns))
        self._items: deque[TResponseInputItem] = deque()
        self._lock = asyncio.Lock()

    async def get_items(self, limit: int | None = None):
        async with self._lock:
            trimmed = self._trim_to_last_turns(list(self._items))
            return trimmed[-limit:] if limit is not None else trimmed

    async def add_items(self, items: list[TResponseInputItem]):
        if not items:
            return
        async with self._lock:
            self._items.extend(items)
            self._items = deque(self._trim_to_last_turns(list(self._items)))

    async def pop_item(self):
        async with self._lock:
            return self._items.pop() if self._items else None

    async def clear_session(self):
        async with self._lock:
            self._items.clear()

    def _trim_to_last_turns(self, items: list[TResponseInputItem]) -> list[TResponseInputItem]:
        """Return the suffix that contains the last `max_turns` user messages."""
        if not items:
            return items
        count = 0
        start_idx = 0
        for i in range(len(items) - 1, -1, -1):
            if _is_user_msg(items[i]):
                count += 1
                if count == self.max_turns:
                    start_idx = i
                    break
        return items[start_idx:]

The _trim_to_last_turns method iterates backward through the deque, counting user messages until it reaches max_turns, then returns the slice from that index forward. This preserves tool calls and assistant responses associated with the retained turns.

Usage with Runner

To use the trimming session, instantiate it and pass it to Runner.run:

session = TrimmingSession("my_session", max_turns=3)

# Run the agent with the custom session

result = await Runner.run(support_agent, "My laptop shows a red blinking light.", session=session)
history = await session.get_items()   # → only the last 3 user turns are kept

Implementing a Summarizing Session

When applications require long-range context beyond the immediate turns, the summarizing session compresses older conversation into synthetic summaries. This approach triggers only when the user turn count exceeds a configurable context_limit.

Core Architecture

The SummarizingSession class (cells 48-118 of session_memory.ipynb) maintains two categories of records:

  1. Verbatim turns – The most recent keep_last_n_turns retained in full
  2. Synthetic summary – A compressed representation of earlier conversation stored as a shadow user message and assistant response pair

Key parameters include:

  • context_limit – Maximum real user turns before summarization triggers
  • keep_last_n_turns – How many recent turns remain uncompressed (must be ≤ context_limit)

Implementation Details

The class tracks metadata separately from messages to identify synthetic entries:

class SummarizingSession:
    _ALLOWED_MSG_KEYS = {"role", "content", "name"}
    
    def __init__(
        self,
        keep_last_n_turns: int = 2,
        context_limit: int = 4,
        summarizer: Optional["Summarizer"] = None,
        session_id: Optional[str] = None,
    ):
        assert context_limit >= 1
        assert 0 <= keep_last_n_turns <= context_limit
        self.keep_last_n_turns = keep_last_n_turns
        self.context_limit = context_limit
        self.summarizer = summarizer
        self.session_id = session_id or "default"
        self._records: deque[Record] = deque()
        self._lock = asyncio.Lock()

The add_items method checks self._summarize_decision_locked() to determine if the conversation exceeds context_limit. When true, it:

  1. Extracts the prefix (turns to be summarized)
  2. Calls the summarizer outside the lock to prevent blocking
  3. Atomically replaces the prefix with synthetic summary messages

LLM Summarizer Integration

The LLMSummarizer class handles the compression logic:

class LLMSummarizer:
    def __init__(self, client, model="gpt-4o", max_tokens=400, tool_trim_limit=600):
        self.client = client
        self.model = model
        self.max_tokens = max_tokens
        self.tool_trim_limit = tool_trim_limit

    async def summarize(self, messages: List[Dict[str, Any]]) -> Tuple[str, str]:
        user_shadow = "Summarize the conversation we had so far."
        TOOL_ROLES = {"tool", "tool_result"}

        def to_snippet(m: Dict[str, Any]) -> Optional[str]:
            role = (m.get("role") or "assistant").lower()
            content = (m.get("content") or "").strip()
            if not content:
                return None
            if role in TOOL_ROLES and len(content) > self.tool_trim_limit:
                content = content[: self.tool_trim_limit] + " …"
            return f"{role.upper()}: {content}"

        snippets = [s for m in messages if (s := to_snippet(m))]
        prompt = [
            {"role": "system", "content": SUMMARY_PROMPT},
            {"role": "user", "content": "\n".join(snippets)},
        ]

        resp = await asyncio.to_thread(
            self.client.responses.create,
            model=self.model,
            input=prompt,
            max_output_tokens=self.max_tokens,
        )
        return user_shadow, resp.output_text

This implementation truncates tool outputs exceeding tool_trim_limit characters to prevent token bloat from verbose API responses.

Complete Usage Pattern

summariser = LLMSummarizer(client)

session = SummarizingSession(
    keep_last_n_turns=2,
    context_limit=4,
    summarizer=summariser,
)

await session.add_items([{"role": "user", "content": "My router disconnects constantly."}])
await session.add_items([{"role": "assistant", "content": "Let’s check the firmware…"}])

# … keep adding items …

history = await session.get_items()   # model-ready list with synthetic summary + last 2 turns

Selecting the Right Memory Strategy

Choose between these approaches based on your application's requirements:

  • TrimmingSession – Use when you need deterministic latency, zero additional API costs, and only recent context matters (e.g., command-line tools, simple Q&A bots).
  • SummarizingSession – Use when conversations require long-range context awareness and you can tolerate occasional summary latency for reduced token usage over time (e.g., customer support agents, long-form writing assistants).

Summary

  • The OpenAI Agents SDK uses SessionABC as the contract for conversation history management, allowing custom implementations to replace the default unlimited session.
  • TrimmingSession provides cost-free short-term memory by keeping only the last N user turns, implemented in examples/agents_sdk/session_memory.ipynb.
  • SummarizingSession compresses older conversation into synthetic summaries when user turns exceed context_limit, preserving long-range context without exceeding token windows.
  • Both implementations use asyncio.Lock() for thread safety and integrate seamlessly with Runner.run and the Responses API.
  • The LLMSummarizer class handles intelligent compression, truncating verbose tool outputs to maintain efficiency.

Frequently Asked Questions

How does the Agents SDK define a conversation turn?

A turn begins with a real user message and includes every subsequent item—assistant responses, tool calls, and tool results—until the next user message appears. Both TrimmingSession and SummarizingSession use this boundary definition to determine which content to retain or compress, checking for the role attribute or synthetic metadata flag to identify true user-initiated turns.

Can I use these custom sessions with the Responses API directly?

Yes. Because both TrimmingSession and SummarizingSession implement SessionABC, they work with any SDK component accepting session objects. You can pass them to Runner.run() for agent execution or use session.get_items() to retrieve formatted message lists for direct insertion into client.responses.create() calls.

What happens if I set keep_last_n_turns equal to context_limit in SummarizingSession?

When keep_last_n_turns equals context_limit, the session never triggers summarization because the number of allowed turns equals the retention threshold. The _summarize_decision_locked() method returns False when len(user_starts) <= self.context_limit, effectively making the session behave like a standard unlimited session until you adjust the parameters or implement additional logic.

Do these session implementations maintain thread safety in multi-agent workflows?

Yes. Both implementations use asyncio.Lock() to protect the internal deque storage during concurrent access. This is essential for multi-agent collaboration scenarios where multiple agents might call add_items() or get_items() simultaneously, as demonstrated in examples/agents_sdk/multi-agent-portfolio-collaboration/multi_agent_portfolio_collaboration.ipynb.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →