How to Implement Session Memory with Agents SDK for Short-Term Memory Management
Replace the default Session with a custom implementation of SessionABC—such as TrimmingSession to retain only the last N user turns or SummarizingSession to compress older context via LLM—to manage short-term memory while staying within token limits.
The OpenAI Agents SDK automatically stores conversation history through its Session abstraction, but long-running chats quickly exceed model context windows. For robust short-term memory management, the openai/openai-cookbook repository demonstrates custom session classes that selectively retain or summarize historical interactions. These implementations conform to the SessionABC interface, allowing seamless integration with Runner.run and the Responses API.
Understanding Session Memory Architecture
The Agents SDK defines a turn as a conversational unit starting with a real user message and including all subsequent items—assistant replies, tool calls, and tool results—until the next user message. The default session keeps every turn indefinitely, which becomes problematic in extended interactions.
Custom session implementations solve this by implementing SessionABC from agents.memory.session and overriding three core methods:
get_items()– Retrieve the current conversation historyadd_items()– Append new messages to the sessionclear_session()– Reset the conversation state
Both strategies maintain async-compatible architectures using asyncio.Lock() to prevent race conditions during concurrent modifications.
Implementing a Trimming Session
The trimming session provides deterministic, low-latency short-term memory by retaining only the most recent max_turns user turns. This approach requires no additional model calls, minimizing both cost and response time.
The TrimmingSession Class
As implemented in examples/agents_sdk/session_memory.ipynb, the TrimmingSession class uses a deque to store items and trims history based on user message boundaries:
from agents.memory.session import SessionABC
from agents.items import TResponseInputItem
from collections import deque
import asyncio
ROLE_USER = "user"
def _is_user_msg(item: TResponseInputItem) -> bool:
"""Detect a user-role message in the generic SDK item."""
if isinstance(item, dict):
if "role" in item:
return item["role"] == ROLE_USER
if item.get("type") == "message":
return item.get("role") == ROLE_USER
return getattr(item, "role", None) == ROLE_USER
class TrimmingSession(SessionABC):
"""Keep only the last `max_turns` user turns."""
def __init__(self, session_id: str, max_turns: int = 8):
self.session_id = session_id
self.max_turns = max(1, int(max_turns))
self._items: deque[TResponseInputItem] = deque()
self._lock = asyncio.Lock()
async def get_items(self, limit: int | None = None):
async with self._lock:
trimmed = self._trim_to_last_turns(list(self._items))
return trimmed[-limit:] if limit is not None else trimmed
async def add_items(self, items: list[TResponseInputItem]):
if not items:
return
async with self._lock:
self._items.extend(items)
self._items = deque(self._trim_to_last_turns(list(self._items)))
async def pop_item(self):
async with self._lock:
return self._items.pop() if self._items else None
async def clear_session(self):
async with self._lock:
self._items.clear()
def _trim_to_last_turns(self, items: list[TResponseInputItem]) -> list[TResponseInputItem]:
"""Return the suffix that contains the last `max_turns` user messages."""
if not items:
return items
count = 0
start_idx = 0
for i in range(len(items) - 1, -1, -1):
if _is_user_msg(items[i]):
count += 1
if count == self.max_turns:
start_idx = i
break
return items[start_idx:]
The _trim_to_last_turns method iterates backward through the deque, counting user messages until it reaches max_turns, then returns the slice from that index forward. This preserves tool calls and assistant responses associated with the retained turns.
Usage with Runner
To use the trimming session, instantiate it and pass it to Runner.run:
session = TrimmingSession("my_session", max_turns=3)
# Run the agent with the custom session
result = await Runner.run(support_agent, "My laptop shows a red blinking light.", session=session)
history = await session.get_items() # → only the last 3 user turns are kept
Implementing a Summarizing Session
When applications require long-range context beyond the immediate turns, the summarizing session compresses older conversation into synthetic summaries. This approach triggers only when the user turn count exceeds a configurable context_limit.
Core Architecture
The SummarizingSession class (cells 48-118 of session_memory.ipynb) maintains two categories of records:
- Verbatim turns – The most recent
keep_last_n_turnsretained in full - Synthetic summary – A compressed representation of earlier conversation stored as a shadow user message and assistant response pair
Key parameters include:
context_limit– Maximum real user turns before summarization triggerskeep_last_n_turns– How many recent turns remain uncompressed (must be ≤context_limit)
Implementation Details
The class tracks metadata separately from messages to identify synthetic entries:
class SummarizingSession:
_ALLOWED_MSG_KEYS = {"role", "content", "name"}
def __init__(
self,
keep_last_n_turns: int = 2,
context_limit: int = 4,
summarizer: Optional["Summarizer"] = None,
session_id: Optional[str] = None,
):
assert context_limit >= 1
assert 0 <= keep_last_n_turns <= context_limit
self.keep_last_n_turns = keep_last_n_turns
self.context_limit = context_limit
self.summarizer = summarizer
self.session_id = session_id or "default"
self._records: deque[Record] = deque()
self._lock = asyncio.Lock()
The add_items method checks self._summarize_decision_locked() to determine if the conversation exceeds context_limit. When true, it:
- Extracts the prefix (turns to be summarized)
- Calls the summarizer outside the lock to prevent blocking
- Atomically replaces the prefix with synthetic summary messages
LLM Summarizer Integration
The LLMSummarizer class handles the compression logic:
class LLMSummarizer:
def __init__(self, client, model="gpt-4o", max_tokens=400, tool_trim_limit=600):
self.client = client
self.model = model
self.max_tokens = max_tokens
self.tool_trim_limit = tool_trim_limit
async def summarize(self, messages: List[Dict[str, Any]]) -> Tuple[str, str]:
user_shadow = "Summarize the conversation we had so far."
TOOL_ROLES = {"tool", "tool_result"}
def to_snippet(m: Dict[str, Any]) -> Optional[str]:
role = (m.get("role") or "assistant").lower()
content = (m.get("content") or "").strip()
if not content:
return None
if role in TOOL_ROLES and len(content) > self.tool_trim_limit:
content = content[: self.tool_trim_limit] + " …"
return f"{role.upper()}: {content}"
snippets = [s for m in messages if (s := to_snippet(m))]
prompt = [
{"role": "system", "content": SUMMARY_PROMPT},
{"role": "user", "content": "\n".join(snippets)},
]
resp = await asyncio.to_thread(
self.client.responses.create,
model=self.model,
input=prompt,
max_output_tokens=self.max_tokens,
)
return user_shadow, resp.output_text
This implementation truncates tool outputs exceeding tool_trim_limit characters to prevent token bloat from verbose API responses.
Complete Usage Pattern
summariser = LLMSummarizer(client)
session = SummarizingSession(
keep_last_n_turns=2,
context_limit=4,
summarizer=summariser,
)
await session.add_items([{"role": "user", "content": "My router disconnects constantly."}])
await session.add_items([{"role": "assistant", "content": "Let’s check the firmware…"}])
# … keep adding items …
history = await session.get_items() # model-ready list with synthetic summary + last 2 turns
Selecting the Right Memory Strategy
Choose between these approaches based on your application's requirements:
- TrimmingSession – Use when you need deterministic latency, zero additional API costs, and only recent context matters (e.g., command-line tools, simple Q&A bots).
- SummarizingSession – Use when conversations require long-range context awareness and you can tolerate occasional summary latency for reduced token usage over time (e.g., customer support agents, long-form writing assistants).
Summary
- The OpenAI Agents SDK uses
SessionABCas the contract for conversation history management, allowing custom implementations to replace the default unlimited session. - TrimmingSession provides cost-free short-term memory by keeping only the last N user turns, implemented in
examples/agents_sdk/session_memory.ipynb. - SummarizingSession compresses older conversation into synthetic summaries when user turns exceed
context_limit, preserving long-range context without exceeding token windows. - Both implementations use
asyncio.Lock()for thread safety and integrate seamlessly withRunner.runand the Responses API. - The
LLMSummarizerclass handles intelligent compression, truncating verbose tool outputs to maintain efficiency.
Frequently Asked Questions
How does the Agents SDK define a conversation turn?
A turn begins with a real user message and includes every subsequent item—assistant responses, tool calls, and tool results—until the next user message appears. Both TrimmingSession and SummarizingSession use this boundary definition to determine which content to retain or compress, checking for the role attribute or synthetic metadata flag to identify true user-initiated turns.
Can I use these custom sessions with the Responses API directly?
Yes. Because both TrimmingSession and SummarizingSession implement SessionABC, they work with any SDK component accepting session objects. You can pass them to Runner.run() for agent execution or use session.get_items() to retrieve formatted message lists for direct insertion into client.responses.create() calls.
What happens if I set keep_last_n_turns equal to context_limit in SummarizingSession?
When keep_last_n_turns equals context_limit, the session never triggers summarization because the number of allowed turns equals the retention threshold. The _summarize_decision_locked() method returns False when len(user_starts) <= self.context_limit, effectively making the session behave like a standard unlimited session until you adjust the parameters or implement additional logic.
Do these session implementations maintain thread safety in multi-agent workflows?
Yes. Both implementations use asyncio.Lock() to protect the internal deque storage during concurrent access. This is essential for multi-agent collaboration scenarios where multiple agents might call add_items() or get_items() simultaneously, as demonstrated in examples/agents_sdk/multi-agent-portfolio-collaboration/multi_agent_portfolio_collaboration.ipynb.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →