Benchmarking Metrics for OAMP vs Naive Flat-History Memory: A Technical Comparison
The Oracle AI Developer Hub benchmark notebook evaluates Oracle Agent Memory (OAMP) against naive flat-history memory using three primary metrics—token consumption, wall-clock latency, and response quality—measured continuously across an 80-turn scripted conversation.
According to the oracle-devrel/oracle-ai-developer-hub repository, the oracle_agent_memory_benchmarks.ipynb notebook provides a quantitative comparison between intelligent memory extraction and simple message appending. The benchmark runs both approaches through an identical 80-turn scripted dialogue, capturing performance data that reveals the practical trade-offs between retrieval-augmented context and growing flat history.
The Three Core Benchmarking Metrics
The benchmark evaluates memory strategies along three distinct axes, each recorded at every turn to build comparative time-series data.
Token Consumption (Cost)
Token consumption measures the number of tokens sent to the LLM per turn, serving as a direct proxy for API costs. In notebooks/agent_memory/oracle_agent_memory_benchmarks.ipynb, the estimate_tokens function approximates token count by dividing the character length of the serialized messages by 4 (lines 85-89).
The notebook stores historical data in dedicated lists:
oamp_token_history– Tracks tokens for the OAMP agent using extracted context cardsnaive_token_history– Tracks tokens for the baseline agent appending full message history
Wall-Clock Latency (Speed)
Wall-clock latency captures two timing measurements: retrieval latency (time to prepare the context) and total end-to-end latency per turn.
For the OAMP implementation, retrieval latency is calculated as t_context_built - t_start, measuring the time required to fetch the relevant context card (lines 33-34). Total latency (t_end - t_start) includes the LLM inference time and is stored in oamp_total_latency (lines 34-35).
The naive agent uses identical measurement patterns (naive_retrieval_latency and naive_total_latency), though its retrieval latency is essentially zero since it simply appends messages to a growing list without semantic search or extraction.
Response Quality (Accuracy)
Response quality is assessed via LLM-as-a-judge methodology. After each turn, both agents' replies are captured in oamp_responses and naive_responses lists (lines 100-103).
A separate evaluation layer (implemented in subsequent cells) compares paired responses to compute win-loss tallies, determining whether the OAMP agent's selective memory retrieval maintains or improves answer accuracy compared to the naive approach of sending complete conversation history.
How Metrics Are Implemented in Code
The benchmark uses precise instrumentation to ensure comparable measurements across both memory strategies.
Token Estimation Logic
The estimate_tokens function provides a lightweight approximation used throughout the benchmark:
import json as _json_lib
def estimate_tokens(messages: list) -> int:
"""Approximate token count as characters/4 (same approach used in the benchmark)."""
return len(_json_lib.dumps(messages)) // 4
# Usage during benchmark execution:
oamp_tokens = estimate_tokens(oamp_messages)
naive_tokens = estimate_tokens(naive_messages)
Latency Instrumentation
Wall-clock timing uses time.perf_counter() for high-precision measurements. For the OAMP agent, the benchmark distinguishes between context preparation and total turn time:
import time
def call_oamp_agent(user_query: str) -> str:
t_start = time.perf_counter()
thread.add_messages([Message(role="user", content=user_query)])
context_card = thread.get_context_card() or "(no prior context)"
t_context_built = time.perf_counter()
# Build prompt with retrieved context
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Relevant memory:\n{context_card}\n\nCurrent question: {user_query}"},
]
oamp_token_history.append(estimate_tokens(messages))
# LLM call and response handling
response = openai_client.chat.completions.create(
model="gpt-5.4",
messages=messages,
)
answer = response.choices[0].message.content or ""
thread.add_messages([Message(role="assistant", content=answer)])
t_end = time.perf_counter()
oamp_retrieval_latency.append(t_context_built - t_start)
oamp_total_latency.append(t_end - t_start)
oamp_responses.append(answer)
return answer
Comparing OAMP vs Naive Implementation
The naive flat-history implementation serves as the baseline, demonstrating the performance characteristics of unbounded context growth:
def call_naive_agent(user_query: str) -> str:
t_start = time.perf_counter()
naive_messages.append({"role": "user", "content": user_query})
t_context_built = time.perf_counter() # No retrieval overhead
naive_token_history.append(estimate_tokens(naive_messages))
response = openai_client.chat.completions.create(
model="gpt-5.4",
messages=naive_messages, # Growing list of all prior messages
)
answer = response.choices[0].message.content or ""
naive_messages.append({"role": "assistant", "content": answer})
t_end = time.perf_counter()
naive_retrieval_latency.append(t_context_built - t_start) # Near zero
naive_total_latency.append(t_end - t_start)
naive_responses.append(answer)
return answer
Key differences revealed by the benchmarking metrics include:
- OAMP incurs retrieval latency (
t_context_built - t_start) but maintains bounded token counts by extracting only relevant memories - Naive approach shows zero retrieval overhead but linearly increasing token consumption as
naive_messagesgrows with each turn - Both approaches are measured against identical 80-turn conversation scripts to ensure fair comparison
Summary
- Token consumption is approximated using
estimate_tokens(characters ÷ 4) and tracked separately inoamp_token_historyandnaive_token_historyto compare API costs. - Wall-clock latency distinguishes between retrieval time (
oamp_retrieval_latency) and total turn time (oamp_total_latency), revealing the overhead of semantic memory extraction versus flat list appending. - Response quality is captured in
oamp_responsesandnaive_responseslists for subsequent LLM-as-a-judge evaluation to determine accuracy trade-offs. - All metrics are collected in
oracle_agent_memory_benchmarks.ipynbusing high-precisiontime.perf_counter()measurements across an 80-turn standardized conversation.
Frequently Asked Questions
What file contains the OAMP benchmarking implementation?
The complete benchmark implementation resides in notebooks/agent_memory/oracle_agent_memory_benchmarks.ipynb within the oracle-devrel/oracle-ai-developer-hub repository. This Jupyter notebook contains the 80-turn test script, metric collection logic, and comparative visualization code.
How does the benchmark measure token consumption without an official tokenizer?
The notebook uses a lightweight approximation via the estimate_tokens function, which calculates token count as the length of the serialized messages divided by 4 (lines 85-89). This provides a consistent relative comparison between OAMP and naive approaches without requiring model-specific tokenizers.
Does OAMP add significant latency compared to naive flat-history?
According to the instrumentation in lines 33-35, OAMP adds measurable retrieval latency (t_context_built - t_start) for fetching context cards, while the naive approach records near-zero retrieval time. However, as conversation length grows, the naive approach may exhibit higher total latency due to increasing token processing, which the oamp_total_latency and naive_total_latency metrics capture end-to-end.
How is response quality evaluated when there is no single correct answer?
The benchmark implements an LLM-as-a-judge pattern where both agents' responses (stored in oamp_responses and naive_responses) are evaluated by a separate LLM instance after the conversation completes. This secondary model assesses which response better addresses the user's query given the conversation context, generating win-loss statistics that quantify accuracy trade-offs between selective memory retrieval and full history context.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →