# Benchmarking Metrics for OAMP vs Naive Flat-History Memory: A Technical Comparison

> Discover OAMP vs naive flat-history memory benchmarks. Learn key metrics like token consumption, latency, and response quality for optimized AI performance.

- Repository: [Oracle Developers/oracle-ai-developer-hub](https://github.com/oracle-devrel/oracle-ai-developer-hub)
- Tags: deep-dive
- Published: 2026-05-10

---

**The Oracle AI Developer Hub benchmark notebook evaluates Oracle Agent Memory (OAMP) against naive flat-history memory using three primary metrics—token consumption, wall-clock latency, and response quality—measured continuously across an 80-turn scripted conversation.**

According to the `oracle-devrel/oracle-ai-developer-hub` repository, the `oracle_agent_memory_benchmarks.ipynb` notebook provides a quantitative comparison between intelligent memory extraction and simple message appending. The benchmark runs both approaches through an identical 80-turn scripted dialogue, capturing performance data that reveals the practical trade-offs between retrieval-augmented context and growing flat history.

## The Three Core Benchmarking Metrics

The benchmark evaluates memory strategies along three distinct axes, each recorded at every turn to build comparative time-series data.

### Token Consumption (Cost)

**Token consumption** measures the number of tokens sent to the LLM per turn, serving as a direct proxy for API costs. In `notebooks/agent_memory/oracle_agent_memory_benchmarks.ipynb`, the `estimate_tokens` function approximates token count by dividing the character length of the serialized messages by 4 ([lines 85-89](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/notebooks/agent_memory/oracle_agent_memory_benchmarks.ipynb#L85-L89)).

The notebook stores historical data in dedicated lists:

- `oamp_token_history` – Tracks tokens for the OAMP agent using extracted context cards
- `naive_token_history` – Tracks tokens for the baseline agent appending full message history

### Wall-Clock Latency (Speed)

**Wall-clock latency** captures two timing measurements: retrieval latency (time to prepare the context) and total end-to-end latency per turn.

For the OAMP implementation, retrieval latency is calculated as `t_context_built - t_start`, measuring the time required to fetch the relevant context card ([lines 33-34](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/notebooks/agent_memory/oracle_agent_memory_benchmarks.ipynb#L33-L34)). Total latency (`t_end - t_start`) includes the LLM inference time and is stored in `oamp_total_latency` ([lines 34-35](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/notebooks/agent_memory/oracle_agent_memory_benchmarks.ipynb#L34-L35)).

The naive agent uses identical measurement patterns (`naive_retrieval_latency` and `naive_total_latency`), though its retrieval latency is essentially zero since it simply appends messages to a growing list without semantic search or extraction.

### Response Quality (Accuracy)

**Response quality** is assessed via LLM-as-a-judge methodology. After each turn, both agents' replies are captured in `oamp_responses` and `naive_responses` lists ([lines 100-103](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/notebooks/agent_memory/oracle_agent_memory_benchmarks.ipynb#L100-L103)).

A separate evaluation layer (implemented in subsequent cells) compares paired responses to compute win-loss tallies, determining whether the OAMP agent's selective memory retrieval maintains or improves answer accuracy compared to the naive approach of sending complete conversation history.

## How Metrics Are Implemented in Code

The benchmark uses precise instrumentation to ensure comparable measurements across both memory strategies.

### Token Estimation Logic

The `estimate_tokens` function provides a lightweight approximation used throughout the benchmark:

```python
import json as _json_lib

def estimate_tokens(messages: list) -> int:
    """Approximate token count as characters/4 (same approach used in the benchmark)."""
    return len(_json_lib.dumps(messages)) // 4

# Usage during benchmark execution:

oamp_tokens = estimate_tokens(oamp_messages)
naive_tokens = estimate_tokens(naive_messages)

```

### Latency Instrumentation

Wall-clock timing uses `time.perf_counter()` for high-precision measurements. For the OAMP agent, the benchmark distinguishes between context preparation and total turn time:

```python
import time

def call_oamp_agent(user_query: str) -> str:
    t_start = time.perf_counter()
    thread.add_messages([Message(role="user", content=user_query)])
    context_card = thread.get_context_card() or "(no prior context)"
    t_context_built = time.perf_counter()
    
    # Build prompt with retrieved context

    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": f"Relevant memory:\n{context_card}\n\nCurrent question: {user_query}"},
    ]
    oamp_token_history.append(estimate_tokens(messages))
    
    # LLM call and response handling

    response = openai_client.chat.completions.create(
        model="gpt-5.4",
        messages=messages,
    )
    answer = response.choices[0].message.content or ""
    thread.add_messages([Message(role="assistant", content=answer)])
    
    t_end = time.perf_counter()
    oamp_retrieval_latency.append(t_context_built - t_start)
    oamp_total_latency.append(t_end - t_start)
    oamp_responses.append(answer)
    return answer

```

## Comparing OAMP vs Naive Implementation

The naive flat-history implementation serves as the baseline, demonstrating the performance characteristics of unbounded context growth:

```python
def call_naive_agent(user_query: str) -> str:
    t_start = time.perf_counter()
    naive_messages.append({"role": "user", "content": user_query})
    t_context_built = time.perf_counter()  # No retrieval overhead

    
    naive_token_history.append(estimate_tokens(naive_messages))
    
    response = openai_client.chat.completions.create(
        model="gpt-5.4",
        messages=naive_messages,  # Growing list of all prior messages

    )
    answer = response.choices[0].message.content or ""
    naive_messages.append({"role": "assistant", "content": answer})
    
    t_end = time.perf_counter()
    naive_retrieval_latency.append(t_context_built - t_start)  # Near zero

    naive_total_latency.append(t_end - t_start)
    naive_responses.append(answer)
    return answer

```

Key differences revealed by the benchmarking metrics include:

- **OAMP** incurs retrieval latency (`t_context_built - t_start`) but maintains bounded token counts by extracting only relevant memories
- **Naive** approach shows zero retrieval overhead but linearly increasing token consumption as `naive_messages` grows with each turn
- Both approaches are measured against identical 80-turn conversation scripts to ensure fair comparison

## Summary

- **Token consumption** is approximated using `estimate_tokens` (characters ÷ 4) and tracked separately in `oamp_token_history` and `naive_token_history` to compare API costs.
- **Wall-clock latency** distinguishes between retrieval time (`oamp_retrieval_latency`) and total turn time (`oamp_total_latency`), revealing the overhead of semantic memory extraction versus flat list appending.
- **Response quality** is captured in `oamp_responses` and `naive_responses` lists for subsequent LLM-as-a-judge evaluation to determine accuracy trade-offs.
- All metrics are collected in `oracle_agent_memory_benchmarks.ipynb` using high-precision `time.perf_counter()` measurements across an 80-turn standardized conversation.

## Frequently Asked Questions

### What file contains the OAMP benchmarking implementation?

The complete benchmark implementation resides in `notebooks/agent_memory/oracle_agent_memory_benchmarks.ipynb` within the `oracle-devrel/oracle-ai-developer-hub` repository. This Jupyter notebook contains the 80-turn test script, metric collection logic, and comparative visualization code.

### How does the benchmark measure token consumption without an official tokenizer?

The notebook uses a lightweight approximation via the `estimate_tokens` function, which calculates token count as the length of the serialized messages divided by 4 ([lines 85-89](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/notebooks/agent_memory/oracle_agent_memory_benchmarks.ipynb#L85-L89)). This provides a consistent relative comparison between OAMP and naive approaches without requiring model-specific tokenizers.

### Does OAMP add significant latency compared to naive flat-history?

According to the instrumentation in [lines 33-35](https://github.com/oracle-devrel/oracle-ai-developer-hub/blob/main/notebooks/agent_memory/oracle_agent_memory_benchmarks.ipynb#L33-L35), OAMP adds measurable retrieval latency (`t_context_built - t_start`) for fetching context cards, while the naive approach records near-zero retrieval time. However, as conversation length grows, the naive approach may exhibit higher total latency due to increasing token processing, which the `oamp_total_latency` and `naive_total_latency` metrics capture end-to-end.

### How is response quality evaluated when there is no single correct answer?

The benchmark implements an LLM-as-a-judge pattern where both agents' responses (stored in `oamp_responses` and `naive_responses`) are evaluated by a separate LLM instance after the conversation completes. This secondary model assesses which response better addresses the user's query given the conversation context, generating win-loss statistics that quantify accuracy trade-offs between selective memory retrieval and full history context.