How to Implement Context Engineering Techniques for Memory-Augmented Agents: A Complete Guide

Context engineering for memory-augmented agents involves orchestrating multiple specialized memory stores—conversational, knowledge-base, toolbox, and summary—to dynamically assemble relevant context while offloading heavy outputs and compressing historical turns to prevent token overflow.

The Oracle AI Developer Hub provides a production-ready reference implementation that demonstrates these techniques using a unified Oracle AI Database backend. The architecture centers on a MemoryManager class that coordinates seven distinct memory types, enabling deterministic context assembly while preserving the LLM's ability to choose when to retrieve or create information. This approach eliminates context window overflow and expensive向量 database synchronization by storing relational, document, and vector data in a single system.

Understanding the Multi-Tier Memory Architecture

Memory-augmented agents require specialized storage strategies for different data lifecycles. The implementation in apps/finance-ai-agent-demo/backend/memory/manager.py defines seven distinct memory types, each optimized for specific access patterns:

  • Conversational Memory: Stored in the CONVERSATIONAL_MEMORY SQL table, this holds unsummarized turn-by-turn chat history (typically ≤ few KB per turn) for immediate short-term context.
  • Knowledge-Base Memory: The SEMANTIC_MEMORY table stores 10⁴–10⁶ document vectors using LangChain's OracleVS abstraction with HNSW indexing for fast approximate nearest-neighbor search.
  • Workflow Memory: Procedural patterns and step-wise solutions reside in WORKFLOW_MEMORY (10³–10⁴ vectors), enabling the agent to recall previously successful multi-step processes.
  • Toolbox Memory: Semantic tool definitions stored in TOOLBOX_MEMORY (10²–10³ vectors) support dynamic function discovery via vector and hybrid text search.
  • Entity Memory: Extracted named entities (people, places, systems) live in ENTITY_MEMORY (10³–10⁴ vectors), allowing the agent to maintain persistent facts about specific objects.
  • Summary Memory: Compressed snapshots of older conversations stored in SUMMARY_MEMORY (10²–10³ vectors) provide just-in-time expansion when historical context is needed.
  • Tool Log: The TOOL_LOG table persists large tool outputs as CLOBs, preventing prompt bloat while maintaining traceability.

All vector-backed memories use LangChain's OracleVS with DistanceStrategy.COSINE, unifying relational and vector operations in a single database connection.

The Six-Step Context Engineering Workflow

The context assembly process follows a deterministic pattern that balances programmatic control with LLM autonomy:

  1. Assemble context – At the start of each turn, the harness loads recent conversational turns, knowledge-base hits, relevant entities, workflow snippets, and summary pointers.
  2. Retrieve tool schemas – Before inference, the system fetches the most relevant tool definitions from Toolbox memory using hybrid vector + Oracle Text search via read_toolbox().
  3. Execute model-driven actions – The LLM decides whether to call a tool, request summary expansion, or continue reasoning.
  4. Off-load heavy outputs – Full tool results are written to TOOL_LOG via write_tool_log() and replaced in the prompt with compact references ([Tool Log: …]).
  5. Persist new information – Assistant responses, new entities, and workflow steps are written back to their respective memory tables.
  6. Compress history – When token counts approach the model limit, older turns are summarized into SUMMARY_MEMORY using write_summary() and marked as processed.

This workflow solves three critical failure modes in agent systems:

Issue Traditional Approach Memory-Engineered Approach
Forgetting Rely on LLM's internal context window (≤ few thousand tokens) Persist every turn in CONVERSATIONAL_MEMORY; retrieve only relevant history
Context overflow Truncate old turns, losing critical details Off-load tool outputs to TOOL_LOG; compress history into SUMMARY_MEMORY
Expensive vector operations Separate vector DB (e.g., Pinecone) with custom sync logic Unified Oracle AI Database handles relational and vector data together

Implementing Context Assembly in Python

The reference notebook memory_context_engineering_agents.ipynb demonstrates initialization and context building. Below is a minimal implementation mirroring the production code:

import oracledb
from langchain_oracledb.vectorstores import OracleVS
from langchain_community.embeddings import HuggingFaceEmbeddings
from langchain_community.vectorstores.utils import DistanceStrategy
from apps.finance-ai-agent-demo.backend.memory.manager import MemoryManager

# Initialize vector-enabled connection

conn = oracledb.connect(
    user="VECTOR",
    password="VectorPwd_2025",
    dsn="127.0.0.1:1521/FREEPDB1",
)

# Configure embedding model

embed = HuggingFaceEmbeddings(model_name="sentence-transformers/paraphrase-mpnet-base-v2")

# Initialize vector stores for semantic memories

knowledge_vs = OracleVS(
    client=conn,
    embedding_function=embed,
    table_name="SEMANTIC_MEMORY",
    distance_strategy=DistanceStrategy.COSINE,
)

toolbox_vs = OracleVS(
    client=conn,
    embedding_function=embed,
    table_name="TOOLBOX_MEMORY",
    distance_strategy=DistanceStrategy.COSINE,
)

# Assemble MemoryManager with all storage backends

mem = MemoryManager(
    conn=conn,
    conversation_table="CONVERSATIONAL_MEMORY",
    knowledge_base_vs=knowledge_vs,
    toolbox_vs=toolbox_vs,
    # ... workflow_vs, entity_vs, summary_vs

    embedding_model=embed,
)

To build the context for a specific conversation thread, aggregate relevant memories programmatically:

def build_context(thread_id: str) -> str:
    """Assemble multi-source context for the LLM prompt."""
    parts = [
        mem.read_conversational_memory(thread_id, limit=10),
        mem.read_knowledge_base("latest regulatory updates", k=3),
        mem.read_entity("recent customers", k=5, thread_id=thread_id),
        mem.read_workflow("payment-routing", k=2, thread_id=thread_id),
        mem.read_summary_context(thread_id=thread_id, k=5),
        mem.read_toolbox("search-web", k=3),
    ]
    return "\n\n".join(parts)

Handling Tool Outputs and Context Compression

Off-loading Heavy Tool Results

When agents invoke tools that return large payloads (web searches, database queries, document retrieval), storing the full output in the prompt consumes valuable token budget. The write_tool_log() method persists raw outputs to the TOOL_LOG table while returning a compact reference:


# Simulate heavy tool execution

raw_output = heavy_web_search(query)  # Potentially megabytes of HTML

# Persist and get compact reference

tool_ref = mem.write_tool_log(
    thread_id=thread_id,
    call_id=tc.id,
    tool_name="web_search",
    arguments=args,
    output=raw_output
)

# Inject reference instead of full content

assistant_msg = f"{base_response}\n\n{tool_ref}"  # [Tool Log: web_search|...]

Just-in-Time Summarization

When estimate_tokens(prompt) approaches 80% of the model's limit (e.g., 128,000 tokens), compress older conversation history:

if estimate_tokens(prompt + assistant_msg) > token_budget * 0.8:
    # Generate summary (typically using a dedicated summarizer LLM)

    summary_text = summarise_conversation(thread_id, mem)
    summary_id = f"sum-{thread_id}-{datetime.utcnow().isoformat()}"
    
    # Store compressed version

    mem.write_summary(
        summary_id=summary_id,
        full_content=prompt,
        summary_text=summary_text,
        description="Compressed view of early turns",
        thread_id=thread_id
    )
    
    # Mark original rows as summarized to exclude from future loads

    unsummed = mem.get_unsummarized_messages(thread_id)
    mem.mark_as_summarized(
        thread_id=thread_id,
        summary_id=summary_id,
        message_ids=[m["id"] for m in unsummed]
    )

The read_summary_context() method returns summary pointers rather than full text, allowing the LLM to request specific historical expansions only when relevant.

Dynamic Tool Discovery with Semantic Retrieval

Hard-coded tool lists limit agent flexibility. The read_toolbox() method implements hybrid retrieval (vector similarity + Oracle Text) to dynamically select relevant tools:


# Retrieve semantic tool definitions

tools = mem.read_toolbox("search-web", k=3)

# Pass to LLM for dynamic function calling

response = client.chat.completions.create(
    model="gpt-5",
    messages=messages,
    tools=tools,  # Only relevant tools included

    temperature=0.2
)

This approach reduces latency by avoiding unnecessary tool descriptions in the prompt while ensuring the LLM has access to domain-specific capabilities stored in TOOLBOX_MEMORY.

Summary

  • Memory segmentation into seven specialized stores (conversational, knowledge-base, workflow, toolbox, entity, summary, and tool log) optimizes retention and retrieval efficiency.
  • Deterministic assembly via build_context() ensures critical data is always present, while LLM-driven expansion allows dynamic summary retrieval.
  • Tool-log off-loading prevents context window ballooning by replacing heavy outputs with database references using write_tool_log().
  • Hybrid search in TOOLBOX_MEMORY combines vector similarity and Oracle Text for low-latency tool discovery without hard-coding function lists.
  • Just-in-time summarization automatically compresses aging conversation turns when token budgets approach limits, maintaining infinite effective context.

Frequently Asked Questions

How does context engineering differ from simple prompt engineering?

Prompt engineering optimizes the text within a single inference call, while context engineering designs the infrastructure that selects, compresses, and retrieves information across multiple turns. According to the Oracle AI Developer Hub implementation, context engineering uses programmatic memory managers (like MemoryManager in manager.py) to orchestrate seven distinct storage types, ensuring the LLM receives relevant data without exceeding token limits, whereas prompt engineering only manipulates the immediate input text.

Can I implement this pattern without using Oracle Database?

Yes, though you lose unified storage benefits. The repository includes apps/finance-ai-agent-demo/backend/memory/sprawl_manager.py, which implements the same MemoryManager API using heterogeneous backends (PostgreSQL, MongoDB, and Qdrant). This "sprawl" approach works across multi-cloud environments but requires more complex synchronization logic compared to the single-database Oracle implementation.

What is the performance impact of HNSW vector indexes on query latency?

The OracleVS implementation using HNSW (Hierarchical Navigable Small World) indexing provides approximate nearest-neighbor search with sub-millisecond latency for typical knowledge-base sizes (10⁴–10⁶ vectors). Unlike external vector databases that require network round-trips, the co-located storage in Oracle AI Database eliminates data movement overhead, making semantic retrieval via read_knowledge_base() or read_entity() comparable in speed to standard SQL queries.

When should I use summary memory versus conversational memory?

Use conversational memory (CONVERSATIONAL_MEMORY) for recent, unprocessed turns that require full fidelity (typically the last 5–10 exchanges). Use summary memory (SUMMARY_MEMORY) for historical context older than the current session's immediate window. The system triggers compression via write_summary() when token counts exceed 80% of budget, moving older turns to compressed summaries that can be expanded on demand via read_summary_context(), effectively providing infinite context depth without infinite token costs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →