How to Implement Deep Research Agents with Tavily Web Search and Oracle Memory Persistence

Implement deep research agents by wiring Tavily's web search API to Oracle's persistent vector memory, ensuring agents retrieve fresh data only when internal knowledge is exhausted and automatically persist findings for future reuse.

The oracle-devrel/oracle-ai-developer-hub repository demonstrates a production-grade pattern for building AI agents that combine real-time web retrieval with durable storage. This approach enables agents to overcome stale parametric knowledge while avoiding redundant API calls through intelligent caching in Oracle-backed Qdrant vector storage.

The Architecture of Retrieval-Augmented Persistence

Modern deep research agents require two complementary capabilities: retrieval-augmented reasoning to pull current web data, and long-term memory to store that data across conversational turns. The reference implementation organizes these concerns into distinct components:

Step 1: Registering the Tavily Search Tool

The search_tavily function (lines 95-118 of tools.py) implements the web retrieval interface. It uses the @toolbox.register_tool(augment=True) decorator to ensure its docstring is stored in toolbox memory, enabling semantic tool discovery:

from tavily import TavilyClient
from datetime import datetime
from toolbox import toolbox

tavily_client = TavilyClient(api_key=os.getenv("TAVILY_API_KEY"))

@toolbox.register_tool(augment=True)
def search_tavily(query: str, max_results: int = 5):
    """
    Search the web with Tavily and persist each result into the
    knowledge-base memory so the agent can reuse it later.
    """
    resp = tavily_client.search(query=query, max_results=max_results)
    for r in resp.get("results", []):
        text = f"Title: {r.get('title')}\nContent: {r.get('content')}\nURL: {r.get('url')}"
        meta = {
            "title": r.get("title", ""),
            "url": r.get("url", ""),
            "score": r.get("score", 0),
            "source_type": "tavily_search",
            "query": query,
            "timestamp": datetime.utcnow().isoformat(),
        }
        memory_manager.write_knowledge_base([text], [meta])
    return resp["results"]

This implementation returns up to five filtered results, formats each as a concise text block with metadata (title, URL, score, timestamp), and immediately persists them to the knowledge base.

Step 2: Persisting Results to Oracle Memory

The write_knowledge_base method (lines 92-100 of sprawl_manager.py) handles the actual vector storage. It generates embeddings via the manager's _embed method and upserts documents into Qdrant:


# Inside SprawlMemoryManager

def write_knowledge_base(self, texts, metadatas):
    """
    Index raw text + metadata in Qdrant's knowledge_base collection.
    """
    embeddings = self._embed(texts)
    points = [
        PointStruct(
            id=str(uuid.uuid4()),
            vector=emb,
            payload={"text": t, **m},
        )
        for t, m, emb in zip(texts, metadatas, embeddings, strict=False)
    ]
    self.qdrant.upsert(collection_name="knowledge_base", points=points)

Each web result is stored with its full context and semantic embedding, enabling future retrieval via memory_manager.search_knowledge_base without additional Tavily API calls.

Step 3: Enforcing the Search Priority Policy

To prevent unnecessary API usage and keep the context window lean, the system prompt (lines 53-55 of system_prompt.py) establishes a strict priority order:

  1. First, query the internal knowledge_base collection
  2. Only if no relevant documents are found, invoke search_tavily

This policy ensures that Tavily acts as a fallback, not a default. The agent checks its persistent Oracle memory before incurring the latency and cost of live web search.

Complete Implementation Workflow

The following pseudo-code illustrates how these components integrate in a production agent loop:

def handle_user_message(message, thread_id):
    # Phase 1: Check persistent memory first

    kb_hits = memory_manager.search_knowledge_base(embed(message), top_k=3)
    
    if kb_hits:
        return synthesize_answer(kb_hits)
    
    # Phase 2: Fallback to Tavily web search

    web_results = tools.search_tavily(message)
    
    # Phase 3: Retrieve the now-persisted results

    persisted = memory_manager.search_knowledge_base(embed(message), top_k=3)
    return synthesize_answer(persisted)

Because search_tavily writes to the knowledge base before returning, subsequent turns automatically benefit from cached web data.

Why This Design Scales

Concern Implementation Solution
Stale LLM knowledge Tavily supplies current, filtered content only when internal data is insufficient
Token budget Only summarized payloads (≈500 tokens per result) are stored, not raw HTML
Repeated web calls Each result is saved in the Qdrant vector store; subsequent queries hit the fast local index
Tool discoverability augment=True injects the tool description into TOOLBOX_MEMORY for semantic retrieval
Observability All tool calls are logged to the tool log table, enabling full audit trails

For a step-by-step tutorial, refer to workshops/agent_memory_workshop/docs/part-5-web-search.md in the repository.

Summary

  • Register tools with @toolbox.register_tool(augment=True) to enable semantic discovery while storing documentation in toolbox memory
  • Persist web results immediately using write_knowledge_base to avoid redundant Tavily API calls across conversation turns
  • Enforce search priority via system prompts that require knowledge-base exhaustion before web search fallback
  • Store metadata-rich vectors in Qdrant to maintain source attribution (URL, timestamp, score) for every web-derived fact
  • Leverage six memory types (conversation, knowledge, workflow, toolbox, entity, summary) through the unified MemoryManager API

Frequently Asked Questions

How does the agent decide when to use Tavily versus internal memory?

The system prompt in apps/finance-ai-agent-demo/backend/agent/system_prompt.py (lines 53-55) encodes a strict priority rule: the agent must first call search_knowledge_base. Only if that returns no relevant documents does the agent invoke search_tavily. This ensures web search is a last resort, preserving API quotas and reducing latency.

What memory types does the Oracle AI Developer Hub support?

The MemoryManager class in apps/finance-ai-agent-demo/backend/memory/manager.py manages six distinct memory types: conversation (turn history), knowledge (vector-stored facts including web results), workflow (process state), toolbox (tool definitions and docstrings), entity (extracted domain objects), and summary (condensed context). Each type is optimized for specific retrieval patterns.

How are web search results formatted before storage?

The search_tavily function formats each result as a structured text block containing the title, content snippet, and URL, augmented with metadata including the Tavily relevance score, source type, original query, and UTC timestamp. This structured approach (implemented in lines 100-110 of tools.py) ensures high-quality embeddings and clear provenance tracking.

Yes. By immediately persisting Tavily results to the Qdrant knowledge_base collection via write_knowledge_base, the agent creates a persistent cache. Subsequent queries on similar topics retrieve data from the local vector store rather than triggering new Tavily API calls, significantly reducing per-session costs while improving response times.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →