How to Implement Automatic Fact Extraction and Entity Tracking in Agent Memory

The Oracle AI Developer Hub implements automatic fact extraction and entity tracking through the MemoryManager class, which uses LLM-driven extraction prompts to populate vector stores for semantic retrieval.

The Oracle AI Developer Hub provides a unified memory layer that enables agents to maintain long-term, searchable context across conversations. By implementing automatic fact extraction and entity tracking in agent memory, developers can ensure that AI agents retain critical information about people, places, and factual knowledge without manual intervention. This capability is centralized in the MemoryManager class located in apps/finance-ai-agent-demo/backend/memory/manager.py.

Architecture of the Memory System

The memory architecture distinguishes between several storage types, each optimized for specific retrieval patterns.

Memory Types and Storage Backends

Memory type Storage backend Typical use
Conversational Oracle AI Database table Chat history per thread
Knowledge-Base Vector-enabled table Searchable documents, factual excerpts
Entity Vector-enabled table Named-entity entries (people, places, instruments, accounts)
Summary Vector-enabled table Compressed snapshots of older conversations
Tool-Log Oracle AI Database table Off-loaded tool outputs

The Entity Extraction Pipeline

The entity memory serves as the foundation for automatic fact extraction. When processing user messages, the MemoryManager.write_entity method orchestrates a three-step pipeline:

  1. LLM Extraction: Calls _extract_entities, which prompts the LLM to return a JSON list of entities with name, type, and description (lines 410-426 in manager.py).
  2. Chunk Transformation: Transforms each entity into a text chunk formatted as "<name>: <description>".
  3. Vector Storage: Persists the chunk in the entity_vs vector store with metadata including entity type, timestamp, and optional thread_id.

The extraction prompt follows this structure:

Extract entities from: "<first 500-char of text>"
Return JSON: [{"name": "X", "type": "PERSON|PLACE|SYSTEM|INSTRUMENT|ACCOUNT", "description": "brief"}]
If none: []

Implementing Entity Tracking

To automatically track entities across conversations, initialize the MemoryManager with a configured entity vector store, then invoke write_entity during message processing.

Initializing the Memory Manager

from apps.finance_ai_agent_demo.backend.memory.manager import MemoryManager
from some_vector_store import OracleVectorStore

# Initialize vector stores

entity_vs = OracleVectorStore(table="ENTITY_MEMORY")
knowledge_vs = OracleVectorStore(table="KNOWLEDGE_BASE")

memory = MemoryManager(
    conn=db_connection,
    conversation_table="CONVERSATION_MEMORY",
    knowledge_base_vs=knowledge_vs,
    entity_vs=entity_vs,
    # ... other stores

)

Extracting and Storing Entities

Call write_entity with the LLM client to automatically extract and persist entities:

def process_message(message: str, thread_id: str, llm_client):
    # Store conversation turn

    memory.write_conversational_memory(
        content=message,
        role="user",
        thread_id=thread_id
    )
    
    # Automatic entity extraction

    memory.write_entity(
        name=None,
        entity_type=None,
        description=None,
        llm_client=llm_client,
        text=message,
        thread_id=thread_id
    )

This routes through write_entity → _extract_entities (lines 382-410), storing results in the entity vector store for later retrieval.

Retrieving Relevant Entities

Perform similarity searches against the entity store to enrich prompts with context:

def build_enriched_prompt(query: str, thread_id: str):
    # Retrieve top-5 relevant entities

    entity_block = memory.read_entity(
        query=query,
        k=5,
        thread_id=thread_id
    )
    
    return f"{BASE_PROMPT}\n\n{entity_block}"

The read_entity method (lines 441-452) executes a similarity search over entity_vs, automatically returning the most semantically relevant entities scoped to the specified thread.

Implementing Fact Extraction

Facts are stored in the knowledge-base memory rather than a dedicated facts table. The read_knowledge_base method (lines 215-223) performs similarity searches against the knowledge_base_vs vector store.

To make fact extraction automatic, implement a dual-write pattern that captures both structured entities and unstructured factual statements:

def store_facts(text: str, source: str, llm_client):
    metadata = {"source": source, "category": "fact"}
    memory.knowledge_base_vs.add_texts([text], [metadata])
    
    # Also extract entities from the same text

    memory.write_entity(
        name=None,
        entity_type=None,
        description=None,
        llm_client=llm_client,
        text=text,
        thread_id=None
    )

Query facts later using:

facts = memory.read_knowledge_base(query=user_question, k=3)

End-to-End Agent Integration

The complete flow implemented in apps/finance-ai-agent-demo/backend/agent/harness.py demonstrates automatic population of agent memory:

  1. Message Arrival: The harness receives a user message.
  2. Conversational Storage: Writes the raw message to conversational memory via write_conversational_memory.
  3. Entity Extraction: Calls MemoryManager.write_entity with the LLM client to extract and store named entities in entity_vs.
  4. Fact Ingestion: Optionally adds the raw text to the knowledge-base vector store.
  5. Context Retrieval: Before generating a response, the agent calls:
    • read_conversational_memory for recent dialogue
    • read_entity for relevant entities
    • read_knowledge_base for factual passages
    • read_summary_memory for compressed historical context
  6. Prompt Assembly: Concatenates all memory blocks into markdown-formatted context blocks injected into the system prompt.

Extending the Pattern

Custom Entity Types

Extend the extraction capabilities by modifying the JSON schema in _extract_entities. Add new enum values such as REGULATORY or PRODUCT to the type field, then update downstream processing logic to handle these categories.

Batch Processing

For bulk document processing, hook into the ingestion pipeline in apps/finance-ai-agent-demo/backend/ingestion/ingestor.py. After the extraction phase and before chunking, insert a call to _extract_entities to pre-populate the entity store with document-wide entities.

Dedicated Fact Stores

If separating facts from general knowledge is required, create a new vector store instance and expose a read_facts method following the pattern established in read_knowledge_base (lines 215-223).

Summary

  • Automatic fact extraction and entity tracking in agent memory relies on the MemoryManager class in apps/finance-ai-agent-demo/backend/memory/manager.py.
  • The _extract_entities method uses structured LLM prompts to identify entities with type classification (PERSON, PLACE, SYSTEM, INSTRUMENT, ACCOUNT).
  • Extracted entities are stored as vector embeddings in entity_vs, enabling semantic similarity search via read_entity.
  • Facts are maintained in the knowledge-base vector store (knowledge_base_vs) and retrieved through read_knowledge_base.
  • The agent harness orchestrates the complete flow: extraction during message processing, retrieval during context building, and injection into system prompts.

Frequently Asked Questions

How does the MemoryManager distinguish between entities and facts?

Entities are structured named items (people, places) extracted via _extract_entities and stored in entity_vs with type metadata. Facts are unstructured text chunks stored in knowledge_base_vs. While entities provide specific identity references, facts provide broader contextual knowledge, and both are retrieved via semantic similarity search.

Can entity extraction be customized for domain-specific types?

Yes. Modify the extraction prompt in _extract_entities (lines 410-426) to include additional enum values in the type field, such as REGULATORY or FINANCIAL_INSTRUMENT. Ensure downstream processing logic handles these new types when formatting retrieval results for prompts.

What is the performance impact of automatic extraction on every message?

The extraction requires one LLM call per message processed. According to the implementation in write_entity, this happens synchronously during message handling. For high-throughput scenarios, consider batching extractions or moving the operation to an asynchronous queue processed by the ingestion pipeline in ingestor.py.

How is entity retrieval scoped to specific conversations?

The write_entity method accepts an optional thread_id parameter that is stored in the vector metadata. When calling read_entity, passing the same thread_id filters the similarity search to entries within that conversation thread, preventing cross-contamination between different user sessions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →