How to Implement Automatic Fact Extraction and Entity Tracking in Agent Memory
The Oracle AI Developer Hub implements automatic fact extraction and entity tracking through the MemoryManager class, which uses LLM-driven extraction prompts to populate vector stores for semantic retrieval.
The Oracle AI Developer Hub provides a unified memory layer that enables agents to maintain long-term, searchable context across conversations. By implementing automatic fact extraction and entity tracking in agent memory, developers can ensure that AI agents retain critical information about people, places, and factual knowledge without manual intervention. This capability is centralized in the MemoryManager class located in apps/finance-ai-agent-demo/backend/memory/manager.py.
Architecture of the Memory System
The memory architecture distinguishes between several storage types, each optimized for specific retrieval patterns.
Memory Types and Storage Backends
| Memory type | Storage backend | Typical use |
|---|---|---|
| Conversational | Oracle AI Database table | Chat history per thread |
| Knowledge-Base | Vector-enabled table | Searchable documents, factual excerpts |
| Entity | Vector-enabled table | Named-entity entries (people, places, instruments, accounts) |
| Summary | Vector-enabled table | Compressed snapshots of older conversations |
| Tool-Log | Oracle AI Database table | Off-loaded tool outputs |
The Entity Extraction Pipeline
The entity memory serves as the foundation for automatic fact extraction. When processing user messages, the MemoryManager.write_entity method orchestrates a three-step pipeline:
- LLM Extraction: Calls
_extract_entities, which prompts the LLM to return a JSON list of entities with name, type, and description (lines 410-426 inmanager.py). - Chunk Transformation: Transforms each entity into a text chunk formatted as
"<name>: <description>". - Vector Storage: Persists the chunk in the
entity_vsvector store with metadata including entity type, timestamp, and optionalthread_id.
The extraction prompt follows this structure:
Extract entities from: "<first 500-char of text>"
Return JSON: [{"name": "X", "type": "PERSON|PLACE|SYSTEM|INSTRUMENT|ACCOUNT", "description": "brief"}]
If none: []
Implementing Entity Tracking
To automatically track entities across conversations, initialize the MemoryManager with a configured entity vector store, then invoke write_entity during message processing.
Initializing the Memory Manager
from apps.finance_ai_agent_demo.backend.memory.manager import MemoryManager
from some_vector_store import OracleVectorStore
# Initialize vector stores
entity_vs = OracleVectorStore(table="ENTITY_MEMORY")
knowledge_vs = OracleVectorStore(table="KNOWLEDGE_BASE")
memory = MemoryManager(
conn=db_connection,
conversation_table="CONVERSATION_MEMORY",
knowledge_base_vs=knowledge_vs,
entity_vs=entity_vs,
# ... other stores
)
Extracting and Storing Entities
Call write_entity with the LLM client to automatically extract and persist entities:
def process_message(message: str, thread_id: str, llm_client):
# Store conversation turn
memory.write_conversational_memory(
content=message,
role="user",
thread_id=thread_id
)
# Automatic entity extraction
memory.write_entity(
name=None,
entity_type=None,
description=None,
llm_client=llm_client,
text=message,
thread_id=thread_id
)
This routes through write_entity → _extract_entities (lines 382-410), storing results in the entity vector store for later retrieval.
Retrieving Relevant Entities
Perform similarity searches against the entity store to enrich prompts with context:
def build_enriched_prompt(query: str, thread_id: str):
# Retrieve top-5 relevant entities
entity_block = memory.read_entity(
query=query,
k=5,
thread_id=thread_id
)
return f"{BASE_PROMPT}\n\n{entity_block}"
The read_entity method (lines 441-452) executes a similarity search over entity_vs, automatically returning the most semantically relevant entities scoped to the specified thread.
Implementing Fact Extraction
Facts are stored in the knowledge-base memory rather than a dedicated facts table. The read_knowledge_base method (lines 215-223) performs similarity searches against the knowledge_base_vs vector store.
To make fact extraction automatic, implement a dual-write pattern that captures both structured entities and unstructured factual statements:
def store_facts(text: str, source: str, llm_client):
metadata = {"source": source, "category": "fact"}
memory.knowledge_base_vs.add_texts([text], [metadata])
# Also extract entities from the same text
memory.write_entity(
name=None,
entity_type=None,
description=None,
llm_client=llm_client,
text=text,
thread_id=None
)
Query facts later using:
facts = memory.read_knowledge_base(query=user_question, k=3)
End-to-End Agent Integration
The complete flow implemented in apps/finance-ai-agent-demo/backend/agent/harness.py demonstrates automatic population of agent memory:
- Message Arrival: The harness receives a user message.
- Conversational Storage: Writes the raw message to conversational memory via
write_conversational_memory. - Entity Extraction: Calls
MemoryManager.write_entitywith the LLM client to extract and store named entities inentity_vs. - Fact Ingestion: Optionally adds the raw text to the knowledge-base vector store.
- Context Retrieval: Before generating a response, the agent calls:
read_conversational_memoryfor recent dialogueread_entityfor relevant entitiesread_knowledge_basefor factual passagesread_summary_memoryfor compressed historical context
- Prompt Assembly: Concatenates all memory blocks into markdown-formatted context blocks injected into the system prompt.
Extending the Pattern
Custom Entity Types
Extend the extraction capabilities by modifying the JSON schema in _extract_entities. Add new enum values such as REGULATORY or PRODUCT to the type field, then update downstream processing logic to handle these categories.
Batch Processing
For bulk document processing, hook into the ingestion pipeline in apps/finance-ai-agent-demo/backend/ingestion/ingestor.py. After the extraction phase and before chunking, insert a call to _extract_entities to pre-populate the entity store with document-wide entities.
Dedicated Fact Stores
If separating facts from general knowledge is required, create a new vector store instance and expose a read_facts method following the pattern established in read_knowledge_base (lines 215-223).
Summary
- Automatic fact extraction and entity tracking in agent memory relies on the
MemoryManagerclass inapps/finance-ai-agent-demo/backend/memory/manager.py. - The
_extract_entitiesmethod uses structured LLM prompts to identify entities with type classification (PERSON, PLACE, SYSTEM, INSTRUMENT, ACCOUNT). - Extracted entities are stored as vector embeddings in
entity_vs, enabling semantic similarity search viaread_entity. - Facts are maintained in the knowledge-base vector store (
knowledge_base_vs) and retrieved throughread_knowledge_base. - The agent harness orchestrates the complete flow: extraction during message processing, retrieval during context building, and injection into system prompts.
Frequently Asked Questions
How does the MemoryManager distinguish between entities and facts?
Entities are structured named items (people, places) extracted via _extract_entities and stored in entity_vs with type metadata. Facts are unstructured text chunks stored in knowledge_base_vs. While entities provide specific identity references, facts provide broader contextual knowledge, and both are retrieved via semantic similarity search.
Can entity extraction be customized for domain-specific types?
Yes. Modify the extraction prompt in _extract_entities (lines 410-426) to include additional enum values in the type field, such as REGULATORY or FINANCIAL_INSTRUMENT. Ensure downstream processing logic handles these new types when formatting retrieval results for prompts.
What is the performance impact of automatic extraction on every message?
The extraction requires one LLM call per message processed. According to the implementation in write_entity, this happens synchronously during message handling. For high-throughput scenarios, consider batching extractions or moving the operation to an asynchronous queue processed by the ingestion pipeline in ingestor.py.
How is entity retrieval scoped to specific conversations?
The write_entity method accepts an optional thread_id parameter that is stored in the vector metadata. When calling read_entity, passing the same thread_id filters the similarity search to entries within that conversation thread, preventing cross-contamination between different user sessions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →