How Hindsight Extracts and Regenerates Entities from Memories: A Complete Technical Guide

Hindsight extracts entities from memories during the retain phase using batched LLM processing and bulk database resolution, then regenerates entity mental models by re-running the reflect pipeline to synthesize updated observations.

The vectorize-io/hindsight open-source project treats entities as first-class objects attached to memory units, enabling sophisticated knowledge graph construction for AI agents. Understanding how entities are extracted and regenerated from memories is essential for developers building systems with persistent, evolving mental models that automatically update as new information arrives.

The Two-Phase Entity Lifecycle

Hindsight manages entities through two distinct operational phases: extraction during memory retention and regeneration when mental models require refreshing.

Phase 1: Extraction During Memory Retention

When a retain request arrives, the system parses raw text into discrete facts. Each fact's text is processed by an LLM to identify entity mentions, which are then resolved to canonical IDs and linked to the originating memory unit.

The extraction pipeline begins in hindsight-api-slim/hindsight_api/engine/retain/entity_processing.py where process_entities_batch prepares per-fact entity lists. This function extracts fact_text and timestamps from ProcessedFact objects:


# entity_processing.py

fact_texts = [fact.fact_text for fact in facts]
fact_dates = [fact.occurred_start if fact.occurred_start is not None else fact.mentioned_at for fact in facts]

The system merges LLM-extracted entities with any user-supplied entities, converting them into a standardized format:


# entity_processing.py

llm_entities = [{"text": entity.name, "type": "CONCEPT"} for entity in (fact.entities or [])]

Finally, extract_entities_batch_optimized in link_utils.py performs bulk resolution by calling resolve_entities_batch and creates the database links via link_units_to_entities_batch using a single INSERT … ON CONFLICT operation.

Phase 2: Regeneration via Mental Model Refresh

When an entity's synthesized view requires updating, clients invoke the refresh mental model workflow. This modern approach replaces the deprecated regenerate_entity_observations endpoint.

The regeneration logic resides in hindsight-api-slim/hindsight_api/engine/memory_engine.py within the refresh_mental_model function. This implementation retrieves the pinned mental model, re-executes reflect_async with the original source_query, and serializes the results:


# memory_engine.py

reflect_result = await self.reflect_async(
    bank_id=bank_id,
    query=mental_model["source_query"],
    request_context=request_context,
    tags=tags,
    tags_match=tags_match,
    exclude_mental_model_ids=[mental_model_id],
    _skip_span=True,
)

Code Implementation Details

Batch Entity Resolution Pipeline

The extraction system optimizes performance through aggressive batching. In entity_processing.py, process_entities_batch coordinates the flow:

  1. Extracts text and temporal metadata from each ProcessedFact
  2. Flattens all entities across facts while preserving event dates
  3. Invokes resolve_entities_batch once for all entities
  4. Creates unit-entity links in bulk via link_units_to_entities_batch

This approach minimizes database round-trips by performing single batched lookups and writes rather than individual queries per entity.

Mental Model Regeneration Architecture

The modern regeneration path avoids the deprecated POST /v1/default/banks/{bank_id}/entities/{entity_id}/regenerate endpoint defined in hindsight-clients/python/hindsight_client_api/api/entities_api.py.

Instead, refresh_mental_model in memory_engine.py orchestrates the process:

  • Fetches the existing mental model using get_mental_model
  • Excludes the current model from reflection results via exclude_mental_model_ids to prevent circular references
  • Updates the content and last_refreshed_at timestamp via update_mental_model

Practical Implementation Examples

Storing Memories with Automatic Entity Extraction

To store a memory and let Hindsight automatically extract entities:

from hindsight_client import HindsightClient

client = HindsightClient(base_url="https://api.hindsight.ai", api_key="YOUR_KEY")

# Retain triggers automatic entity extraction

retain_response = client.retain(
    bank_id="bank-123",
    content="Alice met Bob in Paris. The Eiffel Tower was illuminated."
)

# Inspect generated entity links

for unit in retain_response.units:
    print(f"Unit {unit.id} linked entities: {unit.entities}")

Behind the scenes, this triggers process_entities_batch → extract_entities_batch_optimized, creating the unit_entities rows that link memory units to canonical entity IDs.

Refreshing Entity Mental Models

To update an entity's synthesized observations after new facts arrive:


# Create a mental model summarizing knowledge about "Alice"

mm = client.create_mental_model(
    bank_id="bank-123",
    name="Alice Profile",
    source_query='entity:"Alice"'
)

# Refresh to incorporate new facts

refreshed = client.refresh_mental_model(
    bank_id="bank-123", 
    mental_model_id=mm.id
)

print(refreshed.content)

This call maps to memory_engine.refresh_mental_model, which re-runs reflect_async and atomically replaces the stored content with the new synthesis.

Legacy Regeneration Endpoint

For backward compatibility, the deprecated endpoint remains available in entities_api.py:


# Direct regeneration of observations (legacy API)

obs = client.regenerate_entity_observations(
    bank_id="bank-123",
    entity_id="entity-456"
)

New projects should migrate to refresh_mental_model for better performance and architectural consistency.

Summary

  • Extraction occurs during retention: The process_entities_batch function in entity_processing.py coordinates LLM-based entity identification and bulk resolution via extract_entities_batch_optimized in link_utils.py.
  • Batch operations optimize performance: Entity resolution and unit-entity linking happen in single database operations using resolve_entities_batch and link_units_to_entities_batch.
  • Regeneration uses the reflect pipeline: Modern workflows use refresh_mental_model in memory_engine.py to re-synthesize observations, replacing the deprecated regenerate_entity_observations endpoint.
  • Entity linking creates the knowledge graph: The system maintains bidirectional relationships between memory units and canonical entities, enabling complex graph traversals and mental model construction.

Frequently Asked Questions

How does Hindsight extract entities during the retain phase?

During the retain phase, process_entities_batch processes ProcessedFact objects by extracting text and timestamps, invoking an LLM to identify entity mentions, and merging these with user-supplied entities. The extract_entities_batch_optimized function then performs a single batched database lookup via resolve_entities_batch and creates unit-entity links using link_units_to_entities_batch.

What is the difference between regenerate_entity_observations and refresh_mental_model?

regenerate_entity_observations is a deprecated endpoint in entities_api.py that directly regenerates observations for specific entities. refresh_mental_model is the modern implementation in memory_engine.py that re-runs the reflect_async pipeline with the original source query to synthesize updated content while excluding the existing mental model from results.

Which source files handle entity extraction and linking?

Entity extraction logic resides in hindsight-api-slim/hindsight_api/engine/retain/entity_processing.py (preparation) and hindsight-api-slim/hindsight_api/engine/retain/link_utils.py (batch resolution and linking). Mental model regeneration is implemented in hindsight-api-slim/hindsight_api/engine/memory_engine.py.

How does the batch resolution system improve performance?

The batch resolution system collapses multiple entity lookups into a single resolve_entities_batch call and performs bulk inserts via link_units_to_entities_batch using INSERT … ON CONFLICT operations. This minimizes database round-trips when processing large memory batches containing numerous entity references.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →