How Hindsight's Semantic Memory Architecture Works: A Deep Dive into the Three-Layer Link Graph

Hindsight stores facts as vector-enriched text snippets and connects them through semantic, temporal, and entity link graphs to enable fast, contextual recall across millions of memories.

Hindsight is an open-source semantic memory system developed by Vectorize that transforms raw text into queryable knowledge. According to the vectorize-io/hindsight source code, the architecture centers on "memory units" enriched with embeddings and cross-referenced via three distinct relationship types, enabling retrieval that weighs meaning, time, and identity simultaneously.

The Core Abstraction: Facts as Memory Units

At the heart of Hindsight's semantic memory architecture lies the memory unit. Each unit represents a discrete fact extracted from raw content and stored as a row in the sharded memory_units table.

Every memory unit contains:

  • The original text content and context
  • A vector embedding stored in the embedding column for similarity search
  • Temporal metadata including event_date and timestamps
  • A fact_type classification (world, experience, or opinion)
  • bank_id for multi-tenant isolation
  • References to source documents and chunk IDs

This schema lives in hindsight_api/engine/memory_engine.py, which serves as the primary facade for storage and retrieval operations.

The Retain Pipeline: From Ingestion to Indexed Storage

When content enters the system via hindsight_client.retain(), it flows through a coordinated pipeline orchestrated by retain_batch in hindsight_api/engine/retain/orchestrator.py.

Fact Extraction and Embedding Generation

The pipeline first extracts atomic facts using LLM-driven extraction, then generates embeddings once per fact. These embeddings populate the embedding column of memory_units using pgvector's vector type. The system performs bulk inserts to minimize database round-trips.

Immediately after storage, the system constructs three orthogonal relationship graphs in parallel:

  1. Semantic links based on embedding similarity
  2. Temporal links based on timestamp proximity
  3. Entity links based on shared real-world entities

The functions create_semantic_links_batch, create_temporal_links_batch, and create_causal_links_batch in hindsight_api/engine/retain/link_creation.py handle this bulk creation efficiently.

Hindsight's retrieval power stems from maintaining three distinct link types that capture different dimensions of relatedness.

Semantic connections form when fact embeddings are geometrically close in vector space. In hindsight_api/engine/retain/link_creation.py (lines 35-48), the system compares new batch embeddings against existing ones using the pgvector <=> operator (cosine distance).

Only pairs exceeding a configurable similarity threshold of 0.3 qualify as semantic links. The implementation streams INSERT operations in batches into the memory_links table with link_type='semantic' to maintain performance under load.

Temporal connections capture "when" relationships. The function create_temporal_links_batch_per_fact in hindsight_api/engine/retain/link_utils.py (lines 26-80) queries for candidate neighbors within a default 24-hour window of the event_date.

The system calculates a weight proportional to temporal proximity and stores it on the link, enabling time-aware spreading activation during retrieval.

Entity links bridge facts mentioning the same real-world concepts. After LLM extraction, extract_entities_batch_optimized resolves entities to canonical UUIDs. The system then inserts bidirectional rows into memory_links with link_type='entity'.

To prevent combinatorial explosion, the code caps links per entity to 50 recent units (implemented in hindsight_api/engine/retain/link_utils.py, lines 430-490).

Optimized Storage and Retrieval

Vector search performance relies on partial HNSW (Hierarchical Navigable Small World) indexes defined in the Alembic migration a3b4c5d6e7f8_add_partial_hnsw_indexes.py. The migration creates separate indexes per fact type:

  • idx_mu_emb_world
  • idx_mu_emb_observation
  • idx_mu_emb_experience

Queries use ORDER BY embedding <=> $1::vector LIMIT …, allowing PostgreSQL's planner to automatically select the matching partial index and maintain O(log N) search complexity even with millions of facts.

The Four-Arm Retrieval Strategy

When client.recall() executes, the retrieval driver in hindsight_api/engine/search/retrieval.py (lines 88-115) spawns four parallel retrieval arms:

  1. Semantic arm: ANN vector search combined with BM25 full-text search via retrieve_semantic_bm25_combined, filtered by fact type
  2. BM25 arm: Keyword search over the same fact set for lexical matching
  3. Graph arm: Activation-based traversal using BFS, MPFP (Mean-Field Passing), or link-expansion over the three link graphs
  4. Temporal arm: Time-aware spreading that starts from recent facts and follows temporal links

All arms return RetrievalResult objects. The system merges outputs using Reciprocal Rank Fusion (RRF) and optionally applies a cross-encoder reranker for final ordering.

End-to-End Implementation

Python Client Example

from hindsight_client import HindsightClient

client = HindsightClient(base_url="https://api.hindsight.io", api_key="YOUR_KEY")

# Retain facts with metadata

facts = [
    {"content": "Alice bought a coffee at 9 am.", "tags": ["coffee"]},
    {"content": "Bob arrived at the office at 9:15 am.", "tags": ["office"]},
]

client.retain(
    bank_id="my-bank",
    contents=facts,
    document_id="meeting-2024-04-01",
)

# Recall using semantic similarity

response = client.recall(
    bank_id="my-bank",
    query="Who bought coffee this morning?",
    fact_type="world",
    budget="mid",
    max_tokens=1024,
)

for fact in response["results"]:
    print("-", fact["text"])

Rust Client Example

use hindsight_client::HindsightClient;

let client = HindsightClient::new("https://api.hindsight.io", "YOUR_KEY");

let facts = vec![
    hindsight_client::Fact {
        content: "Alice bought a coffee at 9 am.".into(),
        tags: Some(vec!["coffee".into()]),
        ..Default::default()
    },
];

client.retain("my-bank", facts, Some("meeting-2024-04-01".into())).await?;

let result = client
    .recall(
        "my-bank",
        "Who bought coffee this morning?",
        Some(vec!["world".into()]),
        Some(hindsight_client::Budget::Mid),
        Some(1024),
    )
    .await?;

Summary

  • Facts become memory units stored in memory_units with vector embeddings, timestamps, and fact types (world, experience, opinion).
  • Three link graphs connect facts: semantic (embedding similarity via pgvector <=>), temporal (24-hour window proximity), and entity (shared resolved UUIDs).
  • Bulk processing in retain_batch and link_creation.py ensures efficient ingestion at scale.
  • Partial HNSW indexes per fact type maintain sub-linear vector search complexity.
  • Four-arm retrieval (semantic/BM25, BM25, graph, temporal) feeds into RRF ranking for contextually relevant recall.

Frequently Asked Questions

Hindsight compares vector embeddings using the pgvector cosine distance operator <=> in hindsight_api/engine/retain/link_creation.py. Facts with a distance below the configurable threshold of 0.3 create a semantic link in the memory_links table with link_type='semantic'.

The system caps entity links at 50 recent units per entity (implemented in hindsight_api/engine/retain/link_utils.py lines 430-490). This prevents combinatorial explosion when highly mentioned entities connect to thousands of facts.

Why does Hindsight use partial HNSW indexes instead of one global index?

Partial indexes (idx_mu_emb_world, idx_mu_emb_observation, idx_mu_emb_experience) allow the query planner to filter by fact_type while performing ANN search. This keeps vector search O(log N) even as the table grows to millions of rows, as defined in migration a3b4c5d6e7f8_add_partial_hnsw_indexes.py.

How does the retrieval system balance semantic meaning with temporal relevance?

The four-arm retrieval strategy in hindsight_api/engine/search/retrieval.py runs semantic/BM25 search and temporal spreading activation in parallel. Results merge via Reciprocal Rank Fusion (RRF), which weights reciprocal ranks across all arms, ensuring both topical similarity and temporal proximity influence the final ranking.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →