How Hindsight's Semantic Memory Architecture Works: A Deep Dive into the Three-Layer Link Graph
Hindsight stores facts as vector-enriched text snippets and connects them through semantic, temporal, and entity link graphs to enable fast, contextual recall across millions of memories.
Hindsight is an open-source semantic memory system developed by Vectorize that transforms raw text into queryable knowledge. According to the vectorize-io/hindsight source code, the architecture centers on "memory units" enriched with embeddings and cross-referenced via three distinct relationship types, enabling retrieval that weighs meaning, time, and identity simultaneously.
The Core Abstraction: Facts as Memory Units
At the heart of Hindsight's semantic memory architecture lies the memory unit. Each unit represents a discrete fact extracted from raw content and stored as a row in the sharded memory_units table.
Every memory unit contains:
- The original text content and context
- A vector embedding stored in the
embeddingcolumn for similarity search - Temporal metadata including
event_dateand timestamps - A
fact_typeclassification (world,experience, oropinion) bank_idfor multi-tenant isolation- References to source documents and chunk IDs
This schema lives in hindsight_api/engine/memory_engine.py, which serves as the primary facade for storage and retrieval operations.
The Retain Pipeline: From Ingestion to Indexed Storage
When content enters the system via hindsight_client.retain(), it flows through a coordinated pipeline orchestrated by retain_batch in hindsight_api/engine/retain/orchestrator.py.
Fact Extraction and Embedding Generation
The pipeline first extracts atomic facts using LLM-driven extraction, then generates embeddings once per fact. These embeddings populate the embedding column of memory_units using pgvector's vector type. The system performs bulk inserts to minimize database round-trips.
Bulk Link Creation
Immediately after storage, the system constructs three orthogonal relationship graphs in parallel:
- Semantic links based on embedding similarity
- Temporal links based on timestamp proximity
- Entity links based on shared real-world entities
The functions create_semantic_links_batch, create_temporal_links_batch, and create_causal_links_batch in hindsight_api/engine/retain/link_creation.py handle this bulk creation efficiently.
Three Orthogonal Link Graphs
Hindsight's retrieval power stems from maintaining three distinct link types that capture different dimensions of relatedness.
Semantic Links via pgvector Similarity
Semantic connections form when fact embeddings are geometrically close in vector space. In hindsight_api/engine/retain/link_creation.py (lines 35-48), the system compares new batch embeddings against existing ones using the pgvector <=> operator (cosine distance).
Only pairs exceeding a configurable similarity threshold of 0.3 qualify as semantic links. The implementation streams INSERT operations in batches into the memory_links table with link_type='semantic' to maintain performance under load.
Temporal Links via Time-Window Queries
Temporal connections capture "when" relationships. The function create_temporal_links_batch_per_fact in hindsight_api/engine/retain/link_utils.py (lines 26-80) queries for candidate neighbors within a default 24-hour window of the event_date.
The system calculates a weight proportional to temporal proximity and stores it on the link, enabling time-aware spreading activation during retrieval.
Entity Links via Resolved UUIDs
Entity links bridge facts mentioning the same real-world concepts. After LLM extraction, extract_entities_batch_optimized resolves entities to canonical UUIDs. The system then inserts bidirectional rows into memory_links with link_type='entity'.
To prevent combinatorial explosion, the code caps links per entity to 50 recent units (implemented in hindsight_api/engine/retain/link_utils.py, lines 430-490).
Optimized Storage and Retrieval
Partial HNSW Indexes for Vector Search
Vector search performance relies on partial HNSW (Hierarchical Navigable Small World) indexes defined in the Alembic migration a3b4c5d6e7f8_add_partial_hnsw_indexes.py. The migration creates separate indexes per fact type:
idx_mu_emb_worldidx_mu_emb_observationidx_mu_emb_experience
Queries use ORDER BY embedding <=> $1::vector LIMIT …, allowing PostgreSQL's planner to automatically select the matching partial index and maintain O(log N) search complexity even with millions of facts.
The Four-Arm Retrieval Strategy
When client.recall() executes, the retrieval driver in hindsight_api/engine/search/retrieval.py (lines 88-115) spawns four parallel retrieval arms:
- Semantic arm: ANN vector search combined with BM25 full-text search via
retrieve_semantic_bm25_combined, filtered by fact type - BM25 arm: Keyword search over the same fact set for lexical matching
- Graph arm: Activation-based traversal using BFS, MPFP (Mean-Field Passing), or link-expansion over the three link graphs
- Temporal arm: Time-aware spreading that starts from recent facts and follows temporal links
All arms return RetrievalResult objects. The system merges outputs using Reciprocal Rank Fusion (RRF) and optionally applies a cross-encoder reranker for final ordering.
End-to-End Implementation
Python Client Example
from hindsight_client import HindsightClient
client = HindsightClient(base_url="https://api.hindsight.io", api_key="YOUR_KEY")
# Retain facts with metadata
facts = [
{"content": "Alice bought a coffee at 9 am.", "tags": ["coffee"]},
{"content": "Bob arrived at the office at 9:15 am.", "tags": ["office"]},
]
client.retain(
bank_id="my-bank",
contents=facts,
document_id="meeting-2024-04-01",
)
# Recall using semantic similarity
response = client.recall(
bank_id="my-bank",
query="Who bought coffee this morning?",
fact_type="world",
budget="mid",
max_tokens=1024,
)
for fact in response["results"]:
print("-", fact["text"])
Rust Client Example
use hindsight_client::HindsightClient;
let client = HindsightClient::new("https://api.hindsight.io", "YOUR_KEY");
let facts = vec![
hindsight_client::Fact {
content: "Alice bought a coffee at 9 am.".into(),
tags: Some(vec!["coffee".into()]),
..Default::default()
},
];
client.retain("my-bank", facts, Some("meeting-2024-04-01".into())).await?;
let result = client
.recall(
"my-bank",
"Who bought coffee this morning?",
Some(vec!["world".into()]),
Some(hindsight_client::Budget::Mid),
Some(1024),
)
.await?;
Summary
- Facts become memory units stored in
memory_unitswith vector embeddings, timestamps, and fact types (world,experience,opinion). - Three link graphs connect facts: semantic (embedding similarity via pgvector
<=>), temporal (24-hour window proximity), and entity (shared resolved UUIDs). - Bulk processing in
retain_batchandlink_creation.pyensures efficient ingestion at scale. - Partial HNSW indexes per fact type maintain sub-linear vector search complexity.
- Four-arm retrieval (semantic/BM25, BM25, graph, temporal) feeds into RRF ranking for contextually relevant recall.
Frequently Asked Questions
How does Hindsight determine if two facts are semantically related?
Hindsight compares vector embeddings using the pgvector cosine distance operator <=> in hindsight_api/engine/retain/link_creation.py. Facts with a distance below the configurable threshold of 0.3 create a semantic link in the memory_links table with link_type='semantic'.
What prevents entity links from creating excessive database rows?
The system caps entity links at 50 recent units per entity (implemented in hindsight_api/engine/retain/link_utils.py lines 430-490). This prevents combinatorial explosion when highly mentioned entities connect to thousands of facts.
Why does Hindsight use partial HNSW indexes instead of one global index?
Partial indexes (idx_mu_emb_world, idx_mu_emb_observation, idx_mu_emb_experience) allow the query planner to filter by fact_type while performing ANN search. This keeps vector search O(log N) even as the table grows to millions of rows, as defined in migration a3b4c5d6e7f8_add_partial_hnsw_indexes.py.
How does the retrieval system balance semantic meaning with temporal relevance?
The four-arm retrieval strategy in hindsight_api/engine/search/retrieval.py runs semantic/BM25 search and temporal spreading activation in parallel. Results merge via Reciprocal Rank Fusion (RRF), which weights reciprocal ranks across all arms, ensuring both topical similarity and temporal proximity influence the final ranking.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →