# How Hindsight's Semantic Memory Architecture Works: A Deep Dive into the Three-Layer Link Graph

> Explore Hindsight's semantic memory architecture. Discover how its three layer link graph enables fast, contextual recall across millions of memories using vector-enriched text snippets.

- Repository: [vectorize-io/hindsight](https://github.com/vectorize-io/hindsight)
- Tags: deep-dive
- Published: 2026-03-13

---

**Hindsight stores facts as vector-enriched text snippets and connects them through semantic, temporal, and entity link graphs to enable fast, contextual recall across millions of memories.**

Hindsight is an open-source semantic memory system developed by Vectorize that transforms raw text into queryable knowledge. According to the vectorize-io/hindsight source code, the architecture centers on "memory units" enriched with embeddings and cross-referenced via three distinct relationship types, enabling retrieval that weighs meaning, time, and identity simultaneously.

## The Core Abstraction: Facts as Memory Units

At the heart of Hindsight's semantic memory architecture lies the **memory unit**. Each unit represents a discrete fact extracted from raw content and stored as a row in the sharded `memory_units` table.

Every memory unit contains:
- The original text content and context
- A vector embedding stored in the `embedding` column for similarity search
- Temporal metadata including `event_date` and timestamps
- A `fact_type` classification (`world`, `experience`, or `opinion`)
- `bank_id` for multi-tenant isolation
- References to source documents and chunk IDs

This schema lives in [`hindsight_api/engine/memory_engine.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api/engine/memory_engine.py), which serves as the primary facade for storage and retrieval operations.

## The Retain Pipeline: From Ingestion to Indexed Storage

When content enters the system via `hindsight_client.retain()`, it flows through a coordinated pipeline orchestrated by `retain_batch` in [`hindsight_api/engine/retain/orchestrator.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api/engine/retain/orchestrator.py).

### Fact Extraction and Embedding Generation

The pipeline first extracts atomic facts using LLM-driven extraction, then generates embeddings once per fact. These embeddings populate the `embedding` column of `memory_units` using pgvector's vector type. The system performs bulk inserts to minimize database round-trips.

### Bulk Link Creation

Immediately after storage, the system constructs three orthogonal relationship graphs in parallel:
1. **Semantic links** based on embedding similarity
2. **Temporal links** based on timestamp proximity  
3. **Entity links** based on shared real-world entities

The functions `create_semantic_links_batch`, `create_temporal_links_batch`, and `create_causal_links_batch` in [`hindsight_api/engine/retain/link_creation.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api/engine/retain/link_creation.py) handle this bulk creation efficiently.

## Three Orthogonal Link Graphs

Hindsight's retrieval power stems from maintaining three distinct link types that capture different dimensions of relatedness.

### Semantic Links via pgvector Similarity

Semantic connections form when fact embeddings are geometrically close in vector space. In [`hindsight_api/engine/retain/link_creation.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api/engine/retain/link_creation.py) (lines 35-48), the system compares new batch embeddings against existing ones using the pgvector `<=>` operator (cosine distance).

Only pairs exceeding a **configurable similarity threshold of 0.3** qualify as semantic links. The implementation streams `INSERT` operations in batches into the `memory_links` table with `link_type='semantic'` to maintain performance under load.

### Temporal Links via Time-Window Queries

Temporal connections capture "when" relationships. The function `create_temporal_links_batch_per_fact` in [`hindsight_api/engine/retain/link_utils.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api/engine/retain/link_utils.py) (lines 26-80) queries for candidate neighbors within a **default 24-hour window** of the `event_date`.

The system calculates a weight proportional to temporal proximity and stores it on the link, enabling time-aware spreading activation during retrieval.

### Entity Links via Resolved UUIDs

Entity links bridge facts mentioning the same real-world concepts. After LLM extraction, `extract_entities_batch_optimized` resolves entities to canonical UUIDs. The system then inserts bidirectional rows into `memory_links` with `link_type='entity'`.

To prevent combinatorial explosion, the code caps links per entity to **50 recent units** (implemented in [`hindsight_api/engine/retain/link_utils.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api/engine/retain/link_utils.py), lines 430-490).

## Optimized Storage and Retrieval

### Partial HNSW Indexes for Vector Search

Vector search performance relies on partial HNSW (Hierarchical Navigable Small World) indexes defined in the Alembic migration [`a3b4c5d6e7f8_add_partial_hnsw_indexes.py`](https://github.com/vectorize-io/hindsight/blob/main/a3b4c5d6e7f8_add_partial_hnsw_indexes.py). The migration creates separate indexes per fact type:
- `idx_mu_emb_world`
- `idx_mu_emb_observation`  
- `idx_mu_emb_experience`

Queries use `ORDER BY embedding <=> $1::vector LIMIT …`, allowing PostgreSQL's planner to automatically select the matching partial index and maintain **O(log N)** search complexity even with millions of facts.

### The Four-Arm Retrieval Strategy

When `client.recall()` executes, the retrieval driver in [`hindsight_api/engine/search/retrieval.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api/engine/search/retrieval.py) (lines 88-115) spawns four parallel retrieval arms:

1. **Semantic arm**: ANN vector search combined with BM25 full-text search via `retrieve_semantic_bm25_combined`, filtered by fact type
2. **BM25 arm**: Keyword search over the same fact set for lexical matching
3. **Graph arm**: Activation-based traversal using BFS, MPFP (Mean-Field Passing), or link-expansion over the three link graphs
4. **Temporal arm**: Time-aware spreading that starts from recent facts and follows temporal links

All arms return `RetrievalResult` objects. The system merges outputs using **Reciprocal Rank Fusion (RRF)** and optionally applies a cross-encoder reranker for final ordering.

## End-to-End Implementation

### Python Client Example

```python
from hindsight_client import HindsightClient

client = HindsightClient(base_url="https://api.hindsight.io", api_key="YOUR_KEY")

# Retain facts with metadata

facts = [
    {"content": "Alice bought a coffee at 9 am.", "tags": ["coffee"]},
    {"content": "Bob arrived at the office at 9:15 am.", "tags": ["office"]},
]

client.retain(
    bank_id="my-bank",
    contents=facts,
    document_id="meeting-2024-04-01",
)

# Recall using semantic similarity

response = client.recall(
    bank_id="my-bank",
    query="Who bought coffee this morning?",
    fact_type="world",
    budget="mid",
    max_tokens=1024,
)

for fact in response["results"]:
    print("-", fact["text"])

```

### Rust Client Example

```rust
use hindsight_client::HindsightClient;

let client = HindsightClient::new("https://api.hindsight.io", "YOUR_KEY");

let facts = vec![
    hindsight_client::Fact {
        content: "Alice bought a coffee at 9 am.".into(),
        tags: Some(vec!["coffee".into()]),
        ..Default::default()
    },
];

client.retain("my-bank", facts, Some("meeting-2024-04-01".into())).await?;

let result = client
    .recall(
        "my-bank",
        "Who bought coffee this morning?",
        Some(vec!["world".into()]),
        Some(hindsight_client::Budget::Mid),
        Some(1024),
    )
    .await?;

```

## Summary

- **Facts become memory units** stored in `memory_units` with vector embeddings, timestamps, and fact types (`world`, `experience`, `opinion`).
- **Three link graphs** connect facts: semantic (embedding similarity via pgvector `<=>`), temporal (24-hour window proximity), and entity (shared resolved UUIDs).
- **Bulk processing** in `retain_batch` and [`link_creation.py`](https://github.com/vectorize-io/hindsight/blob/main/link_creation.py) ensures efficient ingestion at scale.
- **Partial HNSW indexes** per fact type maintain sub-linear vector search complexity.
- **Four-arm retrieval** (semantic/BM25, BM25, graph, temporal) feeds into RRF ranking for contextually relevant recall.

## Frequently Asked Questions

### How does Hindsight determine if two facts are semantically related?

Hindsight compares vector embeddings using the pgvector cosine distance operator `<=>` in [`hindsight_api/engine/retain/link_creation.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api/engine/retain/link_creation.py). Facts with a distance below the configurable threshold of 0.3 create a semantic link in the `memory_links` table with `link_type='semantic'`.

### What prevents entity links from creating excessive database rows?

The system caps entity links at 50 recent units per entity (implemented in [`hindsight_api/engine/retain/link_utils.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api/engine/retain/link_utils.py) lines 430-490). This prevents combinatorial explosion when highly mentioned entities connect to thousands of facts.

### Why does Hindsight use partial HNSW indexes instead of one global index?

Partial indexes (`idx_mu_emb_world`, `idx_mu_emb_observation`, `idx_mu_emb_experience`) allow the query planner to filter by `fact_type` while performing ANN search. This keeps vector search **O(log N)** even as the table grows to millions of rows, as defined in migration [`a3b4c5d6e7f8_add_partial_hnsw_indexes.py`](https://github.com/vectorize-io/hindsight/blob/main/a3b4c5d6e7f8_add_partial_hnsw_indexes.py).

### How does the retrieval system balance semantic meaning with temporal relevance?

The four-arm retrieval strategy in [`hindsight_api/engine/search/retrieval.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api/engine/search/retrieval.py) runs semantic/BM25 search and temporal spreading activation in parallel. Results merge via Reciprocal Rank Fusion (RRF), which weights reciprocal ranks across all arms, ensuring both topical similarity and temporal proximity influence the final ranking.