# How the Hindsight Consolidation Process Generates Observations From Memories

> Learn how the Hindsight consolidation process turns raw memories into structured observations using batching knowledge recall and a single LLM prompt for actions.

- Repository: [vectorize-io/hindsight](https://github.com/vectorize-io/hindsight)
- Tags: internals
- Published: 2026-03-13

---

**The Hindsight consolidation process transforms raw memory units into structured observations by batching unconsolidated facts by tags, recalling related existing knowledge, and using a single LLM prompt to generate validated create, update, and delete actions.**

The **Hindsight consolidation process** is the autonomous engine within the `vectorize-io/hindsight` repository that converts ephemeral **memory units** (facts tagged as `world` or `experience`) into durable, queryable **observations**. This asynchronous pipeline runs immediately after retain operations, ensuring that raw captured data becomes structured knowledge while enforcing strict security boundaries through tag-based isolation.

## Selecting Unconsolidated Memories

The consolidation engine begins by querying the database for pending work. In [`hindsight-api-slim/hindsight_api/engine/consolidation/consolidator.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-api-slim/hindsight_api/engine/consolidation/consolidator.py) (lines 15‑24), the system executes:

```python
total_count = await conn.fetchval(
    f"""
    SELECT COUNT(*)
    FROM {fq_table("memory_units")}
    WHERE bank_id = $1
      AND consolidated_at IS NULL
      AND fact_type IN ('experience', 'world')
    """,
    bank_id,
)

```

If `total_count` returns zero, the process exits early with a `no_new_memories` status. Otherwise, the engine logs the pending count and proceeds to batch processing.

## Tag-Based Batching for Security Isolation

To prevent cross-contamination between different security contexts, the engine groups memories by their exact tag sets. As implemented in lines 72‑84 of the consolidator:

```python
tag_groups: dict[tuple[str, ...], list[dict[str, Any]]] = {}
for m in memories:
    tag_key = tuple(sorted(m.get("tags") or []))
    tag_groups.setdefault(tag_key, []).append(dict(m))

```

Each group sharing identical tags is then split into LLM-sized batches respecting `config.consolidation_llm_batch_size`. This **security-first design** guarantees that memories with different tags never share an LLM call, ensuring observations are only influenced by data within their authorized scope.

## Per-Fact Observation Recall

Before invoking the LLM, the system performs a **read-only recall** of existing observations for each memory in the batch. Inside `_process_memory_batch` (lines 10‑24), the engine calls:

```python
recall_tasks = [
    _find_related_observations(
        memory_engine=memory_engine,
        bank_id=bank_id,
        query=m["text"],
        request_context=request_context,
        tags=observation_scope_tags if observation_scope_tags is not None else (m.get("tags") or []),
    )
    for m in memories
]
per_fact_recalls = await asyncio.gather(*recall_tasks)

```

This step retrieves observation IDs that match the memory’s tags (or an override set), creating a validation set used later to gate update and delete actions.

## Building the LLM Prompt and Calling the Model

The core intelligence of the consolidation process resides in `_consolidate_batch_with_llm` (lines 32‑66). The function constructs a prompt containing both the new facts and the union of recalled observations:

```python
if union_observations:
    obs_list = _build_observations_for_llm(union_observations, union_source_facts)
    observations_text = json.dumps(obs_list, indent=2)
else:
    observations_text = "[]"

facts_lines = "\n".join(_fact_line(m) for m in memories)
prompt_template = build_batch_consolidation_prompt(observations_mission)
prompt = prompt_template.format(
    facts_text=facts_lines,
    observations_text=observations_text,
)

```

The LLM receives this context and returns a structured `_ConsolidationBatchResponse` describing **creates**, **updates**, and **deletes**. The system retries the call up to three times if needed.

## Executing Create, Update, and Delete Actions

The engine validates every LLM-suggested action against the recall results before persisting changes.

### Creating New Observations

For each create action, the system aggregates metadata from source memories using `_aggregate_source_fields` (lines 96‑118):

```python
for create in llm_result.creates:
    source_mems = [mem_by_id[fid] for fid in create.source_fact_ids if fid in mem_by_id]
    agg = _aggregate_source_fields(source_mems, tags=fact_tags)
    await _execute_create_action(
        conn=conn,
        memory_engine=memory_engine,
        bank_id=bank_id,
        source_memory_ids=[m["id"] for m in source_mems],
        text=create.text,
        source_fact_tags=agg.tags,
        event_date=agg.event_date,
        occurred_start=agg.occurred_start,
        occurred_end=agg.occurred_end,
        mentioned_at=agg.mentioned_at,
        perf=perf,
    )

```

This merges timestamps (earliest `event_date`, latest `mentioned_at`) and inherits tags before inserting the new observation with `fact_type='observation'`.

### Validating Updates and Deletes

Security validation prevents modifications to observations outside the recall scope. For updates (lines 95‑103):

```python
if not any(update.observation_id in per_fact_obs_ids.get(str(m["id"]), set()) for m in source_mems):
    logger.debug("Batch consolidation: rejected update — observation not in recall")
    continue

```

Similarly, deletes (lines 124‑132) are rejected unless the target observation exists in the `union_observations` set recalled earlier.

### Marking Memories as Consolidated

After executing all valid actions, the engine updates the source memories (lines 85‑88):

```python
await conn.executemany(
    f"UPDATE {fq_table('memory_units')} SET consolidated_at = NOW() WHERE id = $1",
    [(m["id"],) for m in llm_batch],
)

```

This timestamp prevents future runs from reprocessing these units unless they are invalidated by tag changes.

## Running Consolidation Manually

While the consolidation process runs automatically after retain operations, you can trigger it manually using the `run_consolidation_job` entry point:

```python
from hindsight_api.engine.consolidation.consolidator import run_consolidation_job
from hindsight_api.memory_engine import MemoryEngine
from hindsight_api.api.http import RequestContext

# Initialize the memory engine

memory_engine = await MemoryEngine.create(...)

# Create request context

ctx = RequestContext(user_id="admin", auth_token="...")

# Trigger consolidation for a specific bank

result = await run_consolidation_job(
    memory_engine=memory_engine,
    bank_id="my-bank",
    request_context=ctx,
)

print(f"Status: {result['status']}")
print(f"Observations created: {result.get('observations_created', 0)}")

```

This call pulls pending memories, executes the five-stage pipeline, and returns metrics including the number of observations created.

## Key Source Files

| File | Purpose |
|------|---------|
| [`hindsight-api-slim/hindsight_api/engine/consolidation/consolidator.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-api-slim/hindsight_api/engine/consolidation/consolidator.py) | Core orchestration, batch processing, LLM interaction, and action execution. |
| [`hindsight-api-slim/hindsight_api/engine/consolidation/prompts.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-api-slim/hindsight_api/engine/consolidation/prompts.py) | Prompt templates for the consolidation LLM calls. |
| [`hindsight-api-slim/hindsight_api/engine/memory_engine.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-api-slim/hindsight_api/engine/memory_engine.py) | Helper methods for resetting `consolidated_at` and low-level observation writes. |
| [`hindsight-api-slim/hindsight_api/extensions/operation_validator.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-api-slim/hindsight_api/extensions/operation_validator.py) | Security policy validation hooks. |

## Summary

- The **Hindsight consolidation process** automatically converts raw `world` and `experience` memory units into structured `observation` facts.
- **Tag-based batching** ensures security isolation by preventing memories with different tags from sharing LLM contexts.
- **Per-fact recall** retrieves existing observations to provide context and validate subsequent actions.
- The **LLM generates JSON-structured actions** (create, update, delete) based on batched facts and recalled observations.
- **Security gates** reject updates or deletes targeting observations outside the recalled set, preventing cross-tag contamination.
- Processed memories are marked with `consolidated_at` timestamps to ensure idempotent processing.

## Frequently Asked Questions

### What triggers the Hindsight consolidation process?

The consolidation process runs automatically immediately after a **retain operation** completes. The `run_consolidation_job` function serves as the entry point, checking for unconsolidated memories where `consolidated_at IS NULL` and processing them asynchronously.

### How does the consolidation process maintain data security?

Security is enforced through **tag-based isolation**. Memories are grouped by their exact tag sets, and each LLM batch contains only memories sharing identical tags. Additionally, update and delete actions are validated against the **per-fact recall results**, ensuring observations can only be modified by memories within their authorized tag scope.

### What happens if the LLM suggests updating an observation from a different tag scope?

The engine rejects the update. Before executing any update, the system checks that the target `observation_id` exists in the recall set of at least one source memory (lines 95‑103). If the observation was not recalled for the current batch’s tags, the action is logged and discarded, preventing unauthorized cross-tag modifications.

### Can I adjust the batch size for LLM calls?

Yes. The consolidation process respects the `consolidation_llm_batch_size` configuration parameter defined per bank. After grouping memories by tags, each tag group is further subdivided into batches that respect this size limit, allowing you to tune performance and cost based on your LLM provider’s rate limits.