How the Hindsight Consolidation Process Generates Observations From Memories

The Hindsight consolidation process transforms raw memory units into structured observations by batching unconsolidated facts by tags, recalling related existing knowledge, and using a single LLM prompt to generate validated create, update, and delete actions.

The Hindsight consolidation process is the autonomous engine within the vectorize-io/hindsight repository that converts ephemeral memory units (facts tagged as world or experience) into durable, queryable observations. This asynchronous pipeline runs immediately after retain operations, ensuring that raw captured data becomes structured knowledge while enforcing strict security boundaries through tag-based isolation.

Selecting Unconsolidated Memories

The consolidation engine begins by querying the database for pending work. In hindsight-api-slim/hindsight_api/engine/consolidation/consolidator.py (lines 15‑24), the system executes:

total_count = await conn.fetchval(
    f"""
    SELECT COUNT(*)
    FROM {fq_table("memory_units")}
    WHERE bank_id = $1
      AND consolidated_at IS NULL
      AND fact_type IN ('experience', 'world')
    """,
    bank_id,
)

If total_count returns zero, the process exits early with a no_new_memories status. Otherwise, the engine logs the pending count and proceeds to batch processing.

Tag-Based Batching for Security Isolation

To prevent cross-contamination between different security contexts, the engine groups memories by their exact tag sets. As implemented in lines 72‑84 of the consolidator:

tag_groups: dict[tuple[str, ...], list[dict[str, Any]]] = {}
for m in memories:
    tag_key = tuple(sorted(m.get("tags") or []))
    tag_groups.setdefault(tag_key, []).append(dict(m))

Each group sharing identical tags is then split into LLM-sized batches respecting config.consolidation_llm_batch_size. This security-first design guarantees that memories with different tags never share an LLM call, ensuring observations are only influenced by data within their authorized scope.

Per-Fact Observation Recall

Before invoking the LLM, the system performs a read-only recall of existing observations for each memory in the batch. Inside _process_memory_batch (lines 10‑24), the engine calls:

recall_tasks = [
    _find_related_observations(
        memory_engine=memory_engine,
        bank_id=bank_id,
        query=m["text"],
        request_context=request_context,
        tags=observation_scope_tags if observation_scope_tags is not None else (m.get("tags") or []),
    )
    for m in memories
]
per_fact_recalls = await asyncio.gather(*recall_tasks)

This step retrieves observation IDs that match the memory’s tags (or an override set), creating a validation set used later to gate update and delete actions.

Building the LLM Prompt and Calling the Model

The core intelligence of the consolidation process resides in _consolidate_batch_with_llm (lines 32‑66). The function constructs a prompt containing both the new facts and the union of recalled observations:

if union_observations:
    obs_list = _build_observations_for_llm(union_observations, union_source_facts)
    observations_text = json.dumps(obs_list, indent=2)
else:
    observations_text = "[]"

facts_lines = "\n".join(_fact_line(m) for m in memories)
prompt_template = build_batch_consolidation_prompt(observations_mission)
prompt = prompt_template.format(
    facts_text=facts_lines,
    observations_text=observations_text,
)

The LLM receives this context and returns a structured _ConsolidationBatchResponse describing creates, updates, and deletes. The system retries the call up to three times if needed.

Executing Create, Update, and Delete Actions

The engine validates every LLM-suggested action against the recall results before persisting changes.

Creating New Observations

For each create action, the system aggregates metadata from source memories using _aggregate_source_fields (lines 96‑118):

for create in llm_result.creates:
    source_mems = [mem_by_id[fid] for fid in create.source_fact_ids if fid in mem_by_id]
    agg = _aggregate_source_fields(source_mems, tags=fact_tags)
    await _execute_create_action(
        conn=conn,
        memory_engine=memory_engine,
        bank_id=bank_id,
        source_memory_ids=[m["id"] for m in source_mems],
        text=create.text,
        source_fact_tags=agg.tags,
        event_date=agg.event_date,
        occurred_start=agg.occurred_start,
        occurred_end=agg.occurred_end,
        mentioned_at=agg.mentioned_at,
        perf=perf,
    )

This merges timestamps (earliest event_date, latest mentioned_at) and inherits tags before inserting the new observation with fact_type='observation'.

Validating Updates and Deletes

Security validation prevents modifications to observations outside the recall scope. For updates (lines 95‑103):

if not any(update.observation_id in per_fact_obs_ids.get(str(m["id"]), set()) for m in source_mems):
    logger.debug("Batch consolidation: rejected update — observation not in recall")
    continue

Similarly, deletes (lines 124‑132) are rejected unless the target observation exists in the union_observations set recalled earlier.

Marking Memories as Consolidated

After executing all valid actions, the engine updates the source memories (lines 85‑88):

await conn.executemany(
    f"UPDATE {fq_table('memory_units')} SET consolidated_at = NOW() WHERE id = $1",
    [(m["id"],) for m in llm_batch],
)

This timestamp prevents future runs from reprocessing these units unless they are invalidated by tag changes.

Running Consolidation Manually

While the consolidation process runs automatically after retain operations, you can trigger it manually using the run_consolidation_job entry point:

from hindsight_api.engine.consolidation.consolidator import run_consolidation_job
from hindsight_api.memory_engine import MemoryEngine
from hindsight_api.api.http import RequestContext

# Initialize the memory engine

memory_engine = await MemoryEngine.create(...)

# Create request context

ctx = RequestContext(user_id="admin", auth_token="...")

# Trigger consolidation for a specific bank

result = await run_consolidation_job(
    memory_engine=memory_engine,
    bank_id="my-bank",
    request_context=ctx,
)

print(f"Status: {result['status']}")
print(f"Observations created: {result.get('observations_created', 0)}")

This call pulls pending memories, executes the five-stage pipeline, and returns metrics including the number of observations created.

Key Source Files

File Purpose
hindsight-api-slim/hindsight_api/engine/consolidation/consolidator.py Core orchestration, batch processing, LLM interaction, and action execution.
hindsight-api-slim/hindsight_api/engine/consolidation/prompts.py Prompt templates for the consolidation LLM calls.
hindsight-api-slim/hindsight_api/engine/memory_engine.py Helper methods for resetting consolidated_at and low-level observation writes.
hindsight-api-slim/hindsight_api/extensions/operation_validator.py Security policy validation hooks.

Summary

  • The Hindsight consolidation process automatically converts raw world and experience memory units into structured observation facts.
  • Tag-based batching ensures security isolation by preventing memories with different tags from sharing LLM contexts.
  • Per-fact recall retrieves existing observations to provide context and validate subsequent actions.
  • The LLM generates JSON-structured actions (create, update, delete) based on batched facts and recalled observations.
  • Security gates reject updates or deletes targeting observations outside the recalled set, preventing cross-tag contamination.
  • Processed memories are marked with consolidated_at timestamps to ensure idempotent processing.

Frequently Asked Questions

What triggers the Hindsight consolidation process?

The consolidation process runs automatically immediately after a retain operation completes. The run_consolidation_job function serves as the entry point, checking for unconsolidated memories where consolidated_at IS NULL and processing them asynchronously.

How does the consolidation process maintain data security?

Security is enforced through tag-based isolation. Memories are grouped by their exact tag sets, and each LLM batch contains only memories sharing identical tags. Additionally, update and delete actions are validated against the per-fact recall results, ensuring observations can only be modified by memories within their authorized tag scope.

What happens if the LLM suggests updating an observation from a different tag scope?

The engine rejects the update. Before executing any update, the system checks that the target observation_id exists in the recall set of at least one source memory (lines 95‑103). If the observation was not recalled for the current batch’s tags, the action is logged and discarded, preventing unauthorized cross-tag modifications.

Can I adjust the batch size for LLM calls?

Yes. The consolidation process respects the consolidation_llm_batch_size configuration parameter defined per bank. After grouping memories by tags, each tag group is further subdivided into batches that respect this size limit, allowing you to tune performance and cost based on your LLM provider’s rate limits.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →