How Document-Based Memory Storage Works in Hindsight: Architecture and Implementation

Document-based memory storage in Hindsight treats every retained memory as a document within a bank, using a unique document_id to group related content and enable conversation stitching through automatic fetch-and-append logic.

Hindsight organizes persistent memory for LLM applications around a document-centric model. In the vectorize-io/hindsight repository, all information is stored as discrete documents identified by document_id parameters, which support both upsert semantics and scoped retrieval across the Python client and LiteLLM integration layers.

Core Architecture: Documents as Memory Units

Hindsight stores every piece of retained information as a document inside a bank. The document_id serves as the primary grouping mechanism, enabling several key behaviors:

  • Conversation stitching – All messages sharing the same document_id are automatically appended to create a coherent transcript
  • Upsert semantics – Existing documents can be replaced entirely or appended to, depending on integration settings
  • Scope control – Tags attached to documents determine visibility and retrieval permissions
  • Session continuity – The session_id field acts as an alias for document_id, with effective_document_id resolving the active identifier at runtime

The core flow initiates when a client calls retain or retain_batch with a document_id. The integration layer checks effective_document_id, fetches any existing content via _get_existing_document_content(), merges the new content (either concatenating or replacing), and persists the result back to the API under the same identifier.

Configuration: Resolving Document Identifiers

The system resolves document identifiers through a cascading property defined in hindsight_integrations/litellm/hindsight_litellm/config.py.

The HindsightCallSettings dataclass maintains two identifier fields:

@dataclass
class HindsightCallSettings:
    session_id: Optional[str] = None          # Primary – maps to Hindsight document_id

    document_id: Optional[str] = None         # Deprecated alias

    @property
    def effective_document_id(self) -> Optional[str]:
        """Return the ID used for Hindsight API calls."""
        return self.session_id if self.session_id is not None else self.document_id

This effective_document_id property ensures backward compatibility while prioritizing the modern session_id parameter. When either value is present, the LiteLLM wrapper triggers document-aware storage logic rather than creating isolated memory entries.

Fetch-and-Append Implementation

The conversation stitching logic resides in hindsight_integrations/litellm/hindsight_litellm/__init__.py. When effective_document_id is set, the wrapper executes the following sequence:

if defaults.effective_document_id:
    existing_content = _get_existing_document_content(
        defaults.bank_id, defaults.effective_document_id, config.verbose
    )
    if existing_content:
        content_to_store = f"{existing_content}\n\n{conversation_text}"
        if config.verbose:
            _storage_logger.debug(
                f"Appending to existing document: {defaults.effective_document_id}"
            )
retain(
    content=content_to_store,
    bank_id=defaults.bank_id,
    document_id=defaults.effective_document_id,
    # … additional parameters

)

The _get_existing_document_content() function performs an asynchronous lookup against the Documents API:

async def _get_existing_document_content(bank_id: str, document_id: str, verbose: bool):
    from hindsight_client_api.api import documents_api
    docs_api = documents_api.DocumentsApi(client._api_client)
    try:
        doc = await docs_api.get_document(bank_id, document_id)
        return doc.original_text
    except Exception as e:
        if verbose:
            _storage_logger.debug(f"Failed to fetch document: {e}")
        return None

If no existing document is found, the function returns None and the wrapper stores only the new content. If a document exists, the wrapper concatenates the existing text with the new conversation content using double newlines as delimiters.

Server-Side Document API

The underlying HTTP client for document operations is implemented in hindsight_client_api/api/documents_api.py. The DocumentsApi class provides CRUD operations for document-based memory storage:

class DocumentsApi:
    async def get_document(self, bank_id: str, document_id: str) -> DocumentResponse:
        # HTTP GET /banks/{bank_id}/documents/{document_id}

        ...

This API layer handles the actual persistence, retrieving the original_text field that the LiteLLM wrapper uses for append operations. The document response includes metadata tags and content that enable scoped recall operations in subsequent recall or reflect queries.

Practical Usage Patterns

Direct Retention with Document IDs

When using the Python client directly, explicitly specify the document_id to group related memories:

from hindsight_client import Hindsight

client = Hindsight(base_url="http://localhost:8888")
client.retain(
    bank_id="my-bank",
    content="Alice works at Google as a software engineer",
    document_id="user-profile-alice"   # groups later updates under the same doc

)

Batch Operations with Multiple Documents

The retain_batch method allows specifying distinct document identifiers for each item, enabling bulk operations across multiple conversation threads:

client.retain_batch(
    bank_id="my-bank",
    items=[
        {"content": "Alice works at Google", "document_id": "profile-alice"},
        {"content": "Bob is a data scientist", "document_id": "profile-bob"},
    ]
)

LiteLLM Integration for Automatic Stitching

The wrapper automatically manages document IDs for LiteLLM completion calls. Configure defaults once to enable transparent conversation stitching:

import hindsight_litellm

hindsight_litellm.configure(hindsight_api_url="http://localhost:8888")
hindsight_litellm.set_defaults(bank_id="agent-1", session_id="conv-123")
hindsight_litellm.enable()

# Subsequent calls automatically append to document "conv-123"

response = hindsight_litellm.completion(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
)

Upsert Semantics and Content Replacement

Document-based memory storage supports UPSERT operations. When calling retain with an existing document_id, the API replaces the previous content entirely:


# First call creates the document

client.retain(
    bank_id="my-bank",
    content="Initial version of the policy",
    document_id="policy-2024"
)

# Second call replaces the content due to matching ID

client.retain(
    bank_id="my-bank",
    content="Revised policy after legal review",
    document_id="policy-2024"
)

This behavior differs from the LiteLLM wrapper's append logic, which explicitly concatenates content rather than replacing it.

Summary

  • Document-centric model – Hindsight stores all memories as documents within banks, identified by unique document_id values
  • Identifier resolution – The effective_document_id property in config.py prioritizes session_id over the legacy document_id parameter
  • Conversation stitching – The _get_existing_document_content() function in __init__.py enables automatic append logic for continuous transcripts
  • Flexible semantics – Direct API calls support UPSERT (replacement), while the LiteLLM wrapper implements append semantics for conversation history
  • Scoped retrieval – Document-level tagging controls visibility for subsequent recall and reflect operations

Frequently Asked Questions

What is the difference between session_id and document_id in Hindsight?

The session_id parameter is the modern, preferred identifier for document-based memory storage, while document_id serves as a deprecated alias. According to the source code in hindsight_integrations/litellm/hindsight_litellm/config.py, the effective_document_id property returns session_id when present, falling back to document_id only when session_id is None. Both map to the same underlying document storage mechanism.

How does Hindsight handle concurrent writes to the same document ID?

The source code analysis reveals that Hindsight implements last-write-wins semantics for document-based memory storage. When multiple calls specify the same document_id, the retain API treats subsequent calls as replacement operations (UPSERT). For conversation stitching specifically, the LiteLLM wrapper in __init__.py first fetches existing content via _get_existing_document_content(), then concatenates new content before the final storage call, effectively serializing the append operation on the client side.

Can I retrieve specific documents directly without using the recall API?

Yes, the DocumentsApi class in hindsight_client_api/api/documents_api.py provides direct access to document-based memory storage through the get_document(bank_id, document_id) method. This returns a DocumentResponse object containing the original_text field and associated metadata, allowing you to inspect stored content before issuing new retain operations.

What happens if I don't specify a document_id when calling retain?

Without an explicit document_id, Hindsight stores the memory as an isolated entry without grouping. The document-based memory storage system only activates fetch-and-append logic when effective_document_id resolves to a non-null value. Omitting the identifier prevents conversation stitching and scope control, effectively treating each retain call as a distinct, unconnected memory unit.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →