# How Document-Based Memory Storage Works in Hindsight: Architecture and Implementation

> Explore Hindsight's document-based memory storage. Learn how Hindsight treats memories as documents, uses document_id for grouping, and implements fetch-and-append logic for conversation stitching.

- Repository: [vectorize-io/hindsight](https://github.com/vectorize-io/hindsight)
- Tags: architecture
- Published: 2026-03-13

---

**Document-based memory storage in Hindsight treats every retained memory as a document within a bank, using a unique `document_id` to group related content and enable conversation stitching through automatic fetch-and-append logic.**

Hindsight organizes persistent memory for LLM applications around a document-centric model. In the `vectorize-io/hindsight` repository, all information is stored as discrete documents identified by `document_id` parameters, which support both upsert semantics and scoped retrieval across the Python client and LiteLLM integration layers.

## Core Architecture: Documents as Memory Units

Hindsight stores every piece of retained information as a **document** inside a *bank*. The `document_id` serves as the primary grouping mechanism, enabling several key behaviors:

- **Conversation stitching** – All messages sharing the same `document_id` are automatically appended to create a coherent transcript
- **Upsert semantics** – Existing documents can be replaced entirely or appended to, depending on integration settings
- **Scope control** – Tags attached to documents determine visibility and retrieval permissions
- **Session continuity** – The `session_id` field acts as an alias for `document_id`, with `effective_document_id` resolving the active identifier at runtime

The core flow initiates when a client calls `retain` or `retain_batch` with a `document_id`. The integration layer checks `effective_document_id`, fetches any existing content via `_get_existing_document_content()`, merges the new content (either concatenating or replacing), and persists the result back to the API under the same identifier.

## Configuration: Resolving Document Identifiers

The system resolves document identifiers through a cascading property defined in [`hindsight_integrations/litellm/hindsight_litellm/config.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_integrations/litellm/hindsight_litellm/config.py).

The `HindsightCallSettings` dataclass maintains two identifier fields:

```python
@dataclass
class HindsightCallSettings:
    session_id: Optional[str] = None          # Primary – maps to Hindsight document_id

    document_id: Optional[str] = None         # Deprecated alias

    @property
    def effective_document_id(self) -> Optional[str]:
        """Return the ID used for Hindsight API calls."""
        return self.session_id if self.session_id is not None else self.document_id

```

This `effective_document_id` property ensures backward compatibility while prioritizing the modern `session_id` parameter. When either value is present, the LiteLLM wrapper triggers document-aware storage logic rather than creating isolated memory entries.

## Fetch-and-Append Implementation

The conversation stitching logic resides in [`hindsight_integrations/litellm/hindsight_litellm/__init__.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_integrations/litellm/hindsight_litellm/__init__.py). When `effective_document_id` is set, the wrapper executes the following sequence:

```python
if defaults.effective_document_id:
    existing_content = _get_existing_document_content(
        defaults.bank_id, defaults.effective_document_id, config.verbose
    )
    if existing_content:
        content_to_store = f"{existing_content}\n\n{conversation_text}"
        if config.verbose:
            _storage_logger.debug(
                f"Appending to existing document: {defaults.effective_document_id}"
            )
retain(
    content=content_to_store,
    bank_id=defaults.bank_id,
    document_id=defaults.effective_document_id,
    # … additional parameters

)

```

The `_get_existing_document_content()` function performs an asynchronous lookup against the Documents API:

```python
async def _get_existing_document_content(bank_id: str, document_id: str, verbose: bool):
    from hindsight_client_api.api import documents_api
    docs_api = documents_api.DocumentsApi(client._api_client)
    try:
        doc = await docs_api.get_document(bank_id, document_id)
        return doc.original_text
    except Exception as e:
        if verbose:
            _storage_logger.debug(f"Failed to fetch document: {e}")
        return None

```

If no existing document is found, the function returns `None` and the wrapper stores only the new content. If a document exists, the wrapper concatenates the existing text with the new conversation content using double newlines as delimiters.

## Server-Side Document API

The underlying HTTP client for document operations is implemented in [`hindsight_client_api/api/documents_api.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_client_api/api/documents_api.py). The `DocumentsApi` class provides CRUD operations for document-based memory storage:

```python
class DocumentsApi:
    async def get_document(self, bank_id: str, document_id: str) -> DocumentResponse:
        # HTTP GET /banks/{bank_id}/documents/{document_id}

        ...

```

This API layer handles the actual persistence, retrieving the `original_text` field that the LiteLLM wrapper uses for append operations. The document response includes metadata tags and content that enable scoped recall operations in subsequent `recall` or `reflect` queries.

## Practical Usage Patterns

### Direct Retention with Document IDs

When using the Python client directly, explicitly specify the `document_id` to group related memories:

```python
from hindsight_client import Hindsight

client = Hindsight(base_url="http://localhost:8888")
client.retain(
    bank_id="my-bank",
    content="Alice works at Google as a software engineer",
    document_id="user-profile-alice"   # groups later updates under the same doc

)

```

### Batch Operations with Multiple Documents

The `retain_batch` method allows specifying distinct document identifiers for each item, enabling bulk operations across multiple conversation threads:

```python
client.retain_batch(
    bank_id="my-bank",
    items=[
        {"content": "Alice works at Google", "document_id": "profile-alice"},
        {"content": "Bob is a data scientist", "document_id": "profile-bob"},
    ]
)

```

### LiteLLM Integration for Automatic Stitching

The wrapper automatically manages document IDs for LiteLLM completion calls. Configure defaults once to enable transparent conversation stitching:

```python
import hindsight_litellm

hindsight_litellm.configure(hindsight_api_url="http://localhost:8888")
hindsight_litellm.set_defaults(bank_id="agent-1", session_id="conv-123")
hindsight_litellm.enable()

# Subsequent calls automatically append to document "conv-123"

response = hindsight_litellm.completion(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
)

```

### Upsert Semantics and Content Replacement

Document-based memory storage supports UPSERT operations. When calling `retain` with an existing `document_id`, the API replaces the previous content entirely:

```python

# First call creates the document

client.retain(
    bank_id="my-bank",
    content="Initial version of the policy",
    document_id="policy-2024"
)

# Second call replaces the content due to matching ID

client.retain(
    bank_id="my-bank",
    content="Revised policy after legal review",
    document_id="policy-2024"
)

```

This behavior differs from the LiteLLM wrapper's append logic, which explicitly concatenates content rather than replacing it.

## Summary

- **Document-centric model** – Hindsight stores all memories as documents within banks, identified by unique `document_id` values
- **Identifier resolution** – The `effective_document_id` property in [`config.py`](https://github.com/vectorize-io/hindsight/blob/main/config.py) prioritizes `session_id` over the legacy `document_id` parameter
- **Conversation stitching** – The `_get_existing_document_content()` function in [`__init__.py`](https://github.com/vectorize-io/hindsight/blob/main/__init__.py) enables automatic append logic for continuous transcripts
- **Flexible semantics** – Direct API calls support UPSERT (replacement), while the LiteLLM wrapper implements append semantics for conversation history
- **Scoped retrieval** – Document-level tagging controls visibility for subsequent `recall` and `reflect` operations

## Frequently Asked Questions

### What is the difference between `session_id` and `document_id` in Hindsight?

The `session_id` parameter is the modern, preferred identifier for document-based memory storage, while `document_id` serves as a deprecated alias. According to the source code in [`hindsight_integrations/litellm/hindsight_litellm/config.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_integrations/litellm/hindsight_litellm/config.py), the `effective_document_id` property returns `session_id` when present, falling back to `document_id` only when `session_id` is `None`. Both map to the same underlying document storage mechanism.

### How does Hindsight handle concurrent writes to the same document ID?

The source code analysis reveals that Hindsight implements last-write-wins semantics for document-based memory storage. When multiple calls specify the same `document_id`, the `retain` API treats subsequent calls as replacement operations (UPSERT). For conversation stitching specifically, the LiteLLM wrapper in [`__init__.py`](https://github.com/vectorize-io/hindsight/blob/main/__init__.py) first fetches existing content via `_get_existing_document_content()`, then concatenates new content before the final storage call, effectively serializing the append operation on the client side.

### Can I retrieve specific documents directly without using the recall API?

Yes, the `DocumentsApi` class in [`hindsight_client_api/api/documents_api.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_client_api/api/documents_api.py) provides direct access to document-based memory storage through the `get_document(bank_id, document_id)` method. This returns a `DocumentResponse` object containing the `original_text` field and associated metadata, allowing you to inspect stored content before issuing new `retain` operations.

### What happens if I don't specify a document_id when calling retain?

Without an explicit `document_id`, Hindsight stores the memory as an isolated entry without grouping. The document-based memory storage system only activates fetch-and-append logic when `effective_document_id` resolves to a non-null value. Omitting the identifier prevents conversation stitching and scope control, effectively treating each `retain` call as a distinct, unconnected memory unit.