# How to Use Asynchronous Operations for Background Memory Processing in Hindsight

> Learn to use asynchronous operations for background memory processing in Hindsight. Utilize async methods for non-blocking memory writes and maintain application responsiveness.

- Repository: [vectorize-io/hindsight](https://github.com/vectorize-io/hindsight)
- Tags: how-to-guide
- Published: 2026-03-13

---

**Hindsight provides async client methods prefixed with `a` (like `aretain` and `arecall`) and a LiteLLM wrapper that runs retain operations in background threads, enabling non-blocking memory writes while your application continues processing.**

The `vectorize-io/hindsight` repository ships with native asynchronous support for memory operations, allowing you to perform background memory processing without blocking your main execution thread. Whether you are building a FastAPI service or integrating with existing synchronous code, Hindsight offers two distinct patterns for asynchronous operations: native `async`/`await` coroutines and fire-and-forget background threads. This guide covers both approaches with specific implementation details from the source code.

## Native Async Client Methods

The Python client exposes asynchronous versions of all high-level operations through methods prefixed with **`a`**. These coroutines live in [`hindsight-clients/python/hindsight_client/hindsight_client.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-clients/python/hindsight_client/hindsight_client.py) and return the same response models as their synchronous counterparts.

Available async methods include:

- **`aretain`** – Store a single memory asynchronously
- **`aretain_batch`** – Store multiple memories in one async request
- **`arecall`** – Perform semantic search asynchronously  
- **`areflect`** – Generate reasoning responses from stored facts asynchronously

Each method is implemented as an `async def` coroutine that you can `await` directly within an event loop. For example, `aretain` forwards to `aretain_batch` with the appropriate request structure:

```python
async def aretain(
    self,
    bank_id: str,
    content: str,
    # ... additional parameters

) -> RetainResponse:
    """Store a single memory (async)."""
    return await self.aretain_batch(
        bank_id=bank_id,
        items=[{"content": content}],
        # ... other args

    )

```

### Server-Side Background Processing

When you pass **`retain_async=True`** to `aretain` or `aretain_batch`, Hindsight queues the fact-extraction step server-side and returns immediately. This allows the worker process to handle the request later while your client continues execution. According to the source code in [`hindsight_client.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_client.py), the `async_` flag is set based on the `retain_async` argument, triggering background processing on the server.

## LiteLLM Integration with Background Threads

If you are using the LiteLLM integration, the `retain` helper defaults to asynchronous mode (`sync=False`) by spawning a daemon thread. This pattern, implemented in [`hindsight-integrations/litellm/hindsight_litellm/wrappers.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-integrations/litellm/hindsight_litellm/wrappers.py), provides fire-and-forget memory writes without requiring an async event loop in your application.

The background worker `_retain_background` runs the synchronous `_retain_sync` implementation in a separate thread:

```python
def _retain_background(
    content, api_url, target_bank_id, context,
    target_document_id, tags, metadata, verbose, api_key,
):
    """Background thread worker for async retain."""
    try:
        _retain_sync(
            content=content,
            api_url=api_url,
            target_bank_id=target_bank_id,
            context=context,
            # ... other parameters

        )
    except Exception as e:
        with _retain_errors_lock:
            _retain_errors.append(e)

```

### Error Handling for Background Operations

Because background threads execute outside the main flow, errors are captured in a thread-safe list. You can inspect failures by calling `get_pending_retain_errors()`, which retrieves and clears any exceptions that occurred during background processing. This function is defined in [`wrappers.py`](https://github.com/vectorize-io/hindsight/blob/main/wrappers.py) and returns a list of exception objects that you can log or handle as needed.

## Implementation Examples

### Using the Async Client Directly

For FastAPI or other async frameworks, use the native client methods with `await`:

```python
import asyncio
from hindsight_client import Hindsight

async def process_memory():
    client = Hindsight(base_url="http://localhost:8888")
    
    # Store memory with server-side background processing

    await client.aretain(
        bank_id="user-preferences",
        content="User prefers dark mode interface",
        retain_async=True  # Server extracts facts in background

    )
    
    # Search memories asynchronously

    results = await client.arecall(
        bank_id="user-preferences", 
        query="interface preferences"
    )
    
    # Generate reflection

    response = await client.areflect(
        bank_id="user-preferences",
        query="What do we know about the user's UI preferences?"
    )
    print(response.text)

asyncio.run(process_memory())

```

### Fire-and-Forget with LiteLLM

For synchronous applications that need non-blocking writes, use the LiteLLM wrapper:

```python
from hindsight_litellm import configure, retain, get_pending_retain_errors

# Configure once

configure(
    bank_id="agent-memory",
    hindsight_api_url="http://localhost:8888/v1/default",
    verbose=True,
)

# Async retain returns immediately (sync=False by default)

retain(
    content="Customer rated interaction 5 stars",
    context="feedback",
    tags=["customer:alice", "rating:5"],
)

# Check for background errors later

errors = get_pending_retain_errors()
if errors:
    for error in errors:
        print(f"Background retain failed: {error}")

```

### Complete Async Workflow Example

The repository includes a working demonstration in [`hindsight-docs/examples/api/main-methods.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-docs/examples/api/main-methods.py) showing all async methods:

```python
import asyncio
from hindsight import Hindsight

async def async_example():
    client = Hindsight(base_url="http://localhost:8888")
    
    # Store

    await client.aretain(bank_id="demo-bank", content="Async operation example")
    
    # Recall

    results = await client.arecall(bank_id="demo-bank", query="operation")
    for memory in results:
        print(f"- {memory.text}")
    
    # Reflect

    response = await client.areflect(
        bank_id="demo-bank", 
        query="What was stored?"
    )
    print(response.text)

asyncio.run(async_example())

```

## Summary

- **Native async methods** (`aretain`, `arecall`, `areflect`) in [`hindsight_client.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_client.py) provide `async`/`await` coroutines for use in event-loop-based applications.
- **Server-side background processing** via `retain_async=True` queues fact extraction on the server, returning control to the client immediately.
- **LiteLLM wrapper** offers fire-and-forget functionality through daemon threads in [`wrappers.py`](https://github.com/vectorize-io/hindsight/blob/main/wrappers.py), ideal for synchronous codebases.
- **Error inspection** for background threads is available through `get_pending_retain_errors()` to capture and handle failures asynchronously.

## Frequently Asked Questions

### What is the difference between `aretain` and `retain_async=True`?

**`aretain`** is a client-side asynchronous method that uses Python's `async`/`await` pattern to avoid blocking the event loop. **`retain_async=True`** is a server-side flag that tells Hindsight to queue the fact-extraction work and return immediately, processing the memory in a background worker. You can use them together for fully non-blocking operation from client to server.

### How do I handle errors when using the LiteLLM background thread?

Call `get_pending_retain_errors()` to retrieve a list of exceptions that occurred in the background thread. This function is thread-safe and clears the error queue upon reading, allowing you to periodically check for and log any retention failures without blocking your main application flow.

### Can I use the async client with Pydantic-AI or other agent frameworks?

Yes, the async client methods integrate directly with agent frameworks that support async tool calling. The [`hindsight-integrations/pydantic-ai/hindsight_pydantic_ai/tools.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-integrations/pydantic-ai/hindsight_pydantic_ai/tools.py) file demonstrates how to inject `client.aretain`, `client.arecall`, and `client.areflect` into Pydantic-AI tools, enabling asynchronous memory operations within agent workflows.

### When should I use the LiteLLM wrapper instead of the native async client?

Use the **LiteLLM wrapper** when you have synchronous code and want fire-and-forget memory writes without managing an event loop. Use the **native async client** when you are already running in an async environment like FastAPI, Sanic, or asyncio-based workers, and need fine-grained control over coroutine execution and error handling.