How to Use Asynchronous Operations for Background Memory Processing in Hindsight
Hindsight provides async client methods prefixed with a (like aretain and arecall) and a LiteLLM wrapper that runs retain operations in background threads, enabling non-blocking memory writes while your application continues processing.
The vectorize-io/hindsight repository ships with native asynchronous support for memory operations, allowing you to perform background memory processing without blocking your main execution thread. Whether you are building a FastAPI service or integrating with existing synchronous code, Hindsight offers two distinct patterns for asynchronous operations: native async/await coroutines and fire-and-forget background threads. This guide covers both approaches with specific implementation details from the source code.
Native Async Client Methods
The Python client exposes asynchronous versions of all high-level operations through methods prefixed with a. These coroutines live in hindsight-clients/python/hindsight_client/hindsight_client.py and return the same response models as their synchronous counterparts.
Available async methods include:
aretain– Store a single memory asynchronouslyaretain_batch– Store multiple memories in one async requestarecall– Perform semantic search asynchronouslyareflect– Generate reasoning responses from stored facts asynchronously
Each method is implemented as an async def coroutine that you can await directly within an event loop. For example, aretain forwards to aretain_batch with the appropriate request structure:
async def aretain(
self,
bank_id: str,
content: str,
# ... additional parameters
) -> RetainResponse:
"""Store a single memory (async)."""
return await self.aretain_batch(
bank_id=bank_id,
items=[{"content": content}],
# ... other args
)
Server-Side Background Processing
When you pass retain_async=True to aretain or aretain_batch, Hindsight queues the fact-extraction step server-side and returns immediately. This allows the worker process to handle the request later while your client continues execution. According to the source code in hindsight_client.py, the async_ flag is set based on the retain_async argument, triggering background processing on the server.
LiteLLM Integration with Background Threads
If you are using the LiteLLM integration, the retain helper defaults to asynchronous mode (sync=False) by spawning a daemon thread. This pattern, implemented in hindsight-integrations/litellm/hindsight_litellm/wrappers.py, provides fire-and-forget memory writes without requiring an async event loop in your application.
The background worker _retain_background runs the synchronous _retain_sync implementation in a separate thread:
def _retain_background(
content, api_url, target_bank_id, context,
target_document_id, tags, metadata, verbose, api_key,
):
"""Background thread worker for async retain."""
try:
_retain_sync(
content=content,
api_url=api_url,
target_bank_id=target_bank_id,
context=context,
# ... other parameters
)
except Exception as e:
with _retain_errors_lock:
_retain_errors.append(e)
Error Handling for Background Operations
Because background threads execute outside the main flow, errors are captured in a thread-safe list. You can inspect failures by calling get_pending_retain_errors(), which retrieves and clears any exceptions that occurred during background processing. This function is defined in wrappers.py and returns a list of exception objects that you can log or handle as needed.
Implementation Examples
Using the Async Client Directly
For FastAPI or other async frameworks, use the native client methods with await:
import asyncio
from hindsight_client import Hindsight
async def process_memory():
client = Hindsight(base_url="http://localhost:8888")
# Store memory with server-side background processing
await client.aretain(
bank_id="user-preferences",
content="User prefers dark mode interface",
retain_async=True # Server extracts facts in background
)
# Search memories asynchronously
results = await client.arecall(
bank_id="user-preferences",
query="interface preferences"
)
# Generate reflection
response = await client.areflect(
bank_id="user-preferences",
query="What do we know about the user's UI preferences?"
)
print(response.text)
asyncio.run(process_memory())
Fire-and-Forget with LiteLLM
For synchronous applications that need non-blocking writes, use the LiteLLM wrapper:
from hindsight_litellm import configure, retain, get_pending_retain_errors
# Configure once
configure(
bank_id="agent-memory",
hindsight_api_url="http://localhost:8888/v1/default",
verbose=True,
)
# Async retain returns immediately (sync=False by default)
retain(
content="Customer rated interaction 5 stars",
context="feedback",
tags=["customer:alice", "rating:5"],
)
# Check for background errors later
errors = get_pending_retain_errors()
if errors:
for error in errors:
print(f"Background retain failed: {error}")
Complete Async Workflow Example
The repository includes a working demonstration in hindsight-docs/examples/api/main-methods.py showing all async methods:
import asyncio
from hindsight import Hindsight
async def async_example():
client = Hindsight(base_url="http://localhost:8888")
# Store
await client.aretain(bank_id="demo-bank", content="Async operation example")
# Recall
results = await client.arecall(bank_id="demo-bank", query="operation")
for memory in results:
print(f"- {memory.text}")
# Reflect
response = await client.areflect(
bank_id="demo-bank",
query="What was stored?"
)
print(response.text)
asyncio.run(async_example())
Summary
- Native async methods (
aretain,arecall,areflect) inhindsight_client.pyprovideasync/awaitcoroutines for use in event-loop-based applications. - Server-side background processing via
retain_async=Truequeues fact extraction on the server, returning control to the client immediately. - LiteLLM wrapper offers fire-and-forget functionality through daemon threads in
wrappers.py, ideal for synchronous codebases. - Error inspection for background threads is available through
get_pending_retain_errors()to capture and handle failures asynchronously.
Frequently Asked Questions
What is the difference between aretain and retain_async=True?
aretain is a client-side asynchronous method that uses Python's async/await pattern to avoid blocking the event loop. retain_async=True is a server-side flag that tells Hindsight to queue the fact-extraction work and return immediately, processing the memory in a background worker. You can use them together for fully non-blocking operation from client to server.
How do I handle errors when using the LiteLLM background thread?
Call get_pending_retain_errors() to retrieve a list of exceptions that occurred in the background thread. This function is thread-safe and clears the error queue upon reading, allowing you to periodically check for and log any retention failures without blocking your main application flow.
Can I use the async client with Pydantic-AI or other agent frameworks?
Yes, the async client methods integrate directly with agent frameworks that support async tool calling. The hindsight-integrations/pydantic-ai/hindsight_pydantic_ai/tools.py file demonstrates how to inject client.aretain, client.arecall, and client.areflect into Pydantic-AI tools, enabling asynchronous memory operations within agent workflows.
When should I use the LiteLLM wrapper instead of the native async client?
Use the LiteLLM wrapper when you have synchronous code and want fire-and-forget memory writes without managing an event loop. Use the native async client when you are already running in an async environment like FastAPI, Sanic, or asyncio-based workers, and need fine-grained control over coroutine execution and error handling.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →