What Is the Difference Between Recall and Reflect in Hindsight?
Recall retrieves raw memory facts through parallel search strategies, while Reflect executes an agentic reasoning loop that synthesizes interpreted answers according to a bank's disposition and directives.
The Hindsight memory platform by Vectorize provides two distinct operations for accessing stored knowledge. Understanding the architectural distinction between these recall and reflect operations in Hindsight enables developers to choose between high-performance raw retrieval and context-aware intelligent synthesis.
Architectural Goals: Search vs. Synthesis
The fundamental difference between these operations lies in their output intent.
- Recall functions as a high-performance retrieval service. It fetches raw memories, entities, and source chunks that match a query without interpretation.
- Reflect operates as an agentic reasoning system. It thinks through retrieved evidence, applies personality traits and hard rules, and produces a polished answer with citations.
According to the Hindsight source code, Recall is implemented in hindsight-control-plane/src/app/api/recall/route.ts, while Reflect resides in hindsight-control-plane/src/app/api/reflect/route.ts.
How Recall Retrieves Raw Memories
Recall executes four parallel search strategies to maximize relevant memory coverage.
When you invoke the POST /api/recall endpoint or the client.recall() method from the TypeScript SDK (defined at hindsight-clients/typescript/src/index.ts lines 52-98), the system simultaneously runs:
- Semantic search for conceptual similarity
- Keyword search for exact term matching
- Graph search for relational traversal
- Temporal search for time-based relevance
The system fuses these results using Reciprocal Rank Fusion, then reranks them with a cross-encoder. Recall stops fetching when the requested max_tokens budget exhausts, returning a plain list of memory results, optional entities, and source chunks.
This design prioritizes speed and comprehensiveness for fact look-up and context window building.
How Reflect Performs Agentic Reasoning
Reflect implements a hierarchical agentic loop that synthesizes rather than retrieves.
When you call POST /api/reflect or client.reflect() (implemented at hindsight-clients/typescript/src/index.ts lines 101-130), the system initiates a reasoning cycle that repeatedly invokes tools until gathering sufficient evidence.
The tool hierarchy follows this sequence:
- Mental-model search for high-level frameworks
- Observation search for specific instances
- Raw recall for underlying facts
- Expand for drilling into details
Crucially, this loop respects the bank's disposition—a configurable personality trait such as skepticism, literalism, or empathy—and any directives (hard rules that must be followed). The final output includes a synthesized text response, a based_on citation list, and optional trace data for debugging the reasoning path.
API and SDK Usage Examples
Both operations are available through REST endpoints and official SDKs.
TypeScript / JavaScript SDK
import { HindsightClient } from '@vectorize/hindsight';
const client = new HindsightClient({ baseUrl: 'https://api.hindsight.io', apiKey: 'YOUR_KEY' });
// Recall: Raw memory retrieval
const recallResult = await client.recall(
'my-bank',
'What did Alice say about Python?',
{
maxTokens: 1024,
budget: 'mid',
includeChunks: true,
}
);
console.log(recallResult.results);
// Reflect: Synthesized reasoning
const reflectResult = await client.reflect(
'my-bank',
'Should we hire Alice as a senior engineer?',
{
budget: 'mid',
tags: ['team:backend'],
}
);
console.log(reflectResult.text);
console.log(reflectResult.based_on);
Command Line Interface
# Recall raw memories
hindsight memory recall my-bank "What were the deadlines for Project X?" --max-tokens 2048
# Reflect with reasoning
hindsight memory reflect my-bank "Is Project X on track for Q4 launch?" --budget mid
Python Client
from hindsight import HindsightClient
client = HindsightClient(base_url="https://api.hindsight.io", api_key="YOUR_KEY")
# Retrieve facts
recall = client.recall(
bank_id="my-bank",
query="What did Bob mention about the new API version?",
max_tokens=1024,
budget="mid",
include_chunks=True,
)
# Generate analysis
reflect = client.reflect(
bank_id="my-bank",
query="Should we deprecate the old authentication flow?",
budget="high",
)
print(reflect.text)
print(reflect.based_on)
When to Use Recall vs. Reflect
Choose the appropriate operation based on your downstream requirements.
Use Recall when you need:
- Raw facts for external LLM context windows
- Data extraction without interpretation
- Maximum retrieval speed with token budget control
Use Reflect when you need:
- Decision-making recommendations
- Answers framed by specific personality traits or policies
- Citations proving how the conclusion was reached
Summary
- Recall runs parallel semantic, keyword, graph, and temporal searches, fusing results with Reciprocal Rank Fusion and returning raw memories until a token budget exhausts.
- Reflect executes an agentic tool loop (mental-model → observation → recall → expand) that respects bank dispositions and directives to synthesize cited answers.
- Source locations: Recall endpoint at
hindsight-control-plane/src/app/api/recall/route.ts, Reflect endpoint athindsight-control-plane/src/app/api/reflect/route.ts, with SDK methods inhindsight-clients/typescript/src/index.ts. - Documentation: See
hindsight-docs/versioned_docs/version-0.4/developer/retrieval.mdfor Recall architecture andhindsight-docs/versioned_docs/version-0.4/developer/reflect.mdxfor Reflect reasoning details.
Frequently Asked Questions
Can I use recall and reflect together in a single workflow?
Yes. Many applications use Recall to populate a context window for a downstream LLM, while using Reflect when the application requires an interpreted, personality-driven response. Since Reflect internally calls raw recall as part of its tool hierarchy, you can also trace how the reasoning system accessed underlying facts.
How does the token budget affect recall results?
The max_tokens parameter in Recall acts as a hard stop condition. The system fuses and reranks results from all four search strategies, then iteratively adds memories to the response until the cumulative token count approaches the budget. Setting budget: 'high' increases search depth before truncation, while budget: 'low' prioritizes speed.
What are bank dispositions in reflect operations?
Bank dispositions are configurable personality traits stored in the Hindsight memory bank that control how Reflect interprets evidence. Options include skepticism (requiring stronger evidence), literalism (strict adherence to text), and empathy (favoring human-centric interpretations). These dispositions shape the agentic reasoning loop but do not affect Recall, which returns raw data regardless of context.
Is reflect slower than recall due to the reasoning loop?
Yes. Reflect incurs additional latency because it executes multiple tool calls in a hierarchical sequence and applies cross-encoder reranking and disposition filtering during synthesis. Recall is optimized for single-round retrieval with parallel search execution, making it faster for simple fact lookup while Reflect trades speed for interpretive depth.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →