How LanceDB's Memory System Enables Hybrid Retrieval for Schemas and Queries in WrenAI
WrenAI leverages a LanceDB-backed MemoryStore that automatically selects between plain-text schema descriptions for small manifests and semantic vector search for large schemas, while retrieving query history exclusively via embedding-based similarity search.
WrenAI implements an intelligent hybrid retrieval architecture using LanceDB to manage both database schemas and natural language query history. By storing schema fragments and past queries as vector embeddings in dedicated LanceDB tables, the system optimizes LLM context window usage through fast Approximate Nearest Neighbor (ANN) searches. This approach, centered in the MemoryStore class, seamlessly blends deterministic text retrieval with semantic similarity to balance latency and relevance according to the Canner/WrenAI source code.
Dual-Table LanceDB Architecture
The memory system maintains two distinct tables within LanceDB to support hybrid operations. In core/wren/src/wren/memory/store.py, the MemoryStore class defines _SCHEMA_TABLE = "schema_items" for manifest fragments and _QUERY_TABLE = "query_history" for natural language to SQL pairs.
Both tables utilize fixed-size vector columns created with PyArrow specifications (pa.list_(pa.float32, dim)), enabling LanceDB to perform efficient ANN searches across millions of embeddings. As defined in lines 26-28 of the store implementation, this schema ensures that all stored vectors share identical dimensions, which is critical for consistent similarity calculations.
Hybrid Schema Retrieval Logic
The MemoryStore.get_context method implements the core hybrid decision logic, automatically choosing between plain-text retrieval and semantic search based on a configurable size threshold.
Full-Text Fallback for Small Schemas
When the rendered schema description falls below the default threshold of approximately 4 KB, the system employs strategy="full". The method calls describe_schema to generate a complete plain-text representation of the manifest and returns it directly without querying the vector database.
This bypass eliminates embedding computation and vector search latency for small schemas, providing immediate context to the LLM. According to the source code in core/wren/src/wren/memory/store.py (lines 78-82), this path returns the raw schema description when the content size remains under the threshold limit.
Vector Search for Large Schemas
For schemas exceeding the size threshold, the system switches to strategy="search" and delegates to _search_schema. This method embeds the user query using a cached sentence-transformers model, then executes table.search against the schema_items LanceDB table.
The search supports optional filters including mdl_hash, item_type, and model_name, allowing precise retrieval of relevant schema fragments. As implemented in lines 91-107 of core/wren/src/wren/memory/store.py, the system returns the top-k most similar schema items based on vector similarity, significantly reducing context window usage while preserving semantic relevance.
Query History Semantic Recall
Unlike schema retrieval, query history recall operates exclusively through semantic search without a plain-text fallback. The recall_queries method searches the query_history table using vector similarity to find past NL→SQL pairs relevant to the current user question.
This approach is optimal because query history is never rendered as a single text block. The method accepts filters such as datasource to narrow results to specific database connections, returning the most semantically similar historical queries as context for the current generation task.
Embedding Management and Schema Indexing
The memory system handles embedding generation through lazy initialization. The first time vector computation is required, get_embedding_function instantiates a sentence-transformers model and caches both the embedding function and vector dimension. As shown in lines 103-119 of core/wren/src/wren/memory/store.py, subsequent calls reuse these cached components, ensuring all vectors maintain consistent dimensions.
When indexing a new manifest via index_schema, the system extracts individual schema items, computes their embeddings, and upserts them into the schema_items table. Seed NL-SQL pairs undergo the same process, populating the query_history table with pre-computed vectors that enable fast retrieval during runtime.
Code Examples
The following examples demonstrate the hybrid retrieval patterns:
# Automatic strategy selection based on schema size
from wren.memory.store import MemoryStore
store = MemoryStore()
manifest = {...} # Parsed manifest dictionary
user_query = "list all customers"
# Returns either full text or search results based on threshold
context = store.get_context(manifest, user_query, limit=5)
if context["strategy"] == "full":
# Small schema: direct text description
schema_text = context["schema"]
else:
# Large schema: semantic search results
for hit in context["results"]:
print(hit["item_name"], hit["text"])
# Semantic recall of historical queries
store = MemoryStore()
similar = store.recall_queries(
"show orders from last month",
limit=3,
datasource="postgres"
)
for q in similar:
print(q["nl_query"], "→", q["sql_query"])
Summary
- Dual-table architecture: LanceDB stores schema items and query history in separate tables (
schema_itemsandquery_history) with fixed-size float32 vector columns configured incore/wren/src/wren/memory/store.py. - Threshold-based switching: The
get_contextmethod automatically selects plain-text retrieval for schemas under ~4 KB and semantic search for larger manifests. - Lazy embedding initialization: Sentence-transformers models load on first use via
get_embedding_functionand cache their dimensions to ensure vector consistency across the system. - Pure semantic recall: Query history retrieval always uses embedding-based search without plain-text fallback, optimized for finding similar past questions via
recall_queries. - Source locations: Core logic resides in
core/wren/src/wren/memory/store.py, with supporting utilities inembeddings.pyandschema_indexer.py.
Frequently Asked Questions
What triggers the switch from plain-text to semantic retrieval in WrenAI?
The MemoryStore.get_context method compares the rendered schema description size against a configurable threshold (defaulting to approximately 4 KB). When the text exceeds this limit, the system automatically switches from strategy="full" to strategy="search" and performs a vector-based lookup against the schema_items table instead of returning the complete text.
How does LanceDB store schema embeddings in WrenAI's memory system?
Schema embeddings are stored in the schema_items table using PyArrow fixed-size list columns (pa.list_(pa.float32, dim)). During the index_schema process, each schema fragment is embedded using the cached sentence-transformers function and upserted into this table alongside metadata such as item_type and model_name, as defined in the MemoryStore class.
Why doesn't query history use the same hybrid approach as schemas?
Query history retrieval via recall_queries always uses semantic search because historical NL→SQL pairs are never rendered as a single coherent text block suitable for direct LLM consumption. The system treats each query as an independent entity, retrieving the most similar past questions through pure vector similarity without a size-based fallback mechanism.
What embedding model does WrenAI use for the LanceDB memory system?
The system utilizes sentence-transformers models for generating embeddings, initialized lazily through get_embedding_function in core/wren/src/wren/memory/store.py. The specific model is instantiated on first use and cached along with its vector dimension to ensure all embeddings in the LanceDB tables maintain consistent sizes for valid ANN comparisons.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →