How to Filter Semantic Search Results by Wing and Room in MemPalace
To filter semantic search results by wing and room in MemPalace, pass the optional wing and room parameters to the search_memories() function or the CLI, which constructs a ChromaDB where filter that constrains the vector search to matching metadata before re-ranking occurs.
MemPalace organizes user data using a hierarchical metadata system where every stored item (called a drawer) is tagged with a wing (top-level category) and room (sub-category). When you perform semantic search, you can scope results to specific wings, rooms, or combinations of both by leveraging the filter builder in mempalace/searcher.py. This guide explains how the filtering mechanism works and how to use it across the Python API and command-line interface.
Understanding Wing and Room Metadata
In MemPalace, the wing represents a high-level container such as a project, person, or topic, while the room groups related drawers within that wing. These tags are embedded into every drawer's metadata during the mining process and stored in ChromaDB. When you query the palace, the system can pre-filter the candidate pool using these metadata fields before applying semantic similarity or hybrid re-ranking.
How the Where Filter Works Under the Hood
The core filtering logic resides in mempalace/searcher.py within the build_where_filter function. This helper constructs ChromaDB-compatible query filters based on which parameters you provide:
# mempalace/searcher.py – building the filter
def build_where_filter(wing: str = None, room: str = None) -> dict:
"""Build ChromaDB where filter for wing/room filtering."""
if wing and room:
return {"$and": [{"wing": wing}, {"room": room}]}
elif wing:
return {"wing": wing}
elif room:
return {"room": room}
return {}
When both wing and room are specified, the function returns a $and logical operator ensuring both conditions must match. If only one parameter is provided, it returns a simple key-value filter for that field. An empty dictionary (no filter) allows all drawers to be considered.
This filter is injected into the retrieval stage at two entry points in mempalace/searcher.py:
- Lines 20-22 in the
search()function (CLI entry point) - Lines 81-82 in the
search_memories()function (programmatic API)
Because filtering happens at the retrieval stage within ChromaDB (col.query), the subsequent hybrid re-ranking step (_hybrid_rank) operates only on the pre-filtered candidate pool. This guarantees that final relevance scores reflect only drawers belonging to your specified wing and/or room.
Filtering via the Python API
Use the search_memories function to programmatically filter results. Import the function from mempalace.searcher and provide the wing and/or room arguments alongside your query.
from mempalace.searcher import search_memories
results = search_memories(
query="how did I greet the new client",
palace_path="/path/to/palace",
wing="wing_user", # ← filter on wing
room="room_profile", # ← filter on room (optional)
n_results=5,
)
print(results["results"])
If the underlying ChromaDB collection supports native where clauses, the filter executes within the database. Otherwise, MemPalace falls back to _query_drawers_with_filter_fallback, which fetches a superset and applies the filter on the Python side to maintain consistent behavior.
Filtering via the CLI
The mempalace command-line tool exposes the same filtering capabilities through --wing and --room flags. These arguments parse into the same build_where_filter call used by the Python API.
Filter by wing only:
# Show up to 10 matches from the “wing_project” wing, any room
mempalace search \
--query "refactor authentication flow" \
--wing wing_project \
--n-results 10
Filter by both wing and room:
# Show matches from a specific room inside a wing
mempalace search \
--query "daily stand-up notes" \
--wing wing_team \
--room room_meeting
Both commands pass the parsed arguments to the search() function in mempalace/searcher.py, ensuring the filter logic remains consistent across interfaces.
Filter Propagation Through the Architecture
Understanding how the filter travels through the codebase helps debug issues and extend functionality:
mempalace/searcher.py: Containsbuild_where_filter, the publicsearch()function for CLI usage, andsearch_memories()for programmatic access.mempalace/mcp_server.py: Exposessearch_memoriesas themempalace_searchtool for MCP-compatible agents, demonstrating filter propagation to the Model Context Protocol layer.mempalace/palace.py: Provides_open_collection_or_explainand low-level collection helpers that ultimately invokecol.query.mempalace/backends/chroma.py: Implements thewhereclause handling specific to ChromaDB's query syntax.
This architecture ensures that CLI → API → Backend all respect the same wing/room constraints, creating a unified filtering experience regardless of how you interact with MemPalace.
Summary
- Metadata structure: Wings are top-level categories; rooms are sub-categories; both are stored in ChromaDB metadata.
- Filter construction: The
build_where_filterfunction inmempalace/searcher.pytranslates Python arguments into ChromaDBwhereclauses, using$andlogic when both wing and room are specified. - Retrieval-stage filtering: Filters apply before hybrid re-ranking, ensuring only relevant drawers are scored.
- Multiple interfaces: Use
search_memories()in Python or--wing/--roomflags in the CLI; both leverage the same underlying filter builder. - Fallback support: If the backend lacks native filtering,
_query_drawers_with_filter_fallbackhandles filtering in Python.
Frequently Asked Questions
Can I filter by room without specifying a wing?
Yes. The build_where_filter function in mempalace/searcher.py supports filtering by room independently. Pass only the room parameter to the API or CLI, and the function returns {"room": "your_room_value"}, which matches drawers across all wings that reside in that specific room.
What happens if no drawers match the wing and room filters?
If the ChromaDB where clause returns no matches, the col.query call yields an empty result set. The hybrid re-ranking step receives no candidates, and both the Python API and CLI return empty results. No error is raised; the response simply contains zero matches.
Does filtering affect the semantic similarity scores?
Filtering occurs before the hybrid re-ranking stage (_hybrid_rank). The scores reflect only the semantic similarity and keyword overlap of drawers within your filtered subset. This means scores are calculated relative to the constrained candidate pool, not the entire palace.
How do I search across multiple wings or rooms simultaneously?
The current implementation in mempalace/searcher.py supports single-value equality matching for wing and room. To search across multiple wings, you would need to execute separate queries for each wing (or room) and merge results manually, as the build_where_filter does not generate $or operators or list-based $in clauses by default.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →