How Entity Detection and Registry Handle Name Disambiguation in MemPalace

MemPalace resolves ambiguous names by flagging common English words during detection and applying regex-based context pattern matching in the registry to determine whether a token refers to a person or a concept.

MemPalace implements a deterministic, privacy-first approach to distinguishing personal names from ordinary English words that collide in text. The system separates concerns between detection and storage, using a two-module pipeline that flags potential ambiguities during extraction and resolves them via contextual analysis during lookup. This ensures that words like "ever" or "grace" are correctly identified as either personal entities or generic concepts based on surrounding context.

The Two-Module Architecture

MemPalace splits the detect-then-store workflow into two cooperating components: the detection module that scans source material, and the registry module that persists and resolves entities.

Entity Detection Phase

The entity_detector.py module scans source files and extracts candidate capitalized tokens. During this phase, it scores tokens as either person or project entities. When a candidate appears in the built-in common-English-words list, the detector marks it as potentially ambiguous before passing it to the registry.

Registry Management Phase

The entity_registry.py module maintains the persistent registry at ~/.mempalace/entity_registry.json. This file stores not only the entity mappings but also a dedicated ambiguous_flags array that tracks which words require contextual disambiguation during lookup operations.

Flagging Ambiguous Tokens

The system identifies ambiguous candidates early in the pipeline to prevent false positives when common words are used as names.

The Common English Words List

The registry defines a COMMON_ENGLISH_WORDS constant containing high-frequency English words that often collide with personal names (e.g., "ever", "may", "grace"). According to the source code in entity_registry.py (lines 31-70), this list serves as the baseline for ambiguity detection.

Persistent Flag Storage

During the seeding process, entity_registry.py (lines 37-44) copies any matching entries from the common words list into the ambiguous_flags array within the JSON registry. This persistent flag ensures that subsequent lookups recognize these tokens as requiring special handling, even across application restarts.

Context-Based Disambiguation Logic

When the registry encounters a flagged word, it employs rule-based context matching rather than statistical models to preserve user privacy.

The Lookup Method

The lookup() method in entity_registry.py (lines 48-61) checks if a requested name exists in ambiguous_flags. If the flag is present and a surrounding sentence is supplied, the method invokes the private _disambiguate() method to resolve the ambiguity before returning the result.

Pattern Matching Algorithm

The _disambiguate() method (defined in entity_registry.py lines 89-125) runs two competing sets of regex patterns against the supplied context:

  • Person-context patterns: Match name usage indicating proper nouns (e.g., r"\bhey\s+{name}\b", r"{name}\s+said")
  • Concept-context patterns: Match generic linguistic usage (e.g., r"\bhave\s+you\s+{name}\b", r"{name}\s+since")

Scoring and Resolution

The algorithm counts matches for each pattern category and compares the scores. If the person score exceeds the concept score, it returns a resolved person entry; otherwise, it returns a concept entry. When scores tie, the system falls back to the original registration (treated as a person). This guarantees that "ever" is treated as a person only when context strongly suggests a name (e.g., "Hey ever, thank-you!"), otherwise remaining a generic concept.

Practical Usage Example

The following example demonstrates how the same ambiguous word resolves differently based on context:

from mempalace.entity_registry import EntityRegistry

# Load the persisted registry (it already contains the ambiguous flags)

registry = EntityRegistry.load()

# 1️⃣ Ambiguous word used as a name

result = registry.lookup("ever", context="Hey ever, thanks for the help!")

# → {'type': 'person', 'confidence': 0.75, 'source': 'onboarding',

#    'name': 'ever', 'needs_disambiguation': False}

# 2️⃣ Same word used as a common concept

result = registry.lookup("ever", context="Have you ever tried meditation?")

# → {'type': 'concept', 'confidence': 0.80, 'source': 'inferred',

#    'name': 'ever', 'needs_disambiguation': False}

In the first case, the pattern r"\bhey\s+{name}\b" triggers a person classification. In the second, r"\bhave\s+you\s+{name}\b" indicates a concept usage.

Summary

  • MemPalace uses entity_detector.py to identify capitalized candidates and flag common English words as ambiguous during the detection phase.
  • The entity_registry.py module maintains persistent flags in ~/.mempalace/entity_registry.json and resolves ambiguities during lookup.
  • The _disambiguate() method applies competing person-context and concept-context regex patterns to determine entity type.
  • Disambiguation relies on deterministic rule-based scoring rather than machine learning models, ensuring privacy and reproducibility.

Frequently Asked Questions

How does MemPalace identify which words need disambiguation?

MemPalace references the COMMON_ENGLISH_WORDS constant defined in entity_registry.py (lines 31-70). During the seeding process (lines 37-44), any candidate entity that matches this list gets added to the ambiguous_flags array in the persistent registry, marking it for contextual evaluation during future lookups.

What happens if the context patterns produce a tie?

When the person-context patterns and concept-context patterns produce equal match counts, the algorithm falls back to the original registration default, treating the token as a person entity. This conservative approach prioritizes recognizing personal names over generic concepts when evidence is inconclusive.

Can the disambiguation logic handle multi-word names?

The current implementation in entity_registry.py focuses on single-token disambiguation using the lookup() method. Multi-word entities are typically handled upstream in entity_detector.py before reaching the registry's ambiguity resolution stage.

Where are the context patterns defined?

The regex patterns for person and concept contexts are defined directly in entity_registry.py (lines 89-125) as PERSON_CONTEXT_PATTERNS and CONCEPT_CONTEXT_PATTERNS. These patterns are interpolated with the target name at runtime to evaluate surrounding text.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →