# How Entity Detection and Registry Handle Name Disambiguation in MemPalace

> Discover how MemPalace handles name disambiguation using entity detection and registry. Learn about flagging common words and regex context matching.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: architecture
- Published: 2026-06-06

---

**MemPalace resolves ambiguous names by flagging common English words during detection and applying regex-based context pattern matching in the registry to determine whether a token refers to a person or a concept.**

MemPalace implements a deterministic, privacy-first approach to distinguishing personal names from ordinary English words that collide in text. The system separates concerns between detection and storage, using a two-module pipeline that flags potential ambiguities during extraction and resolves them via contextual analysis during lookup. This ensures that words like "ever" or "grace" are correctly identified as either personal entities or generic concepts based on surrounding context.

## The Two-Module Architecture

MemPalace splits the *detect-then-store* workflow into two cooperating components: the detection module that scans source material, and the registry module that persists and resolves entities.

### Entity Detection Phase

The **[`entity_detector.py`](https://github.com/MemPalace/mempalace/blob/main/entity_detector.py)** module scans source files and extracts candidate capitalized tokens. During this phase, it scores tokens as either *person* or *project* entities. When a candidate appears in the built-in **common-English-words list**, the detector marks it as *potentially ambiguous* before passing it to the registry.

### Registry Management Phase

The **[`entity_registry.py`](https://github.com/MemPalace/mempalace/blob/main/entity_registry.py)** module maintains the persistent registry at `~/.mempalace/entity_registry.json`. This file stores not only the entity mappings but also a dedicated `ambiguous_flags` array that tracks which words require contextual disambiguation during lookup operations.

## Flagging Ambiguous Tokens

The system identifies ambiguous candidates early in the pipeline to prevent false positives when common words are used as names.

### The Common English Words List

The registry defines a **`COMMON_ENGLISH_WORDS`** constant containing high-frequency English words that often collide with personal names (e.g., "ever", "may", "grace"). According to the source code in [`entity_registry.py`](https://github.com/MemPalace/mempalace/blob/main/entity_registry.py) (lines 31-70), this list serves as the baseline for ambiguity detection.

### Persistent Flag Storage

During the seeding process, [`entity_registry.py`](https://github.com/MemPalace/mempalace/blob/main/entity_registry.py) (lines 37-44) copies any matching entries from the common words list into the `ambiguous_flags` array within the JSON registry. This persistent flag ensures that subsequent lookups recognize these tokens as requiring special handling, even across application restarts.

## Context-Based Disambiguation Logic

When the registry encounters a flagged word, it employs rule-based context matching rather than statistical models to preserve user privacy.

### The Lookup Method

The **`lookup()`** method in [`entity_registry.py`](https://github.com/MemPalace/mempalace/blob/main/entity_registry.py) (lines 48-61) checks if a requested name exists in `ambiguous_flags`. If the flag is present and a surrounding sentence is supplied, the method invokes the private **`_disambiguate()`** method to resolve the ambiguity before returning the result.

### Pattern Matching Algorithm

The **`_disambiguate()`** method (defined in [`entity_registry.py`](https://github.com/MemPalace/mempalace/blob/main/entity_registry.py) lines 89-125) runs two competing sets of regex patterns against the supplied context:

- **Person-context patterns**: Match name usage indicating proper nouns (e.g., `r"\bhey\s+{name}\b"`, `r"{name}\s+said"`)
- **Concept-context patterns**: Match generic linguistic usage (e.g., `r"\bhave\s+you\s+{name}\b"`, `r"{name}\s+since"`)

### Scoring and Resolution

The algorithm counts matches for each pattern category and compares the scores. If the **person score** exceeds the **concept score**, it returns a resolved *person* entry; otherwise, it returns a *concept* entry. When scores tie, the system falls back to the original registration (treated as a person). This guarantees that "ever" is treated as a person only when context strongly suggests a name (e.g., "Hey ever, thank-you!"), otherwise remaining a generic concept.

## Practical Usage Example

The following example demonstrates how the same ambiguous word resolves differently based on context:

```python
from mempalace.entity_registry import EntityRegistry

# Load the persisted registry (it already contains the ambiguous flags)

registry = EntityRegistry.load()

# 1️⃣ Ambiguous word used as a name

result = registry.lookup("ever", context="Hey ever, thanks for the help!")

# → {'type': 'person', 'confidence': 0.75, 'source': 'onboarding',

#    'name': 'ever', 'needs_disambiguation': False}

# 2️⃣ Same word used as a common concept

result = registry.lookup("ever", context="Have you ever tried meditation?")

# → {'type': 'concept', 'confidence': 0.80, 'source': 'inferred',

#    'name': 'ever', 'needs_disambiguation': False}

```

In the first case, the pattern `r"\bhey\s+{name}\b"` triggers a **person** classification. In the second, `r"\bhave\s+you\s+{name}\b"` indicates a **concept** usage.

## Summary

- MemPalace uses **[`entity_detector.py`](https://github.com/MemPalace/mempalace/blob/main/entity_detector.py)** to identify capitalized candidates and flag common English words as ambiguous during the detection phase.
- The **[`entity_registry.py`](https://github.com/MemPalace/mempalace/blob/main/entity_registry.py)** module maintains persistent flags in `~/.mempalace/entity_registry.json` and resolves ambiguities during lookup.
- The **`_disambiguate()`** method applies competing **person-context** and **concept-context** regex patterns to determine entity type.
- Disambiguation relies on deterministic rule-based scoring rather than machine learning models, ensuring privacy and reproducibility.

## Frequently Asked Questions

### How does MemPalace identify which words need disambiguation?

MemPalace references the **`COMMON_ENGLISH_WORDS`** constant defined in [`entity_registry.py`](https://github.com/MemPalace/mempalace/blob/main/entity_registry.py) (lines 31-70). During the seeding process (lines 37-44), any candidate entity that matches this list gets added to the `ambiguous_flags` array in the persistent registry, marking it for contextual evaluation during future lookups.

### What happens if the context patterns produce a tie?

When the **person-context patterns** and **concept-context patterns** produce equal match counts, the algorithm falls back to the original registration default, treating the token as a **person** entity. This conservative approach prioritizes recognizing personal names over generic concepts when evidence is inconclusive.

### Can the disambiguation logic handle multi-word names?

The current implementation in [`entity_registry.py`](https://github.com/MemPalace/mempalace/blob/main/entity_registry.py) focuses on single-token disambiguation using the **`lookup()`** method. Multi-word entities are typically handled upstream in [`entity_detector.py`](https://github.com/MemPalace/mempalace/blob/main/entity_detector.py) before reaching the registry's ambiguity resolution stage.

### Where are the context patterns defined?

The regex patterns for person and concept contexts are defined directly in **[`entity_registry.py`](https://github.com/MemPalace/mempalace/blob/main/entity_registry.py)** (lines 89-125) as `PERSON_CONTEXT_PATTERNS` and `CONCEPT_CONTEXT_PATTERNS`. These patterns are interpolated with the target name at runtime to evaluate surrounding text.