# Customizing Entity Detection for Domain-Specific Content in MemPalace

> Customize MemPalace entity detection for your domain. Tailor people, project and topic extraction with our flexible two-pass pipeline supporting multilingual patterns and JSON filtering.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: how-to-guide
- Published: 2026-06-07

---

**MemPalace extracts people, projects, and topics from prose files using a customizable two-pass pipeline that supports domain-specific compounds, multilingual patterns, and configurable filtering via JSON data files.**

MemPalace is an open-source knowledge management system that automatically identifies entities in your writing. Customizing entity detection for domain-specific content in MemPalace requires understanding its detection pipeline in [`mempalace/entity_detector.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/entity_detector.py) and modifying the locale-specific JSON files or data lexicons that drive the extraction patterns.

## How MemPalace Detects Entities

The detection pipeline runs in three distinct phases, implemented in [`mempalace/entity_detector.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/entity_detector.py). Understanding these phases reveals exactly where to inject domain-specific customizations.

### Stage 1: Candidate Extraction with Language-Aware Patterns

The first pass scans text for capitalized proper-noun candidates using regexes loaded from each locale’s `entity` section. For English, these patterns reside in [`mempalace/i18n/en.json`](https://github.com/MemPalace/mempalace/blob/main/mempalace/i18n/en.json) and are retrieved via `mempalace.i18n.get_entity_patterns`.

The `get_entity_patterns` function merges patterns from all requested languages and expands `\b` boundaries to support scripts with combining marks, handled by `_script_boundary` in [`i18n/__init__.py`](https://github.com/MemPalace/mempalace/blob/main/i18n/__init__.py). This ensures accurate tokenization for languages like Devanagari or CJK scripts.

Before single-word extraction, the system runs `_apply_known_systems_prepass` using [`mempalace/data/known_systems.json`](https://github.com/MemPalace/mempalace/blob/main/mempalace/data/known_systems.json). This tier-3 detection identifies multi-word product names (e.g., *Claude Code*, *GitHub Copilot*) and masks matched spans, preventing the single-word extractor from splitting these compounds.

### Stage 2: Linguistic Cleanup and Filtering

After compound pre-pass masking, single-word candidates face two filtering layers:

- **Stopwords** – The union of language-specific stopword lists (e.g., English stopwords from [`en.json`](https://github.com/MemPalace/mempalace/blob/main/en.json)).
- **COCA Content-Word Filter** – A curated list in [`mempalace/data/coca_content_words.json`](https://github.com/MemPalace/mempalace/blob/main/mempalace/data/coca_content_words.json) containing common nouns and verbs that frequently appear capitalized but are not proper nouns (e.g., "Code", "Note").

This filter is case-insensitive and optional; if [`coca_content_words.json`](https://github.com/MemPalace/mempalace/blob/main/coca_content_words.json) is missing, the detector degrades gracefully without failing.

### Stage 3: Scoring and Classification

For each surviving candidate, `score_entity` invokes `_build_patterns` to construct per-entity regexes, then counts matches across:

- Dialogue markers
- Person-verb patterns
- Pronoun proximity
- Direct address formats
- Project verbs
- Version strings
- File-reference patterns

Finally, `classify_entity` combines person vs. project scores, frequency, and signal diversity to assign types (`person`, `project`, `uncertain`). The function also supports corpus-origin reclassification into an `agent_personas` bucket when the origin JSON contains known agent names.

## Customizing Entity Detection for Your Domain

MemPalace provides three primary customization points that require no core code changes—only JSON edits or locale additions.

### Add Domain-Specific Compounds via known_systems.json

Extend [`mempalace/data/known_systems.json`](https://github.com/MemPalace/mempalace/blob/main/mempalace/data/known_systems.json) to include multi-word product names specific to your industry. The `_apply_known_systems_prepass` function will treat these as atomic entities, preventing false-positive splits during single-word extraction.

```json
{
  "compounds": [
    "Acme SuperWidget",
    "Acme SuperWidget Pro",
    "NeuralNet Framework"
  ]
}

```

After updating this file, `detect_entities` automatically recognizes these compounds in subsequent runs.

### Create Custom Language Patterns

To support a new language, create a `<lang>.json` file in `mempalace/i18n/` containing an `entity` section with the following keys:

- `candidate_pattern`: Regex for single-word proper nouns
- `multi_word_pattern`: Regex for multi-word proper nouns
- `person_verb_patterns`: Array of regex templates using `{name}` placeholder
- `pronoun_patterns`: Pronoun detection regexes
- `dialogue_patterns`: Dialogue marker templates with `{name}` placeholder
- `project_verb_patterns`: Project action verb templates
- `stopwords`: Array of common words to exclude

For example, to add Hindi support in [`mempalace/i18n/hi.json`](https://github.com/MemPalace/mempalace/blob/main/mempalace/i18n/hi.json):

```json
{
  "lang": "hi",
  "entity": {
    "candidate_pattern": "[\\u0900-\\u097F][\\u0900-\\u097F]+",
    "multi_word_pattern": "[\\u0900-\\u097F]+(?:\\s+[\\u0900-\\u097F]+)+",
    "person_verb_patterns": ["\\b{name}\\s+कहते\\b"],
    "pronoun_patterns": ["\\bवह\\b"],
    "dialogue_patterns": ["^>{name}:\\s"],
    "project_verb_patterns": ["\\bबनाते\\s+{name}\\b"],
    "stopwords": ["और", "पर", "से"]
  }
}

```

The loader in [`i18n/__init__.py`](https://github.com/MemPalace/mempalace/blob/main/i18n/__init__.py) merges these automatically with existing locales.

### Refine Filtering with Content-Word Lists

If your domain capitalizes common terms that trigger false positives (e.g., "Server", "Database" in technical documentation), modify [`mempalace/data/coca_content_words.json`](https://github.com/MemPalace/mempalace/blob/main/mempalace/data/coca_content_words.json). Add terms that should be excluded from entity detection despite their capitalization. The filter applies case-insensitively across all processed text.

## Practical Implementation Examples

### Scanning a Project with Multiple Languages

To detect entities using both English and Portuguese-Brazil patterns:

```python
from mempalace.entity_detector import detect_entities, confirm_entities
from mempalace.i18n import load_lang

# Load additional locale

load_lang("pt-br")

# Gather prose files from target directory

files = scan_for_detection("/path/to/project", max_files=10)

# Detect using English + Portuguese patterns

detected = detect_entities(files, languages=("en", "pt-br"))

# Interactive confirmation (or pass yes=True to auto-accept)

confirmed = confirm_entities(detected, yes=False)
print("Confirmed entities:", confirmed)

```

### Extending Domain Compounds Programmatically

When distributing custom configurations across teams, you can programmatically extend the known systems list before detection:

```python
import json
from pathlib import Path

# Extend known systems for automotive domain

systems_path = Path("mempalace/data/known_systems.json")
systems = json.loads(systems_path.read_text())

# Add automotive-specific compounds

domain_compounds = [
    "Tesla Autopilot",
    "Waymo Driver",
    "GM Super Cruise"
]
systems["compounds"].extend(domain_compounds)

# Write back (or maintain as separate overlay)

systems_path.write_text(json.dumps(systems, indent=2))

```

## Summary

- **MemPalace uses a two-pass pipeline** in [`mempalace/entity_detector.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/entity_detector.py) combining compound detection, regex-based candidate extraction, and linguistic scoring.
- **Domain-specific compounds** are added to [`mempalace/data/known_systems.json`](https://github.com/MemPalace/mempalace/blob/main/mempalace/data/known_systems.json) to prevent splitting of multi-word product names.
- **Multilingual support** requires creating JSON locale files in `mempalace/i18n/` with regex patterns for candidates, verbs, pronouns, and stopwords.
- **Content filtering** is controlled via [`mempalace/data/coca_content_words.json`](https://github.com/MemPalace/mempalace/blob/main/mempalace/data/coca_content_words.json), allowing suppression of commonly capitalized but non-proper terms.
- **No code changes are required** for most customizations; the `get_entity_patterns` loader merges configurations automatically.

## Frequently Asked Questions

### How do I prevent MemPalace from splitting "Machine Learning" into two separate entities?

Add "Machine Learning" to the `compounds` array in [`mempalace/data/known_systems.json`](https://github.com/MemPalace/mempalace/blob/main/mempalace/data/known_systems.json). The `_apply_known_systems_prepass` function masks these spans before single-word extraction runs, ensuring the phrase is treated as a single atomic entity during scoring and classification.

### Can I use MemPalace with non-Latin scripts like Cyrillic or Devanagari?

Yes. Create a new locale file (e.g., [`mempalace/i18n/ru.json`](https://github.com/MemPalace/mempalace/blob/main/mempalace/i18n/ru.json) for Russian) defining the `candidate_pattern` using Unicode ranges (e.g., `[\\u0400-\\u04FF]+` for Cyrillic). The `_script_boundary` function in [`i18n/__init__.py`](https://github.com/MemPalace/mempalace/blob/main/i18n/__init__.py) handles word boundaries for scripts with combining marks, and `get_entity_patterns` merges your new patterns automatically.

### Why are common words like "Code" or "Notes" being detected as entities?

These are likely passing through the COCA content-word filter. Edit [`mempalace/data/coca_content_words.json`](https://github.com/MemPalace/mempalace/blob/main/mempalace/data/coca_content_words.json) to include "code" and "notes" (case-insensitive). This file contains frequently capitalized common nouns that should be excluded from entity detection. Alternatively, add them to your language-specific stopwords list in the corresponding `i18n/<lang>.json` file.

### How do I add support for domain-specific verbs that indicate project ownership?

Extend the `project_verb_patterns` array in your language's JSON configuration (e.g., [`mempalace/i18n/en.json`](https://github.com/MemPalace/mempalace/blob/main/mempalace/i18n/en.json)). Use the `{name}` placeholder where the entity name would appear in the regex. For example, `"\\bdeveloped\\s+{name}\\b"` captures "developed ProjectX" as a project signal. These patterns are compiled by `_build_patterns` during the scoring phase.