Customizing Entity Detection for Domain-Specific Content in MemPalace

MemPalace extracts people, projects, and topics from prose files using a customizable two-pass pipeline that supports domain-specific compounds, multilingual patterns, and configurable filtering via JSON data files.

MemPalace is an open-source knowledge management system that automatically identifies entities in your writing. Customizing entity detection for domain-specific content in MemPalace requires understanding its detection pipeline in mempalace/entity_detector.py and modifying the locale-specific JSON files or data lexicons that drive the extraction patterns.

How MemPalace Detects Entities

The detection pipeline runs in three distinct phases, implemented in mempalace/entity_detector.py. Understanding these phases reveals exactly where to inject domain-specific customizations.

Stage 1: Candidate Extraction with Language-Aware Patterns

The first pass scans text for capitalized proper-noun candidates using regexes loaded from each locale’s entity section. For English, these patterns reside in mempalace/i18n/en.json and are retrieved via mempalace.i18n.get_entity_patterns.

The get_entity_patterns function merges patterns from all requested languages and expands \b boundaries to support scripts with combining marks, handled by _script_boundary in i18n/__init__.py. This ensures accurate tokenization for languages like Devanagari or CJK scripts.

Before single-word extraction, the system runs _apply_known_systems_prepass using mempalace/data/known_systems.json. This tier-3 detection identifies multi-word product names (e.g., Claude Code, GitHub Copilot) and masks matched spans, preventing the single-word extractor from splitting these compounds.

Stage 2: Linguistic Cleanup and Filtering

After compound pre-pass masking, single-word candidates face two filtering layers:

  • Stopwords – The union of language-specific stopword lists (e.g., English stopwords from en.json).
  • COCA Content-Word Filter – A curated list in mempalace/data/coca_content_words.json containing common nouns and verbs that frequently appear capitalized but are not proper nouns (e.g., "Code", "Note").

This filter is case-insensitive and optional; if coca_content_words.json is missing, the detector degrades gracefully without failing.

Stage 3: Scoring and Classification

For each surviving candidate, score_entity invokes _build_patterns to construct per-entity regexes, then counts matches across:

  • Dialogue markers
  • Person-verb patterns
  • Pronoun proximity
  • Direct address formats
  • Project verbs
  • Version strings
  • File-reference patterns

Finally, classify_entity combines person vs. project scores, frequency, and signal diversity to assign types (person, project, uncertain). The function also supports corpus-origin reclassification into an agent_personas bucket when the origin JSON contains known agent names.

Customizing Entity Detection for Your Domain

MemPalace provides three primary customization points that require no core code changes—only JSON edits or locale additions.

Add Domain-Specific Compounds via known_systems.json

Extend mempalace/data/known_systems.json to include multi-word product names specific to your industry. The _apply_known_systems_prepass function will treat these as atomic entities, preventing false-positive splits during single-word extraction.

{
  "compounds": [
    "Acme SuperWidget",
    "Acme SuperWidget Pro",
    "NeuralNet Framework"
  ]
}

After updating this file, detect_entities automatically recognizes these compounds in subsequent runs.

Create Custom Language Patterns

To support a new language, create a <lang>.json file in mempalace/i18n/ containing an entity section with the following keys:

  • candidate_pattern: Regex for single-word proper nouns
  • multi_word_pattern: Regex for multi-word proper nouns
  • person_verb_patterns: Array of regex templates using {name} placeholder
  • pronoun_patterns: Pronoun detection regexes
  • dialogue_patterns: Dialogue marker templates with {name} placeholder
  • project_verb_patterns: Project action verb templates
  • stopwords: Array of common words to exclude

For example, to add Hindi support in mempalace/i18n/hi.json:

{
  "lang": "hi",
  "entity": {
    "candidate_pattern": "[\\u0900-\\u097F][\\u0900-\\u097F]+",
    "multi_word_pattern": "[\\u0900-\\u097F]+(?:\\s+[\\u0900-\\u097F]+)+",
    "person_verb_patterns": ["\\b{name}\\s+कहते\\b"],
    "pronoun_patterns": ["\\bवह\\b"],
    "dialogue_patterns": ["^>{name}:\\s"],
    "project_verb_patterns": ["\\bबनाते\\s+{name}\\b"],
    "stopwords": ["और", "पर", "से"]
  }
}

The loader in i18n/__init__.py merges these automatically with existing locales.

Refine Filtering with Content-Word Lists

If your domain capitalizes common terms that trigger false positives (e.g., "Server", "Database" in technical documentation), modify mempalace/data/coca_content_words.json. Add terms that should be excluded from entity detection despite their capitalization. The filter applies case-insensitively across all processed text.

Practical Implementation Examples

Scanning a Project with Multiple Languages

To detect entities using both English and Portuguese-Brazil patterns:

from mempalace.entity_detector import detect_entities, confirm_entities
from mempalace.i18n import load_lang

# Load additional locale

load_lang("pt-br")

# Gather prose files from target directory

files = scan_for_detection("/path/to/project", max_files=10)

# Detect using English + Portuguese patterns

detected = detect_entities(files, languages=("en", "pt-br"))

# Interactive confirmation (or pass yes=True to auto-accept)

confirmed = confirm_entities(detected, yes=False)
print("Confirmed entities:", confirmed)

Extending Domain Compounds Programmatically

When distributing custom configurations across teams, you can programmatically extend the known systems list before detection:

import json
from pathlib import Path

# Extend known systems for automotive domain

systems_path = Path("mempalace/data/known_systems.json")
systems = json.loads(systems_path.read_text())

# Add automotive-specific compounds

domain_compounds = [
    "Tesla Autopilot",
    "Waymo Driver",
    "GM Super Cruise"
]
systems["compounds"].extend(domain_compounds)

# Write back (or maintain as separate overlay)

systems_path.write_text(json.dumps(systems, indent=2))

Summary

  • MemPalace uses a two-pass pipeline in mempalace/entity_detector.py combining compound detection, regex-based candidate extraction, and linguistic scoring.
  • Domain-specific compounds are added to mempalace/data/known_systems.json to prevent splitting of multi-word product names.
  • Multilingual support requires creating JSON locale files in mempalace/i18n/ with regex patterns for candidates, verbs, pronouns, and stopwords.
  • Content filtering is controlled via mempalace/data/coca_content_words.json, allowing suppression of commonly capitalized but non-proper terms.
  • No code changes are required for most customizations; the get_entity_patterns loader merges configurations automatically.

Frequently Asked Questions

How do I prevent MemPalace from splitting "Machine Learning" into two separate entities?

Add "Machine Learning" to the compounds array in mempalace/data/known_systems.json. The _apply_known_systems_prepass function masks these spans before single-word extraction runs, ensuring the phrase is treated as a single atomic entity during scoring and classification.

Can I use MemPalace with non-Latin scripts like Cyrillic or Devanagari?

Yes. Create a new locale file (e.g., mempalace/i18n/ru.json for Russian) defining the candidate_pattern using Unicode ranges (e.g., [\\u0400-\\u04FF]+ for Cyrillic). The _script_boundary function in i18n/__init__.py handles word boundaries for scripts with combining marks, and get_entity_patterns merges your new patterns automatically.

Why are common words like "Code" or "Notes" being detected as entities?

These are likely passing through the COCA content-word filter. Edit mempalace/data/coca_content_words.json to include "code" and "notes" (case-insensitive). This file contains frequently capitalized common nouns that should be excluded from entity detection. Alternatively, add them to your language-specific stopwords list in the corresponding i18n/<lang>.json file.

How do I add support for domain-specific verbs that indicate project ownership?

Extend the project_verb_patterns array in your language's JSON configuration (e.g., mempalace/i18n/en.json). Use the {name} placeholder where the entity name would appear in the regex. For example, "\\bdeveloped\\s+{name}\\b" captures "developed ProjectX" as a project signal. These patterns are compiled by _build_patterns during the scoring phase.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →