Customizing Entity Detection for Domain-Specific Content in MemPalace
MemPalace extracts people, projects, and topics from prose files using a customizable two-pass pipeline that supports domain-specific compounds, multilingual patterns, and configurable filtering via JSON data files.
MemPalace is an open-source knowledge management system that automatically identifies entities in your writing. Customizing entity detection for domain-specific content in MemPalace requires understanding its detection pipeline in mempalace/entity_detector.py and modifying the locale-specific JSON files or data lexicons that drive the extraction patterns.
How MemPalace Detects Entities
The detection pipeline runs in three distinct phases, implemented in mempalace/entity_detector.py. Understanding these phases reveals exactly where to inject domain-specific customizations.
Stage 1: Candidate Extraction with Language-Aware Patterns
The first pass scans text for capitalized proper-noun candidates using regexes loaded from each locale’s entity section. For English, these patterns reside in mempalace/i18n/en.json and are retrieved via mempalace.i18n.get_entity_patterns.
The get_entity_patterns function merges patterns from all requested languages and expands \b boundaries to support scripts with combining marks, handled by _script_boundary in i18n/__init__.py. This ensures accurate tokenization for languages like Devanagari or CJK scripts.
Before single-word extraction, the system runs _apply_known_systems_prepass using mempalace/data/known_systems.json. This tier-3 detection identifies multi-word product names (e.g., Claude Code, GitHub Copilot) and masks matched spans, preventing the single-word extractor from splitting these compounds.
Stage 2: Linguistic Cleanup and Filtering
After compound pre-pass masking, single-word candidates face two filtering layers:
- Stopwords – The union of language-specific stopword lists (e.g., English stopwords from
en.json). - COCA Content-Word Filter – A curated list in
mempalace/data/coca_content_words.jsoncontaining common nouns and verbs that frequently appear capitalized but are not proper nouns (e.g., "Code", "Note").
This filter is case-insensitive and optional; if coca_content_words.json is missing, the detector degrades gracefully without failing.
Stage 3: Scoring and Classification
For each surviving candidate, score_entity invokes _build_patterns to construct per-entity regexes, then counts matches across:
- Dialogue markers
- Person-verb patterns
- Pronoun proximity
- Direct address formats
- Project verbs
- Version strings
- File-reference patterns
Finally, classify_entity combines person vs. project scores, frequency, and signal diversity to assign types (person, project, uncertain). The function also supports corpus-origin reclassification into an agent_personas bucket when the origin JSON contains known agent names.
Customizing Entity Detection for Your Domain
MemPalace provides three primary customization points that require no core code changes—only JSON edits or locale additions.
Add Domain-Specific Compounds via known_systems.json
Extend mempalace/data/known_systems.json to include multi-word product names specific to your industry. The _apply_known_systems_prepass function will treat these as atomic entities, preventing false-positive splits during single-word extraction.
{
"compounds": [
"Acme SuperWidget",
"Acme SuperWidget Pro",
"NeuralNet Framework"
]
}
After updating this file, detect_entities automatically recognizes these compounds in subsequent runs.
Create Custom Language Patterns
To support a new language, create a <lang>.json file in mempalace/i18n/ containing an entity section with the following keys:
candidate_pattern: Regex for single-word proper nounsmulti_word_pattern: Regex for multi-word proper nounsperson_verb_patterns: Array of regex templates using{name}placeholderpronoun_patterns: Pronoun detection regexesdialogue_patterns: Dialogue marker templates with{name}placeholderproject_verb_patterns: Project action verb templatesstopwords: Array of common words to exclude
For example, to add Hindi support in mempalace/i18n/hi.json:
{
"lang": "hi",
"entity": {
"candidate_pattern": "[\\u0900-\\u097F][\\u0900-\\u097F]+",
"multi_word_pattern": "[\\u0900-\\u097F]+(?:\\s+[\\u0900-\\u097F]+)+",
"person_verb_patterns": ["\\b{name}\\s+कहते\\b"],
"pronoun_patterns": ["\\bवह\\b"],
"dialogue_patterns": ["^>{name}:\\s"],
"project_verb_patterns": ["\\bबनाते\\s+{name}\\b"],
"stopwords": ["और", "पर", "से"]
}
}
The loader in i18n/__init__.py merges these automatically with existing locales.
Refine Filtering with Content-Word Lists
If your domain capitalizes common terms that trigger false positives (e.g., "Server", "Database" in technical documentation), modify mempalace/data/coca_content_words.json. Add terms that should be excluded from entity detection despite their capitalization. The filter applies case-insensitively across all processed text.
Practical Implementation Examples
Scanning a Project with Multiple Languages
To detect entities using both English and Portuguese-Brazil patterns:
from mempalace.entity_detector import detect_entities, confirm_entities
from mempalace.i18n import load_lang
# Load additional locale
load_lang("pt-br")
# Gather prose files from target directory
files = scan_for_detection("/path/to/project", max_files=10)
# Detect using English + Portuguese patterns
detected = detect_entities(files, languages=("en", "pt-br"))
# Interactive confirmation (or pass yes=True to auto-accept)
confirmed = confirm_entities(detected, yes=False)
print("Confirmed entities:", confirmed)
Extending Domain Compounds Programmatically
When distributing custom configurations across teams, you can programmatically extend the known systems list before detection:
import json
from pathlib import Path
# Extend known systems for automotive domain
systems_path = Path("mempalace/data/known_systems.json")
systems = json.loads(systems_path.read_text())
# Add automotive-specific compounds
domain_compounds = [
"Tesla Autopilot",
"Waymo Driver",
"GM Super Cruise"
]
systems["compounds"].extend(domain_compounds)
# Write back (or maintain as separate overlay)
systems_path.write_text(json.dumps(systems, indent=2))
Summary
- MemPalace uses a two-pass pipeline in
mempalace/entity_detector.pycombining compound detection, regex-based candidate extraction, and linguistic scoring. - Domain-specific compounds are added to
mempalace/data/known_systems.jsonto prevent splitting of multi-word product names. - Multilingual support requires creating JSON locale files in
mempalace/i18n/with regex patterns for candidates, verbs, pronouns, and stopwords. - Content filtering is controlled via
mempalace/data/coca_content_words.json, allowing suppression of commonly capitalized but non-proper terms. - No code changes are required for most customizations; the
get_entity_patternsloader merges configurations automatically.
Frequently Asked Questions
How do I prevent MemPalace from splitting "Machine Learning" into two separate entities?
Add "Machine Learning" to the compounds array in mempalace/data/known_systems.json. The _apply_known_systems_prepass function masks these spans before single-word extraction runs, ensuring the phrase is treated as a single atomic entity during scoring and classification.
Can I use MemPalace with non-Latin scripts like Cyrillic or Devanagari?
Yes. Create a new locale file (e.g., mempalace/i18n/ru.json for Russian) defining the candidate_pattern using Unicode ranges (e.g., [\\u0400-\\u04FF]+ for Cyrillic). The _script_boundary function in i18n/__init__.py handles word boundaries for scripts with combining marks, and get_entity_patterns merges your new patterns automatically.
Why are common words like "Code" or "Notes" being detected as entities?
These are likely passing through the COCA content-word filter. Edit mempalace/data/coca_content_words.json to include "code" and "notes" (case-insensitive). This file contains frequently capitalized common nouns that should be excluded from entity detection. Alternatively, add them to your language-specific stopwords list in the corresponding i18n/<lang>.json file.
How do I add support for domain-specific verbs that indicate project ownership?
Extend the project_verb_patterns array in your language's JSON configuration (e.g., mempalace/i18n/en.json). Use the {name} placeholder where the entity name would appear in the regex. For example, "\\bdeveloped\\s+{name}\\b" captures "developed ProjectX" as a project signal. These patterns are compiled by _build_patterns during the scoring phase.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →