How the MemPalace Entity Registry Maps Real Names to AAAK Codes
The EntityRegistry class stores canonical names, relationships, and confidence scores, but does not store AAAK compression codes; instead, the onboarding process in mempalace/onboarding.py generates these codes by extracting unique three-letter prefixes and persists them to ~/.mempalace/aaak_entities.md for use by the dialect encoder.
The MemPalace system uses a dual-registry architecture to separate semantic entity management from compression encoding. While the EntityRegistry maintains the authoritative source for person identities and relationship metadata, the actual AAAK codes—the compressed three-to-four-letter identifiers used in diary entries—are generated during the initial onboarding phase and maintained in a separate user-visible file.
The Separation Between Entity Storage and AAAK Compression
The EntityRegistry class, implemented in mempalace/entity_registry.py, persists canonical entity data including real names, project associations, and confidence scores. However, it explicitly does not maintain AAAK compression mappings. This design enforces a strict separation of concerns: the registry handles identity resolution and disambiguation, while the AAAK entity registry handles compression-specific code mapping. When the system compresses a diary entry to AAAK format, the dialect encoder consults the pre-generated registry file rather than querying the EntityRegistry for lookups.
How the Onboarding Process Generates AAAK Codes
During the seed() operation in mempalace/onboarding.py, specifically within the _generate_aaak_bootstrap function (lines 300-311), the system iterates over every person supplied by the user to construct the compression dictionary.
The Prefix Extraction Algorithm
For each person in the input list, the algorithm extracts the first three characters of the name, converts them to uppercase, and assigns this as the provisional code:
# _generate_aaak_bootstrap (excerpt from mempalace/onboarding.py)
entity_codes = {}
for p in people:
name = p["name"]
code = name[:3].upper() # first three letters, upper-case
# handle collisions
while code in entity_codes.values():
code = name[:4].upper() # fall back to four letters
entity_codes[name] = code
Collision Handling and Uniqueness
If two names share the same three-letter prefix (for example, "Riley" and "Ripley"), the algorithm extends the prefix length until uniqueness is achieved. The while loop increments the character slice to four letters, and potentially beyond, ensuring no duplicate codes exist in the final mapping. This guarantees that every person receives a unique AAAK identifier regardless of naming collisions.
The AAAK Entity Registry File Structure
After generating the mapping, the onboarding process writes each entry to ~/.mempalace/aaak_entities.md using a structured format. Each line contains the code, an equals sign, the full name, and optionally the relationship in parentheses:
RIL=Riley (daughter)
MAR=Marcus (colleague)
This file serves as the authoritative lookup table for the compression dialect. The space-padded formatting ensures consistent readability while the simple CODE=Name structure allows for quick parsing by the encoder.
How the Dialect Encoder Uses the Mapping
The dialect encoder (mempalace/dialect.py) reads the AAAK entity registry when compressing diary entries into AAAK format. It builds an in-memory representation of the entity_codes dictionary from aaak_entities.md, then substitutes full names with their corresponding codes during the conversion process. The EntityRegistry is not queried during compression; the dialect operates solely against the pre-computed AAAK registry, ensuring fast O(1) lookup and consistent encoding.
Code Example: Generating and Using AAAK Codes
The following example demonstrates the onboarding phase that creates the registry:
from mempalace.entity_registry import EntityRegistry
registry = EntityRegistry.load()
registry.seed(mode="personal",
people=[{"name": "Riley", "relationship": "daughter", "context": "personal"}],
projects=[],
aliases={})
# This creates ~/.mempalace/aaak_entities.md containing:
# RIL=Riley (daughter)
When encoding a diary entry, the dialect substitutes the name with the generated code:
# Example AAAK diary entry using the mapped code
# entity_codes dict contains: {"Riley": "RIL"}
aaak_entry = "SESSION:2026-04-04|built.palace.graph+diary.tools|ALC.req:agent.diaries.in.aaak|★★★|RIL"
Summary
- The EntityRegistry maintains semantic entity data but does not store AAAK compression codes.
- The onboarding process generates unique AAAK codes by extracting three-letter prefixes from names, extending to four letters when necessary to avoid collisions.
- Codes are persisted to
~/.mempalace/aaak_entities.mdin a structured text format. - The dialect encoder utilizes this separate registry file for compression, never querying the EntityRegistry for code lookups.
Frequently Asked Questions
Does EntityRegistry store AAAK codes internally?
No. The EntityRegistry class in mempalace/entity_registry.py stores canonical names, relationships, confidence scores, and aliases, but it explicitly excludes AAAK compression mappings. These codes exist only in the separate aaak_entities.md file generated during the onboarding phase.
How does the system handle name collisions during AAAK generation?
The _generate_aaak_bootstrap function in mempalace/onboarding.py detects collisions using a while loop that checks against existing values in the entity_codes dictionary. When a duplicate three-letter prefix is found, the algorithm extends the prefix to four characters (or more) until the code becomes unique.
Where are AAAK codes persisted?
AAAK codes are written to ~/.mempalace/aaak_entities.md, a user-visible markdown file in the MemPalace configuration directory. Each line follows the format CODE=Name (relationship), allowing both human inspection and machine parsing by the dialect encoder.
What happens if I add new people after onboarding?
New people added after the initial onboarding require regeneration of the AAAK entity registry, as the current implementation generates the compression mapping once during the seed() operation. The dialect encoder relies on the static aaak_entities.md file for lookups, so updates to the entity list necessitate running the onboarding process again to rebuild the code mappings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →