# How the MemPalace Entity Registry Maps Real Names to AAAK Codes

> Discover how the MemPalace Entity Registry maps real names to AAAK codes. Learn about the onboarding process generating unique prefixes for the dialect encoder.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: internals
- Published: 2026-06-06

---

**The EntityRegistry class stores canonical names, relationships, and confidence scores, but does not store AAAK compression codes; instead, the onboarding process in [`mempalace/onboarding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/onboarding.py) generates these codes by extracting unique three-letter prefixes and persists them to `~/.mempalace/aaak_entities.md` for use by the dialect encoder.**

The MemPalace system uses a dual-registry architecture to separate semantic entity management from compression encoding. While the **EntityRegistry** maintains the authoritative source for person identities and relationship metadata, the actual **AAAK codes**—the compressed three-to-four-letter identifiers used in diary entries—are generated during the initial onboarding phase and maintained in a separate user-visible file.

## The Separation Between Entity Storage and AAAK Compression

The `EntityRegistry` class, implemented in [`mempalace/entity_registry.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/entity_registry.py), persists canonical entity data including real names, project associations, and confidence scores. However, it explicitly does **not** maintain AAAK compression mappings. This design enforces a strict separation of concerns: the registry handles identity resolution and disambiguation, while the **AAAK entity registry** handles compression-specific code mapping. When the system compresses a diary entry to AAAK format, the dialect encoder consults the pre-generated registry file rather than querying the EntityRegistry for lookups.

## How the Onboarding Process Generates AAAK Codes

During the `seed()` operation in [`mempalace/onboarding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/onboarding.py), specifically within the `_generate_aaak_bootstrap` function (lines 300-311), the system iterates over every person supplied by the user to construct the compression dictionary.

### The Prefix Extraction Algorithm

For each person in the input list, the algorithm extracts the first three characters of the name, converts them to uppercase, and assigns this as the provisional code:

```python

# _generate_aaak_bootstrap (excerpt from mempalace/onboarding.py)

entity_codes = {}
for p in people:
    name = p["name"]
    code = name[:3].upper()                # first three letters, upper-case

    # handle collisions

    while code in entity_codes.values():
        code = name[:4].upper()            # fall back to four letters

    entity_codes[name] = code

```

### Collision Handling and Uniqueness

If two names share the same three-letter prefix (for example, "Riley" and "Ripley"), the algorithm extends the prefix length until uniqueness is achieved. The `while` loop increments the character slice to four letters, and potentially beyond, ensuring no duplicate codes exist in the final mapping. This guarantees that every person receives a unique AAAK identifier regardless of naming collisions.

## The AAAK Entity Registry File Structure

After generating the mapping, the onboarding process writes each entry to `~/.mempalace/aaak_entities.md` using a structured format. Each line contains the code, an equals sign, the full name, and optionally the relationship in parentheses:

```text
  RIL=Riley (daughter)
  MAR=Marcus (colleague)

```

This file serves as the authoritative lookup table for the compression dialect. The space-padded formatting ensures consistent readability while the simple `CODE=Name` structure allows for quick parsing by the encoder.

## How the Dialect Encoder Uses the Mapping

The **dialect encoder** ([`mempalace/dialect.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/dialect.py)) reads the AAAK entity registry when compressing diary entries into AAAK format. It builds an in-memory representation of the `entity_codes` dictionary from [`aaak_entities.md`](https://github.com/MemPalace/mempalace/blob/main/aaak_entities.md), then substitutes full names with their corresponding codes during the conversion process. The EntityRegistry is not queried during compression; the dialect operates solely against the pre-computed AAAK registry, ensuring fast O(1) lookup and consistent encoding.

## Code Example: Generating and Using AAAK Codes

The following example demonstrates the onboarding phase that creates the registry:

```python
from mempalace.entity_registry import EntityRegistry

registry = EntityRegistry.load()
registry.seed(mode="personal",
              people=[{"name": "Riley", "relationship": "daughter", "context": "personal"}],
              projects=[],
              aliases={})

# This creates ~/.mempalace/aaak_entities.md containing:

#   RIL=Riley (daughter)

```

When encoding a diary entry, the dialect substitutes the name with the generated code:

```python

# Example AAAK diary entry using the mapped code

# entity_codes dict contains: {"Riley": "RIL"}

aaak_entry = "SESSION:2026-04-04|built.palace.graph+diary.tools|ALC.req:agent.diaries.in.aaak|★★★|RIL"

```

## Summary

- The **EntityRegistry** maintains semantic entity data but does not store AAAK compression codes.
- The **onboarding process** generates unique AAAK codes by extracting three-letter prefixes from names, extending to four letters when necessary to avoid collisions.
- Codes are persisted to **`~/.mempalace/aaak_entities.md`** in a structured text format.
- The **dialect encoder** utilizes this separate registry file for compression, never querying the EntityRegistry for code lookups.

## Frequently Asked Questions

### Does EntityRegistry store AAAK codes internally?

No. The `EntityRegistry` class in [`mempalace/entity_registry.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/entity_registry.py) stores canonical names, relationships, confidence scores, and aliases, but it explicitly excludes AAAK compression mappings. These codes exist only in the separate [`aaak_entities.md`](https://github.com/MemPalace/mempalace/blob/main/aaak_entities.md) file generated during the onboarding phase.

### How does the system handle name collisions during AAAK generation?

The `_generate_aaak_bootstrap` function in [`mempalace/onboarding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/onboarding.py) detects collisions using a `while` loop that checks against existing values in the `entity_codes` dictionary. When a duplicate three-letter prefix is found, the algorithm extends the prefix to four characters (or more) until the code becomes unique.

### Where are AAAK codes persisted?

AAAK codes are written to **`~/.mempalace/aaak_entities.md`**, a user-visible markdown file in the MemPalace configuration directory. Each line follows the format `CODE=Name (relationship)`, allowing both human inspection and machine parsing by the dialect encoder.

### What happens if I add new people after onboarding?

New people added after the initial onboarding require regeneration of the AAAK entity registry, as the current implementation generates the compression mapping once during the `seed()` operation. The dialect encoder relies on the static [`aaak_entities.md`](https://github.com/MemPalace/mempalace/blob/main/aaak_entities.md) file for lookups, so updates to the entity list necessitate running the onboarding process again to rebuild the code mappings.