# What Entity Codes Does the AAAK Compression Dialect Use? A Technical Guide to MemPalace Compression

> Discover the entity codes used by the AAAK compression dialect. This guide details its hybrid system of explicit mappings, dictionary lookups, and three-letter abbreviations for efficient entity name compression.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: deep-dive
- Published: 2026-06-06

---

**The AAAK compression dialect uses a hybrid system of explicit user-defined mappings, case-insensitive dictionary lookups, and deterministic three-letter abbreviations to compress entity names into short codes.**

The MemPalace/mempalace repository implements the AAAK dialect to compactly represent zettelkasten notes. Understanding how **entity codes** function is essential for configuring the compression pipeline, as these short identifiers replace full entity names to minimize storage while preserving semantic references.

## How Entity Codes Are Generated

The encoding logic resides in [`mempalace/dialect.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/dialect.py), specifically within the `Dialect` class and its `encode_entity` method. The dialect resolves entity names through a prioritized fallback system that balances custom configuration with automatic abbreviation.

### Explicit Custom Mappings

When creating a `Dialect` instance, you supply a dictionary of full names to short codes via the `entities` parameter. The constructor stores these mappings in `self.entity_codes` (lines 34-38), and they take precedence during the encoding process.

### Case-Insensitive Lookup

The constructor automatically registers lower-cased variants of each mapped name (lines 36-38). This ensures that lookups succeed regardless of input casing—both `"Alice"` and `"alice"` resolve to the same configured code.

### Auto-Generated Three-Letter Codes

If a name lacks an explicit mapping, the dialect falls back to an **auto-generated code** consisting of the first three characters of the name, converted to uppercase. This logic appears in `encode_entity` at lines 99-101:

```python
return name[:3].upper()

```

### Skip Lists for Exclusion

Names listed in the optional `skip_names` parameter are filtered from output entirely. The `encode_entity` method exits early (lines 92-94) and returns `None` for these entries, effectively removing them from the compressed representation.

## The Entity Encoding Workflow

When processing a name `n`, the `encode_entity` method follows this deterministic resolution order:

1. **Skip Check** – If `n` matches an entry in `skip_names`, return `None`
2. **Exact Match** – If `n` exists in `self.entity_codes`, return the mapped code
3. **Case-Insensitive Match** – If `n.lower()` exists in the mapping, return the mapped code
4. **Substring Match** – If any known name is a substring of `n`, return that name’s code
5. **Auto-Generation** – Return the first three characters of `n` upper-cased

This workflow ensures that user preferences always override automatic behavior while providing deterministic fallbacks for unconfigured entities.

## Practical Implementation Examples

The following examples demonstrate entity code configuration and usage:

```python
from mempalace.dialect import Dialect

# Custom explicit mapping

custom_entities = {"Alice": "ALC", "Bob": "BOB"}
dial = Dialect(entities=custom_entities)

print(dial.encode_entity("Alice"))   # → "ALC"

print(dial.encode_entity("bob"))     # → "BOB" (case-insensitive)

print(dial.encode_entity("Charlie")) # → "CHA" (auto-coded)

```

Loading entity definitions from configuration files:

```python

# entities.json contents:

# {"entities": {"Mona Lisa": "MLA", "Leonardo": "LEO"}}

dial = Dialect.from_config("entities.json")
print(dial.encode_entity("Mona Lisa"))          # → "MLA"

print(dial.encode_entity("Leonardo da Vinci")) # → "LEO"

```

Excluding specific names from encoding:

```python
dial = Dialect(skip_names=["Gandalf"])
print(dial.encode_entity("Gandalf"))  # → None (skipped)

```

## Compression Output Format

When the dialect compresses full text via the `compress()` method (lines 80-82), it collects all detected entity codes and formats them in the entity section of the output line. The compressed format appears as:

```

0:ALC+BOB+CHA|topic_keywords|"key sentence"|...

```

Here, `ALC`, `BOB`, and `CHA` represent the entity codes for Alice, Bob, and Charlie respectively, separated by plus signs within the entity segment (prefixed by `0:`).

## Summary

- **Explicit mappings** defined at initialization take precedence and are stored in `self.entity_codes` (lines 34-38)
- **Case-insensitive resolution** allows flexible matching regardless of input casing (lines 36-38)
- **Auto-coding** generates three-letter uppercase abbreviations for unknown entities via `name[:3].upper()` (lines 99-101)
- **Skip lists** exclude specified names from output entirely (lines 92-94)
- The `compress()` method assembles codes into the final output format at lines 80-82 of [`mempalace/dialect.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/dialect.py)

## Frequently Asked Questions

### What happens if an entity name is not in the explicit mapping?

The dialect automatically generates a code using the first three characters of the name converted to uppercase. For example, `"Charlemagne"` becomes `"CHA"`. This fallback ensures every entity receives a code without requiring manual configuration of every possible name.

### Where is the entity encoding logic implemented in the MemPalace codebase?

The core implementation resides in [`mempalace/dialect.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/dialect.py) within the `Dialect` class. The `encode_entity` method (lines 92-101) contains the resolution logic, while the constructor (lines 34-38) handles initial configuration of custom mappings and case normalization.

### Can entity codes be configured via external configuration files?

Yes. The `Dialect` class provides a `from_config()` class method that loads entity mappings from JSON files. This allows you to maintain dictionaries of entity codes (such as `{"Mona Lisa": "MLA"}`) separate from your application code for easier maintenance and sharing.

### Are entity code lookups case-sensitive?

No. The dialect normalizes lookup keys to lowercase during initialization (lines 36-38), ensuring that `"Alice"`, `"alice"`, and `"ALICE"` all resolve to the same configured code. The auto-generation fallback also standardizes output to uppercase regardless of input casing.