What Entity Codes Does the AAAK Compression Dialect Use? A Technical Guide to MemPalace Compression

The AAAK compression dialect uses a hybrid system of explicit user-defined mappings, case-insensitive dictionary lookups, and deterministic three-letter abbreviations to compress entity names into short codes.

The MemPalace/mempalace repository implements the AAAK dialect to compactly represent zettelkasten notes. Understanding how entity codes function is essential for configuring the compression pipeline, as these short identifiers replace full entity names to minimize storage while preserving semantic references.

How Entity Codes Are Generated

The encoding logic resides in mempalace/dialect.py, specifically within the Dialect class and its encode_entity method. The dialect resolves entity names through a prioritized fallback system that balances custom configuration with automatic abbreviation.

Explicit Custom Mappings

When creating a Dialect instance, you supply a dictionary of full names to short codes via the entities parameter. The constructor stores these mappings in self.entity_codes (lines 34-38), and they take precedence during the encoding process.

Case-Insensitive Lookup

The constructor automatically registers lower-cased variants of each mapped name (lines 36-38). This ensures that lookups succeed regardless of input casing—both "Alice" and "alice" resolve to the same configured code.

Auto-Generated Three-Letter Codes

If a name lacks an explicit mapping, the dialect falls back to an auto-generated code consisting of the first three characters of the name, converted to uppercase. This logic appears in encode_entity at lines 99-101:

return name[:3].upper()

Skip Lists for Exclusion

Names listed in the optional skip_names parameter are filtered from output entirely. The encode_entity method exits early (lines 92-94) and returns None for these entries, effectively removing them from the compressed representation.

The Entity Encoding Workflow

When processing a name n, the encode_entity method follows this deterministic resolution order:

  1. Skip Check – If n matches an entry in skip_names, return None
  2. Exact Match – If n exists in self.entity_codes, return the mapped code
  3. Case-Insensitive Match – If n.lower() exists in the mapping, return the mapped code
  4. Substring Match – If any known name is a substring of n, return that name’s code
  5. Auto-Generation – Return the first three characters of n upper-cased

This workflow ensures that user preferences always override automatic behavior while providing deterministic fallbacks for unconfigured entities.

Practical Implementation Examples

The following examples demonstrate entity code configuration and usage:

from mempalace.dialect import Dialect

# Custom explicit mapping

custom_entities = {"Alice": "ALC", "Bob": "BOB"}
dial = Dialect(entities=custom_entities)

print(dial.encode_entity("Alice"))   # → "ALC"

print(dial.encode_entity("bob"))     # → "BOB" (case-insensitive)

print(dial.encode_entity("Charlie")) # → "CHA" (auto-coded)

Loading entity definitions from configuration files:


# entities.json contents:

# {"entities": {"Mona Lisa": "MLA", "Leonardo": "LEO"}}

dial = Dialect.from_config("entities.json")
print(dial.encode_entity("Mona Lisa"))          # → "MLA"

print(dial.encode_entity("Leonardo da Vinci")) # → "LEO"

Excluding specific names from encoding:

dial = Dialect(skip_names=["Gandalf"])
print(dial.encode_entity("Gandalf"))  # → None (skipped)

Compression Output Format

When the dialect compresses full text via the compress() method (lines 80-82), it collects all detected entity codes and formats them in the entity section of the output line. The compressed format appears as:


0:ALC+BOB+CHA|topic_keywords|"key sentence"|...

Here, ALC, BOB, and CHA represent the entity codes for Alice, Bob, and Charlie respectively, separated by plus signs within the entity segment (prefixed by 0:).

Summary

  • Explicit mappings defined at initialization take precedence and are stored in self.entity_codes (lines 34-38)
  • Case-insensitive resolution allows flexible matching regardless of input casing (lines 36-38)
  • Auto-coding generates three-letter uppercase abbreviations for unknown entities via name[:3].upper() (lines 99-101)
  • Skip lists exclude specified names from output entirely (lines 92-94)
  • The compress() method assembles codes into the final output format at lines 80-82 of mempalace/dialect.py

Frequently Asked Questions

What happens if an entity name is not in the explicit mapping?

The dialect automatically generates a code using the first three characters of the name converted to uppercase. For example, "Charlemagne" becomes "CHA". This fallback ensures every entity receives a code without requiring manual configuration of every possible name.

Where is the entity encoding logic implemented in the MemPalace codebase?

The core implementation resides in mempalace/dialect.py within the Dialect class. The encode_entity method (lines 92-101) contains the resolution logic, while the constructor (lines 34-38) handles initial configuration of custom mappings and case normalization.

Can entity codes be configured via external configuration files?

Yes. The Dialect class provides a from_config() class method that loads entity mappings from JSON files. This allows you to maintain dictionaries of entity codes (such as {"Mona Lisa": "MLA"}) separate from your application code for easier maintenance and sharing.

Are entity code lookups case-sensitive?

No. The dialect normalizes lookup keys to lowercase during initialization (lines 36-38), ensuring that "Alice", "alice", and "ALICE" all resolve to the same configured code. The auto-generation fallback also standardizes output to uppercase regardless of input casing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →