Understanding the AAAK Compression Dialect for Index Scanning in MemPalace
The AAAK (A-Annotated-A-K) compression dialect transforms verbose documents into compact, human-readable summaries—typically ~10 tokens per moment—enabling large language models to scan massive indices instantly without ingesting raw text corpora.
The AAAK compression dialect serves as the cornerstone of MemPalace's index-scanning architecture, converting transcripts, markdown notes, and JSON zettels into lossy but semantically dense representations. Instead of feeding every verbatim drawer to a language model, MemPalace creates a fast "lookup table" that captures the who, what, and why of content while discarding expendable wording. This implementation lives primarily in mempalace/dialect.py within the MemPalace/mempalace repository.
Core Architecture of the Dialect Class
The Dialect class orchestrates the entire compression pipeline, handling entity encoding, emotion detection, flag assignment, and topic extraction before assembling the final AAAK line. Defined at line 300 in mempalace/dialect.py, this class serves as the primary entry point for transforming raw text into indexable summaries.
Entity Encoding and Emotion Detection
The dialect replaces full names with short 3-letter entity codes to minimize token count. The encode_entity() method (lines 89-100) manages this compression, while the EMOTION_CODES dictionary (lines 53-94) provides a lookup table for emotional tagging. The _detect_emotions() method (lines 330-340) scans text for emotional keywords and produces up to three emoji-style tags that signal the sentiment weight of a moment.
Flag Detection and Topic Extraction
Structural importance is signaled through flags such as ORIGIN, CORE, PIVOT, and TECHNICAL. The _FLAG_SIGNALS dictionary (lines 122-158) defines these markers, while _detect_flags() (lines 341-350) applies them based on keyword presence. For thematic indexing, _extract_topics() (lines 352-378) employs a frequency-based word picker that drops stop-words and boosts proper nouns to identify the subject matter.
Key-Sentence Extraction
To preserve the most decision-oriented content, _extract_key_sentence() (lines 380-426) selects the pivotal sentence fragment that best summarizes the moment. This extracted quote becomes the semantic anchor within the compressed line.
The AAAK Format Specification
The AAAK dialect uses a deliberately simple plain-text structure that any LLM can parse without custom decoders. The format supports optional metadata headers and mandatory content lines:
HEADER (optional): wing|room|date|title
CONTENT line: 0:ENTITIES|topic_keywords|"key_quote"|EMOTIONS|FLAGS
A concrete example illustrates the compression ratio achieved:
2023-07-15|ALC+BOB|2023-07-15|project-kickoff
0:ALC+BOB|api_design| "We decided to use GraphQL instead of REST" |joy+trust |ORIGIN+CORE
The entity codes (ALC, BOB) identify actors instantly; topic keywords (api_design) provide thematic tags; the quoted fragment supplies the pivotal sentence; and emotions and flags signal importance for ranking algorithms.
Index Scanning Workflow
The dialect integrates into a four-stage retrieval pipeline that minimizes latency while maximizing recall accuracy.
Ingestion and Compression
When new data arrives—whether plain text or JSON zettels—MemPalace routes it through compress() (lines 560-625) for raw text or encode_file() for zettel formats. The compress() method glues together headers, entities, topics, key sentences, emotions, and flags into the final AAAK line.
Storage in ChromaDB
The resulting one-line strings are stored as drawers in ChromaDB via mempalace/backends/chroma.py. Because each compressed line averages approximately 10 tokens, massive collections can be scanned with negligible latency.
Layer-1 Generation for Instant Recall
The generate_layer1() method (lines 806-879) aggregates the highest-weight moments—those marked with ORIGIN or CORE flags or carrying high emotional weight—into a "wake-up" file. When a user submits a query, the search stack first feeds this Layer-1 file to the LLM, providing instant context without loading the full corpus.
Fine-Grained Retrieval
If deeper detail is required, the LLM requests the specific full drawer referenced by the AAAK line's entity code or ZID (in zettel format). This drill-down approach ensures the LLM never processes unnecessary raw text.
Practical Implementation Examples
Compressing Plain Text via Python API
from mempalace.dialect import Dialect
dialect = Dialect()
text = (
"We decided to use GraphQL instead of REST because it reduces over-fetching. "
"The team (Alice, Bob) is thrilled."
)
compressed = dialect.compress(
text,
metadata={
"wing": "ENG",
"room": "2023-07-15",
"date": "2023-07-15",
"source_file": "notes.txt"
}
)
print(compressed)
This outputs:
ENG|2023-07-15|2023-07-15|notes
0:ALC+BOB|graphQL_rest| "We decided to use GraphQL instead of REST" |joy+trust |DECISION+CORE
Generating Layer-1 Files via CLI
python -m mempalace.dialect --layer1 ./zettels/
This command reads every *.json file in ./zettels/, extracts the highest-weight moments using generate_layer1(), and writes the aggregated LAYER1.aaak file for immediate LLM consumption.
Encoding Zettel JSON Files
dialect = Dialect()
compressed = dialect.compress_file(
"zettels/file_001.json",
output_path="file_001.aaak"
)
print(compressed[:200]) # preview first few lines
The compress_file() method (lines 777-785) loads JSON structures, processes each zettel through encode_file(), and persists the AAAK representation to disk.
Estimating Token Savings
original = open("large_transcript.txt").read()
compressed = Dialect().compress(original)
stats = Dialect().compression_stats(original, compressed)
print(f"Size ratio: {stats['size_ratio']}×")
The count_tokens() helper (lines 553-566) provides the heuristic used for these compression ratio calculations.
Summary
- The AAAK compression dialect reduces documents to ~10-token summaries containing entities, topics, key quotes, emotions, and flags.
- Entity encoding (
encode_entity()) and emotion detection (_detect_emotions()) compress semantic meaning into machine-readable tags. - Layer-1 generation (
generate_layer1()) creates a prioritized "wake-up" file that LLMs scan first for instant recall. - The format is model-agnostic and human-readable, requiring no custom decoders for Claude, GPT-4, Gemini, or Llama models.
- Integration with ChromaDB and the searcher pipeline enables sub-10ms index scanning before fine-grained retrieval.
Frequently Asked Questions
How does the AAAK compression dialect handle entity disambiguation?
The encode_entity() method in mempalace/dialect.py (lines 89-100) maps full names to unique 3-letter codes based on a configurable registry. If collisions occur, the dialect appends numeric suffixes to maintain uniqueness while preserving the 3-character constraint.
Can the AAAK format be customized for domain-specific terminology?
Yes. While the base Dialect class provides standard emotion codes and flag signals (lines 53-158), you can extend the class to override EMOTION_CODES or _FLAG_SIGNALS before instantiation. The topic extraction logic in _extract_topics() (lines 352-378) also accepts custom stop-word lists via the constructor.
What performance gains does Layer-1 generation provide?
According to the implementation in generate_layer1() (lines 806-879), a Layer-1 file containing dozens of compressed moments can be read by an LLM in under 10 milliseconds. This compares to potentially seconds of latency when processing raw transcripts, enabling the "instant recall" promise of the MemPalace architecture.
How does the dialect integrate with the broader MemPalace retrieval system?
The Dialect class works alongside mempalace/searcher.py for hybrid BM25 and vector search, and mempalace/palace.py for high-level orchestration. Compressed AAAK lines are stored as ChromaDB drawers via mempalace/backends/chroma.py, creating a unified pipeline from ingestion through retrieval.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →