What Is the AAAK Compression Dialect in MemPalace?
The AAAK (Abstract Algebraic Annotation Kernel) dialect is a lossy summarization format in MemPalace that encodes entities, topics, emotions, and flags into compact LLM-readable strings and does not achieve 30× lossless compression.
The AAAK compression dialect provides a symbolic abstraction layer over large memory footprints in the MemPalace repository. Implemented primarily in mempalace/dialect.py, it trades verbatim fidelity for retrieval speed by emitting dense summaries that large language models can skim quickly. Although early documentation mistakenly advertised a "30× loss-less ratio," the source code explicitly states that AAAK is a lossy format that discards fine-grained information.
How the AAAK Compression Dialect Works
The dialect follows a multi-stage pipeline. Each stage discards surface text in favor of an abstract token meant for fast matching.
Entity Coding via encode_entity
The encode_entity method in mempalace/dialect.py (lines 33‑40) replaces proper names with short 3‑letter codes or user‑provided mappings. This substitution shrinks the token footprint of repeated references.
Topic Extraction via _extract_topics
Inside _extract_topics (lines 52‑78), the engine ranks non‑stop‑word tokens by frequency and joins the top candidates with underscores. The result is a dense topic slug such as graphQL_rest that stands in for the original discussion.
Key-Sentence Selection via _extract_key_sentence
The _extract_key_sentence helper (lines 81‑102) scans sentences for decision words, scores each candidate, and extracts the highest‑scoring verbatim fragment. This quote becomes the single representative sentence for the input block.
Emotion and Flag Detection
Two private helpers, _detect_emotions and _detect_flags (lines 30‑49), consult keyword lists named _EMOTION_SIGNALS and _FLAG_SIGNALS. Plain‑text cues are mapped to short tokens such as joy, fear, and CORE, producing a symbolic emotional and priority profile.
Output Assembly in compress
The compress method (lines 60‑89) pulls the results together into a compact, delimited record. An optional header line can include metadata, and the remaining fields follow the entity, topic, quote, emotion, and flag sequence.
Why AAAK Is Lossy, Not Lossless
The AAAK compression dialect does not support reconstruction of the original text. In mempalace/dialect.py, the file header at lines 12‑14 explicitly states that "AAAK is NOT lossless compression." What survives the pipeline is a summary pointer to the original memory "drawers," not a compressed bitstream that can be decompressed back to the source string.
The mistaken claim of a "30× loss‑less ratio" originated in early documentation. The test suite in tests/test_readme_claims.py (lines 474‑511) now enforces that the README must contain the phrase "AAAK is lossy, not lossless." Any ratio reported by compression_stats (lines 67‑74) is merely an estimated size ratio of summary tokens versus original tokens, not a true lossless compression metric.
Working with the AAAK Dialect in Python
You can invoke the encoder through a simple sequence:
- Import
Dialectfrommempalace.dialect. - Call
compress()for strings orcompress_file()for JSON zettels. - Check
compression_stats()for estimated token reduction. - Use
generate_layer1()to aggregate a directory into a single skim file.
The same operations are exposed via the CLI and the MCP server registered in mempalace/mcp_server.py.
Compress Plain Text
from mempalace.dialect import Dialect
dialect = Dialect()
text = "We decided to use GraphQL instead of REST because it reduces over‑fetching."
compressed = dialect.compress(text)
print(compressed)
Sample output:
0:WE+? graphQL_rest "We decided to use GraphQL instead of REST because it reduces over‑fetching." determ+prefer DECISION+CORE
Compress a JSON Zettel File
dialect.compress_file("zettels/file_001.json", output_path="file_001.aaak")
Inspect Compression Statistics
original = "Long paragraph …"
summary = dialect.compress(original)
stats = dialect.compression_stats(original, summary)
print(stats)
The returned dictionary includes original_tokens_est, summary_tokens_est, and size_ratio:
{'original_tokens_est': 120, 'summary_tokens_est': 30, 'size_ratio': 4.0, ...}
Generate a Layer‑1 Wake‑Up File
dialect.generate_layer1("zettels/", output_path="LAYER1.aaak")
This aggregates an entire directory into a single file that an LLM can read quickly.
Summary
- The AAAK compression dialect is a lossy summarization engine living in
mempalace/dialect.py. - It encodes entities as 3‑letter codes, extracts topics from token frequency, selects a key sentence, and tags emotions and flags.
- The output is a compact, symbolic string designed for LLM skimming, not for lossless reconstruction.
- The source code explicitly rejects the "30× lossless" label, and
tests/test_readme_claims.pyenforces accurate documentation. - Practical access is available through the Python API, CLI, and the MCP server tool
aaak_dialect.
Frequently Asked Questions
What does AAAK stand for?
AAAK stands for Abstract Algebraic Annotation Kernel. It is the internal name for the summarization protocol used by MemPalace to abstract memory contents into compact annotations.
Does AAAK achieve 30× lossless compression?
No. The AAAK compression dialect is lossy. Early documentation contained a mistaken claim of a "30× loss‑less ratio," but the implementation in mempalace/dialect.py explicitly states that AAAK is not lossless, and automated tests in tests/test_readme_claims.py prevent the README from repeating that error.
How do I generate AAAK summaries from JSON files?
Call dialect.compress_file("path/to/file.json", output_path="output.aaak") on a Dialect instance. You can also batch an entire directory into a single Layer‑1 wake‑up file with dialect.generate_layer1("zettels/", output_path="LAYER1.aaak").
Where is the AAAK dialect exposed to external agents?
The MCP server defined in mempalace/mcp_server.py registers an aaak_dialect tool that returns the specification to agents. The same functionality is available through the Python API documented in website/reference/python-api.md and the human‑readable concept guide at website/concepts/aaak-dialect.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →