How Wenyan Compression Works in Caveman: Rule-Based Classical Chinese Text Transformation

Wenyan compression in Caveman rewrites generated text into Classical Chinese (文言文) using deterministic rule-based transformations rather than external translation models, reducing token count by 80–90% while preserving semantic meaning.

The Caveman project by JuliusBrussee implements a specialized text compression pipeline that transforms modern technical writing into Classical Chinese through its Wenyan compression mode. This unique approach leverages predefined rewrite rules stored in skills/caveman/SKILL.md to achieve dramatic token reduction while maintaining semantic accuracy. Understanding how Wenyan compression works requires examining the mode selection logic, the rule-based transformation pipeline, and the session persistence mechanism.

Mode Selection and Canonicalization

Users activate Wenyan compression via the /caveman command followed by the desired intensity level. The system accepts three variants: wenyan-lite, wenyan-full, and wenyan-ultra, with the alias wenyan automatically canonicalizing to wenyan-full through the mode-tracker hook.

In src/hooks/caveman-mode-tracker.js, the middleware intercepts user commands and normalizes mode aliases to their canonical forms. This ensures that invoking /caveman wenyan triggers the same ruleset as explicitly requesting /caveman wenyan-full, providing a consistent interface while supporting shorthand notation.

The Rule-Based Rewriting Pipeline

The core of Wenyan compression resides in skills/caveman/SKILL.md, which defines the transformation rules for each intensity level. Unlike machine translation approaches, Caveman applies deterministic text transformations that modify syntax, vocabulary, and structure according to Classical Chinese grammatical patterns.

Syntax Reordering and Particle Insertion

The transformation rules systematically reorder modern Chinese syntax to match Classical Chinese conventions. This includes placing verbs before objects, omitting explicit subjects where context permits, and inserting classical particles such as 之, 乃, 為, and 其 to create grammatically correct Wenyan prose.

Compression Levels and Aggressive Optimization

Each Wenyan level applies progressively aggressive compression:

  • wenyan-lite: Drops filler words and hedging language while maintaining modern sentence structure.
  • wenyan-full: Reorders syntax and compresses clauses using classical particles, achieving approximately 80% token reduction.
  • wenyan-ultra: Aggressively abbreviates sentences through idiomatic compression, potentially reducing token count by 90%.

For example, the phrase "Component frequently re-renders, causing each new object to reference the previous one" transforms to "組件頻重繪,以每繪新生對象參照故" at the lite level, compresses further to "每繪新生對象參照,故重繪" at full intensity, and reduces to "新參照則重繪" in ultra mode.

Session Persistence and State Management

Selected compression modes persist across the entire session through the .caveman-active flag file stored in the repository root. This file contains the active mode identifier—such as wenyan-ultra or wenyan-full—ensuring consistent text transformation until the user explicitly switches modes.

The persistence mechanism is validated in tests/verify_repo.py, which confirms that the flag file correctly stores and retrieves all seven compression levels, including the three Wenyan variants. This stateful approach allows developers to maintain their preferred compression intensity throughout extended coding sessions without repeated mode selection.

Activating Wenyan Compression

To enable Wenyan compression in your Caveman session:


# Activate standard Classical Chinese mode

/caveman wenyan

# Activate lite compression (minimal transformation)

/caveman wenyan-lite

# Activate maximum compression (ultra-abbreviated)

/caveman wenyan-ultra

Once activated, all LLM-generated output passes through the deterministic transformer defined in skills/caveman/SKILL.md, resulting in Classical Chinese text that preserves technical meaning while minimizing token consumption.

Summary

  • Wenyan compression transforms text into Classical Chinese using deterministic rule-based transformations rather than external translation APIs.
  • The mode-tracker in src/hooks/caveman-mode-tracker.js canonicalizes aliases like wenyan to wenyan-full.
  • Three intensity levels—wenyan-lite, wenyan-full, and wenyan-ultra—provide progressive compression from 80% to 90% token reduction.
  • Transformation rules in skills/caveman/SKILL.md handle syntax reordering, particle insertion, and idiom compression.
  • Session persistence via .caveman-active maintains the selected mode across interactions.

Frequently Asked Questions

What is the difference between wenyan-lite and wenyan-ultra?

Wenyan-lite applies basic transformations that remove filler words and modern Chinese colloquialisms while maintaining recognizable sentence structures. Wenyan-ultra aggressively compresses text using classical idioms and extreme abbreviation, often reducing sentences to their minimal semantic components by removing grammatical markers and employing classical ellipses.

Does Wenyan compression require an external translation API?

No, Wenyan compression operates entirely through local rule-based transformations defined in skills/caveman/SKILL.md. The system applies deterministic regex and string manipulation patterns rather than calling external LLMs or translation services, ensuring deterministic output and offline functionality.

How does Caveman preserve technical accuracy when compressing to Classical Chinese?

The transformation rules in skills/caveman/SKILL.md specifically account for technical terminology by preserving key nouns and code references (such as useMemo or function names) while compressing surrounding explanatory text. The rule-based approach ensures that technical concepts remain intact even as grammatical structures shift to Classical Chinese patterns.

Can I switch between Wenyan levels mid-session?

Yes, you can switch between any of the seven Caveman compression levels—including the three Wenyan variants—by issuing a new /caveman command. The system updates the .caveman-active flag file immediately, and subsequent LLM responses will use the newly selected transformation rules without requiring session restart.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →