# How Wenyan Compression Works in Caveman: Rule-Based Classical Chinese Text Transformation

> Discover Wenyan compression in Caveman. Learn how rule-based classical Chinese text transformation reduces token count by 80-90% while preserving meaning without external translation models.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: how-to-guide
- Published: 2026-07-09

---

**Wenyan compression in Caveman rewrites generated text into Classical Chinese (文言文) using deterministic rule-based transformations rather than external translation models, reducing token count by 80–90% while preserving semantic meaning.**

The Caveman project by JuliusBrussee implements a specialized text compression pipeline that transforms modern technical writing into Classical Chinese through its Wenyan compression mode. This unique approach leverages predefined rewrite rules stored in [`skills/caveman/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman/SKILL.md) to achieve dramatic token reduction while maintaining semantic accuracy. Understanding how Wenyan compression works requires examining the mode selection logic, the rule-based transformation pipeline, and the session persistence mechanism.

## Mode Selection and Canonicalization

Users activate Wenyan compression via the `/caveman` command followed by the desired intensity level. The system accepts three variants: `wenyan-lite`, `wenyan-full`, and `wenyan-ultra`, with the alias `wenyan` automatically canonicalizing to `wenyan-full` through the mode-tracker hook.

In [`src/hooks/caveman-mode-tracker.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/hooks/caveman-mode-tracker.js), the middleware intercepts user commands and normalizes mode aliases to their canonical forms. This ensures that invoking `/caveman wenyan` triggers the same ruleset as explicitly requesting `/caveman wenyan-full`, providing a consistent interface while supporting shorthand notation.

## The Rule-Based Rewriting Pipeline

The core of Wenyan compression resides in [`skills/caveman/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman/SKILL.md), which defines the transformation rules for each intensity level. Unlike machine translation approaches, Caveman applies deterministic text transformations that modify syntax, vocabulary, and structure according to Classical Chinese grammatical patterns.

### Syntax Reordering and Particle Insertion

The transformation rules systematically reorder modern Chinese syntax to match Classical Chinese conventions. This includes placing verbs before objects, omitting explicit subjects where context permits, and inserting classical particles such as `之`, `乃`, `為`, and `其` to create grammatically correct Wenyan prose.

### Compression Levels and Aggressive Optimization

Each Wenyan level applies progressively aggressive compression:

- **wenyan-lite**: Drops filler words and hedging language while maintaining modern sentence structure.
- **wenyan-full**: Reorders syntax and compresses clauses using classical particles, achieving approximately 80% token reduction.
- **wenyan-ultra**: Aggressively abbreviates sentences through idiomatic compression, potentially reducing token count by 90%.

For example, the phrase "Component frequently re-renders, causing each new object to reference the previous one" transforms to "組件頻重繪，以每繪新生對象參照故" at the lite level, compresses further to "每繪新生對象參照，故重繪" at full intensity, and reduces to "新參照則重繪" in ultra mode.

## Session Persistence and State Management

Selected compression modes persist across the entire session through the `.caveman-active` flag file stored in the repository root. This file contains the active mode identifier—such as `wenyan-ultra` or `wenyan-full`—ensuring consistent text transformation until the user explicitly switches modes.

The persistence mechanism is validated in [`tests/verify_repo.py`](https://github.com/JuliusBrussee/caveman/blob/main/tests/verify_repo.py), which confirms that the flag file correctly stores and retrieves all seven compression levels, including the three Wenyan variants. This stateful approach allows developers to maintain their preferred compression intensity throughout extended coding sessions without repeated mode selection.

## Activating Wenyan Compression

To enable Wenyan compression in your Caveman session:

```bash

# Activate standard Classical Chinese mode

/caveman wenyan

```

```bash

# Activate lite compression (minimal transformation)

/caveman wenyan-lite

```

```bash

# Activate maximum compression (ultra-abbreviated)

/caveman wenyan-ultra

```

Once activated, all LLM-generated output passes through the deterministic transformer defined in [`skills/caveman/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman/SKILL.md), resulting in Classical Chinese text that preserves technical meaning while minimizing token consumption.

## Summary

- Wenyan compression transforms text into Classical Chinese using deterministic rule-based transformations rather than external translation APIs.
- The mode-tracker in [`src/hooks/caveman-mode-tracker.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/hooks/caveman-mode-tracker.js) canonicalizes aliases like `wenyan` to `wenyan-full`.
- Three intensity levels—`wenyan-lite`, `wenyan-full`, and `wenyan-ultra`—provide progressive compression from 80% to 90% token reduction.
- Transformation rules in [`skills/caveman/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman/SKILL.md) handle syntax reordering, particle insertion, and idiom compression.
- Session persistence via `.caveman-active` maintains the selected mode across interactions.

## Frequently Asked Questions

### What is the difference between wenyan-lite and wenyan-ultra?

Wenyan-lite applies basic transformations that remove filler words and modern Chinese colloquialisms while maintaining recognizable sentence structures. Wenyan-ultra aggressively compresses text using classical idioms and extreme abbreviation, often reducing sentences to their minimal semantic components by removing grammatical markers and employing classical ellipses.

### Does Wenyan compression require an external translation API?

No, Wenyan compression operates entirely through local rule-based transformations defined in [`skills/caveman/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman/SKILL.md). The system applies deterministic regex and string manipulation patterns rather than calling external LLMs or translation services, ensuring deterministic output and offline functionality.

### How does Caveman preserve technical accuracy when compressing to Classical Chinese?

The transformation rules in [`skills/caveman/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman/SKILL.md) specifically account for technical terminology by preserving key nouns and code references (such as `useMemo` or function names) while compressing surrounding explanatory text. The rule-based approach ensures that technical concepts remain intact even as grammatical structures shift to Classical Chinese patterns.

### Can I switch between Wenyan levels mid-session?

Yes, you can switch between any of the seven Caveman compression levels—including the three Wenyan variants—by issuing a new `/caveman` command. The system updates the `.caveman-active` flag file immediately, and subsequent LLM responses will use the newly selected transformation rules without requiring session restart.