# How Caveman Handles Multi-Language Conversations: Style Compression Without Translation

> Learn how Caveman achieves multi-language conversation compression without translation. It preserves original language, URLs, and code blocks, maintaining linguistic identity in your markdown. Read more!

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: deep-dive
- Published: 2026-07-12

---

**Caveman compresses natural-language markdown by shortening prose and removing filler words while explicitly preserving the original language, URLs, headings, and code blocks, ensuring multilingual content retains its linguistic identity.**

The JuliusBrussee/caveman repository provides a Claude-based compression tool that processes markdown files across any human language. Unlike translation tools, Caveman's multi-language conversation handling focuses exclusively on stylistic reduction—condensing verbose text into concise "caveman" format without converting Spanish to English or Portuguese to Japanese. The system achieves this through a language-agnostic detection pipeline and carefully constructed prompts that forbid linguistic transformation.

## Language-Agnostic Detection Pipeline

Caveman treats all human-written text equally, regardless of the character set or grammar rules involved.

### Content Classification in detect.py

The entry point for multi-language support begins in [`skills/caveman-compress/scripts/detect.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/detect.py). This module classifies files into four categories: `natural_language`, `code`, `config`, or `unknown`. Only files marked as `natural_language` proceed to compression, meaning a French blog post, a German technical guide, or a Japanese tutorial all receive identical processing pipelines. The detection logic relies on content patterns rather than language-specific dictionaries, ensuring universal compatibility.

### Protecting Metadata During Compression

Before the LLM processes any content, the [`compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/compress.py) script strips YAML front-matter from the document. According to the source code in [`skills/caveman-compress/scripts/compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/compress.py) (lines 68-74), the front-matter is stored separately and re-attached unchanged after compression completes. This protection mechanism ensures that language-specific metadata—such as `lang: pt-BR` or custom locale settings—remains intact throughout the transformation process.

## The Compression Prompt Strategy

The core language-preserving logic resides in how Caveman instructs the underlying model to transform text.

### build_compress_prompt and Strict Preservation Rules

Inside [`skills/caveman-compress/scripts/compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/compress.py), the `build_compress_prompt` function (lines 71-81) constructs a Claude prompt containing explicit constraints. The instructions direct the model to **"Compress this markdown into caveman format"** while mandating that it **"Preserve ALL URLs exactly"**, **"Preserve ALL headings exactly"**, and **"Preserve file paths and commands"**. Crucially, the prompt contains no translation instructions, creating a context where the model understands its role is strictly editorial shortening, not linguistic conversion.

### Style-Only Transformation Constraints

The prompt's **STRICT RULES** explicitly forbid modifying inline code, backticks, URLs, headings, or commands. The only permitted transformation is shortening natural-language prose. This constraint architecture ensures that when processing a Portuguese paragraph, the LLM outputs a shorter Portuguese paragraph; when processing Spanish content, the output remains Spanish. The compression removes filler words and rephrases sentences while maintaining the original vocabulary and grammatical structure.

## The Wenyan Exception: When Translation Is Intentional

Caveman includes one deliberate translation mode: the `wenyan` level. As documented in [`README.md`](https://github.com/JuliusBrussee/caveman/blob/main/README.md) (lines 34-36), invoking `/caveman wenyan` rewrites content in classical Chinese (wenyan). This represents the sole built-in translation feature in the codebase. All other compression levels—including `ultra` and the default mode—operate as language-preserving style compressors that respect the source document's original tongue.

## Practical Examples

Process a French markdown file while preserving the language:

```bash

# Compress a markdown file written in French – output stays French

/caveman-compress notes-fr.md

```

Switch to high compression without translation:

```bash

# Switch Caveman to the “ultra” level (high compression) – still French

/caveman ultra

```

The intentional translation exception:

```bash

# Use the special “wenyan” mode – the same French content will be rewritten in classical Chinese

/caveman wenyan

```

Programmatic usage works identically across all languages:

```python
from caveman_compress.scripts.compress import compress_file
from pathlib import Path

# Programmatic usage – works for any language

compress_file(Path("docs/portuguese_guide.md"))

```

## Summary

- **Language-agnostic detection**: The [`detect.py`](https://github.com/JuliusBrussee/caveman/blob/main/detect.py) script classifies files as `natural_language` without regard to specific languages, routing all human text to compression.
- **Metadata protection**: YAML front-matter is stripped and re-attached unchanged in [`compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/compress.py), preserving language-specific settings.
- **Explicit preservation rules**: The `build_compress_prompt` function mandates exact preservation of URLs, headings, and code while forbidding translation.
- **Style-only constraints**: The LLM may only shorten natural-language prose, ensuring Portuguese stays Portuguese and Spanish stays Spanish.
- **Single exception**: Only the `wenyan` mode performs translation, converting content to classical Chinese as the sole linguistic transformation offered.

## Frequently Asked Questions

### Does Caveman automatically detect the language of the content?

Caveman does not identify specific languages such as French or Japanese. Instead, [`skills/caveman-compress/scripts/detect.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/detect.py) categorizes files as `natural_language`, `code`, `config`, or `unknown`. Any file classified as natural language—regardless of whether it contains English, Arabic, or Korean text—proceeds to the compression stage with identical treatment.

### Will Caveman translate my English documents to another language?

No. Unless you explicitly invoke the `wenyan` mode, Caveman never translates content. The compression prompt in [`skills/caveman-compress/scripts/compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/compress.py) explicitly forbids linguistic transformation, directing the model to preserve the original wording while shortening sentence structures only.

### How does Caveman handle mixed-language documents?

Mixed-language documents containing, for example, both English and Spanish paragraphs are processed as single `natural_language` files. The style compression rules apply uniformly across the entire document, shortening each language section independently without cross-translating between them. URLs, code blocks, and headings remain untouched regardless of the language surrounding them.

### What happens to YAML front-matter in non-English files?

The [`compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/compress.py) script strips YAML front-matter before compression and re-attaches it afterward exactly as written. This ensures that language-specific metadata—such as `lang: de` declarations or locale-specific tags—survives the compression process unmodified, maintaining the document's original linguistic configuration.