How Caveman Handles Multi-Language Conversations: Style Compression Without Translation
Caveman compresses natural-language markdown by shortening prose and removing filler words while explicitly preserving the original language, URLs, headings, and code blocks, ensuring multilingual content retains its linguistic identity.
The JuliusBrussee/caveman repository provides a Claude-based compression tool that processes markdown files across any human language. Unlike translation tools, Caveman's multi-language conversation handling focuses exclusively on stylistic reduction—condensing verbose text into concise "caveman" format without converting Spanish to English or Portuguese to Japanese. The system achieves this through a language-agnostic detection pipeline and carefully constructed prompts that forbid linguistic transformation.
Language-Agnostic Detection Pipeline
Caveman treats all human-written text equally, regardless of the character set or grammar rules involved.
Content Classification in detect.py
The entry point for multi-language support begins in skills/caveman-compress/scripts/detect.py. This module classifies files into four categories: natural_language, code, config, or unknown. Only files marked as natural_language proceed to compression, meaning a French blog post, a German technical guide, or a Japanese tutorial all receive identical processing pipelines. The detection logic relies on content patterns rather than language-specific dictionaries, ensuring universal compatibility.
Protecting Metadata During Compression
Before the LLM processes any content, the compress.py script strips YAML front-matter from the document. According to the source code in skills/caveman-compress/scripts/compress.py (lines 68-74), the front-matter is stored separately and re-attached unchanged after compression completes. This protection mechanism ensures that language-specific metadata—such as lang: pt-BR or custom locale settings—remains intact throughout the transformation process.
The Compression Prompt Strategy
The core language-preserving logic resides in how Caveman instructs the underlying model to transform text.
build_compress_prompt and Strict Preservation Rules
Inside skills/caveman-compress/scripts/compress.py, the build_compress_prompt function (lines 71-81) constructs a Claude prompt containing explicit constraints. The instructions direct the model to "Compress this markdown into caveman format" while mandating that it "Preserve ALL URLs exactly", "Preserve ALL headings exactly", and "Preserve file paths and commands". Crucially, the prompt contains no translation instructions, creating a context where the model understands its role is strictly editorial shortening, not linguistic conversion.
Style-Only Transformation Constraints
The prompt's STRICT RULES explicitly forbid modifying inline code, backticks, URLs, headings, or commands. The only permitted transformation is shortening natural-language prose. This constraint architecture ensures that when processing a Portuguese paragraph, the LLM outputs a shorter Portuguese paragraph; when processing Spanish content, the output remains Spanish. The compression removes filler words and rephrases sentences while maintaining the original vocabulary and grammatical structure.
The Wenyan Exception: When Translation Is Intentional
Caveman includes one deliberate translation mode: the wenyan level. As documented in README.md (lines 34-36), invoking /caveman wenyan rewrites content in classical Chinese (wenyan). This represents the sole built-in translation feature in the codebase. All other compression levels—including ultra and the default mode—operate as language-preserving style compressors that respect the source document's original tongue.
Practical Examples
Process a French markdown file while preserving the language:
# Compress a markdown file written in French – output stays French
/caveman-compress notes-fr.md
Switch to high compression without translation:
# Switch Caveman to the “ultra” level (high compression) – still French
/caveman ultra
The intentional translation exception:
# Use the special “wenyan” mode – the same French content will be rewritten in classical Chinese
/caveman wenyan
Programmatic usage works identically across all languages:
from caveman_compress.scripts.compress import compress_file
from pathlib import Path
# Programmatic usage – works for any language
compress_file(Path("docs/portuguese_guide.md"))
Summary
- Language-agnostic detection: The
detect.pyscript classifies files asnatural_languagewithout regard to specific languages, routing all human text to compression. - Metadata protection: YAML front-matter is stripped and re-attached unchanged in
compress.py, preserving language-specific settings. - Explicit preservation rules: The
build_compress_promptfunction mandates exact preservation of URLs, headings, and code while forbidding translation. - Style-only constraints: The LLM may only shorten natural-language prose, ensuring Portuguese stays Portuguese and Spanish stays Spanish.
- Single exception: Only the
wenyanmode performs translation, converting content to classical Chinese as the sole linguistic transformation offered.
Frequently Asked Questions
Does Caveman automatically detect the language of the content?
Caveman does not identify specific languages such as French or Japanese. Instead, skills/caveman-compress/scripts/detect.py categorizes files as natural_language, code, config, or unknown. Any file classified as natural language—regardless of whether it contains English, Arabic, or Korean text—proceeds to the compression stage with identical treatment.
Will Caveman translate my English documents to another language?
No. Unless you explicitly invoke the wenyan mode, Caveman never translates content. The compression prompt in skills/caveman-compress/scripts/compress.py explicitly forbids linguistic transformation, directing the model to preserve the original wording while shortening sentence structures only.
How does Caveman handle mixed-language documents?
Mixed-language documents containing, for example, both English and Spanish paragraphs are processed as single natural_language files. The style compression rules apply uniformly across the entire document, shortening each language section independently without cross-translating between them. URLs, code blocks, and headings remain untouched regardless of the language surrounding them.
What happens to YAML front-matter in non-English files?
The compress.py script strips YAML front-matter before compression and re-attaches it afterward exactly as written. This ensures that language-specific metadata—such as lang: de declarations or locale-specific tags—survives the compression process unmodified, maintaining the document's original linguistic configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →