How Pixel Mode Converts Skill Text to PNG Pages and When It Fires
Pixel mode is a lossy compression feature in the Caveman CLI that transforms Markdown skill bodies into PNG images named SKILL.px1.png, SKILL.px2.png, etc., triggered during installation or conversion when specific heuristic and entitlement conditions are met.
Pixel mode reduces token consumption by rendering skill documentation as images rather than raw text, transmitting base64-encoded PNGs to compatible models. Implemented in the JuliusBrussee/caveman repository, this feature integrates with the CLI's installation workflow to automatically compress skills only when the conversion yields smaller token payloads.
The Pixel Conversion Pipeline
The pixel engine—located at caveman-engine pixel render—processes skills through three distinct stages defined in the engine/pixel package.
Text Preparation and Page Segmentation
In engine/pixel/textprep.go, the raw Markdown body is split into discrete pages based on line density limits. Each page is padded or trimmed to meet a target line count before being passed to the rendering layer as a raster canvas.
PNG Rendering and Encoding
The engine/pixel/render.go module rasterizes each text page onto an image canvas and encodes it as a PNG file. The resulting bytes are base64-encoded for transmission to the model, while local copies persist as SKILL.pxN.png files (where N is the sequential page number) in the skill directory.
Skill Stub Creation
After successful rendering, the original SKILL.md body is replaced with a JSON stub that tells the model to read the generated PNG pages. The stub preserves the skill name and front-matter intact—ensuring discovery mechanisms remain functional—while containing image_url references to the base64-encoded images.
When Pixel Mode Fires (Activation Conditions)
Pixel mode activation is gated by five specific conditions enforced in engine/pixel/gate.go and the CLI proxy layer. All conditions must be satisfied for conversion to proceed.
CLI Flag Invocation – The user must explicitly pass --pixel to caveman install or caveman convert commands. According to the CLI implementation in packages/cli/src/index.ts (around line 4963), this flag routes the request to the pixel engine.
Model Allow-List Validation – The target model must appear in the CAVE_PIXEL_MODELS environment variable (default: claude-fable-5,gpt-5.6). The proxy validates this against its documented allow-list before permitting pixel mode entry.
Size Heuristic Evaluation – The engine estimates token counts for both the original text and the prospective PNG pages. Conversion proceeds only when the PNG pages are predicted to consume fewer tokens—the "pages beat the text" rule documented in README.md.
Successful Render Completion – If any step in the PNG rendering pipeline fails, the engine falls back to a byte-identical pass-through, leaving the skill unchanged and recording the specific gate that aborted the conversion.
Entitlement Verification – For gated compression features (e.g., Claude or OpenAI endpoints), a valid account entitlement must be present as verified in wrap-gate.runtime.mjs; otherwise, the proxy rejects pixel mode.
Practical Usage and Examples
CLI Installation with Pixel Conversion
Invoke pixel mode during skill installation or manual conversion:
# Install a skill and automatically convert it to PNG pages
caveman install my-skill --pixel
# Convert an existing skill directory in place
caveman convert ./my-skill --pixel
Engine Execution Flow
As implemented in packages/cli/src/index.ts, the CLI resolves the engine binary and invokes the pixel render command with the skill directory:
await execFile(engineBin, [
"pixel", "render",
"--dir", skillDir,
"--density", "max", // optional density setting
]);
Resulting File Structure
After conversion, the skill directory contains the stub and generated images:
my-skill/
├─ SKILL.md # JSON stub referencing PNG pages
├─ SKILL.px1.png # First rendered page
├─ SKILL.px2.png # Second rendered page (if any)
└─ …
The stub content follows this structure:
{
"media_type": "image/png",
"url": "data:image/png;base64,iVBORw0KGgoAAA…"
}
Summary
- Pixel mode converts skill Markdown to PNG sequences via
engine/pixel/textprep.goandengine/pixel/render.goto reduce token usage for model interactions. - Conversion fires only when the
--pixelflag is present, the model is in theCAVE_PIXEL_MODELSallow-list, the "pages beat the text" size heuristic passes, and valid entitlements exist. - The system replaces original skill bodies with JSON stubs containing base64-encoded image references while preserving metadata for discovery.
- Failed renders trigger a byte-identical pass-through that leaves the skill unchanged, ensuring no data loss on conversion errors.
Frequently Asked Questions
What file naming convention does pixel mode use for generated images?
Pixel mode generates files following the pattern SKILL.pxN.png, where SKILL matches the skill identifier and N represents the sequential page number (e.g., SKILL.px1.png, SKILL.px2.png). These files persist locally while their base64-encoded contents transmit to the model.
Why would pixel mode refuse to convert a skill even with the --pixel flag?
The conversion aborts if the size heuristic determines that PNG pages would exceed the original text's token count, if the target model is not in the CAVE_PIXEL_MODELS allow-list, or if the account lacks required entitlements for model-gated compression features. In these cases, the engine falls back to byte-identical pass-through.
Is the original skill content preserved after pixel conversion?
No, the original Markdown body is replaced with a JSON stub containing image references. However, the stub retains the skill name and front-matter to preserve discovery functionality. Note that conversion only proceeds after successful PNG generation; failures trigger a pass-through that leaves the original content completely unchanged.
Which models support pixel mode by default?
The default CAVE_PIXEL_MODELS allow-list includes claude-fable-5 and gpt-5.6. Administrators can modify this environment variable to enable pixel mode for additional model endpoints, subject to entitlement verification in the proxy layer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →