How Pixel Mode Converts Skill Text to PNG Pages and When It Fires

Pixel mode is a lossy compression feature in the Caveman CLI that transforms Markdown skill bodies into PNG images named SKILL.px1.png, SKILL.px2.png, etc., triggered during installation or conversion when specific heuristic and entitlement conditions are met.

Pixel mode reduces token consumption by rendering skill documentation as images rather than raw text, transmitting base64-encoded PNGs to compatible models. Implemented in the JuliusBrussee/caveman repository, this feature integrates with the CLI's installation workflow to automatically compress skills only when the conversion yields smaller token payloads.

The Pixel Conversion Pipeline

The pixel engine—located at caveman-engine pixel render—processes skills through three distinct stages defined in the engine/pixel package.

Text Preparation and Page Segmentation

In engine/pixel/textprep.go, the raw Markdown body is split into discrete pages based on line density limits. Each page is padded or trimmed to meet a target line count before being passed to the rendering layer as a raster canvas.

PNG Rendering and Encoding

The engine/pixel/render.go module rasterizes each text page onto an image canvas and encodes it as a PNG file. The resulting bytes are base64-encoded for transmission to the model, while local copies persist as SKILL.pxN.png files (where N is the sequential page number) in the skill directory.

Skill Stub Creation

After successful rendering, the original SKILL.md body is replaced with a JSON stub that tells the model to read the generated PNG pages. The stub preserves the skill name and front-matter intact—ensuring discovery mechanisms remain functional—while containing image_url references to the base64-encoded images.

When Pixel Mode Fires (Activation Conditions)

Pixel mode activation is gated by five specific conditions enforced in engine/pixel/gate.go and the CLI proxy layer. All conditions must be satisfied for conversion to proceed.

CLI Flag Invocation – The user must explicitly pass --pixel to caveman install or caveman convert commands. According to the CLI implementation in packages/cli/src/index.ts (around line 4963), this flag routes the request to the pixel engine.

Model Allow-List Validation – The target model must appear in the CAVE_PIXEL_MODELS environment variable (default: claude-fable-5,gpt-5.6). The proxy validates this against its documented allow-list before permitting pixel mode entry.

Size Heuristic Evaluation – The engine estimates token counts for both the original text and the prospective PNG pages. Conversion proceeds only when the PNG pages are predicted to consume fewer tokens—the "pages beat the text" rule documented in README.md.

Successful Render Completion – If any step in the PNG rendering pipeline fails, the engine falls back to a byte-identical pass-through, leaving the skill unchanged and recording the specific gate that aborted the conversion.

Entitlement Verification – For gated compression features (e.g., Claude or OpenAI endpoints), a valid account entitlement must be present as verified in wrap-gate.runtime.mjs; otherwise, the proxy rejects pixel mode.

Practical Usage and Examples

CLI Installation with Pixel Conversion

Invoke pixel mode during skill installation or manual conversion:


# Install a skill and automatically convert it to PNG pages

caveman install my-skill --pixel

# Convert an existing skill directory in place

caveman convert ./my-skill --pixel

Engine Execution Flow

As implemented in packages/cli/src/index.ts, the CLI resolves the engine binary and invokes the pixel render command with the skill directory:

await execFile(engineBin, [
  "pixel", "render",
  "--dir", skillDir,
  "--density", "max",   // optional density setting
]);

Resulting File Structure

After conversion, the skill directory contains the stub and generated images:


my-skill/
├─ SKILL.md          # JSON stub referencing PNG pages

├─ SKILL.px1.png    # First rendered page

├─ SKILL.px2.png    # Second rendered page (if any)

└─ …

The stub content follows this structure:

{
  "media_type": "image/png",
  "url": "data:image/png;base64,iVBORw0KGgoAAA…"
}

Summary

  • Pixel mode converts skill Markdown to PNG sequences via engine/pixel/textprep.go and engine/pixel/render.go to reduce token usage for model interactions.
  • Conversion fires only when the --pixel flag is present, the model is in the CAVE_PIXEL_MODELS allow-list, the "pages beat the text" size heuristic passes, and valid entitlements exist.
  • The system replaces original skill bodies with JSON stubs containing base64-encoded image references while preserving metadata for discovery.
  • Failed renders trigger a byte-identical pass-through that leaves the skill unchanged, ensuring no data loss on conversion errors.

Frequently Asked Questions

What file naming convention does pixel mode use for generated images?

Pixel mode generates files following the pattern SKILL.pxN.png, where SKILL matches the skill identifier and N represents the sequential page number (e.g., SKILL.px1.png, SKILL.px2.png). These files persist locally while their base64-encoded contents transmit to the model.

Why would pixel mode refuse to convert a skill even with the --pixel flag?

The conversion aborts if the size heuristic determines that PNG pages would exceed the original text's token count, if the target model is not in the CAVE_PIXEL_MODELS allow-list, or if the account lacks required entitlements for model-gated compression features. In these cases, the engine falls back to byte-identical pass-through.

Is the original skill content preserved after pixel conversion?

No, the original Markdown body is replaced with a JSON stub containing image references. However, the stub retains the skill name and front-matter to preserve discovery functionality. Note that conversion only proceeds after successful PNG generation; failures trigger a pass-through that leaves the original content completely unchanged.

Which models support pixel mode by default?

The default CAVE_PIXEL_MODELS allow-list includes claude-fable-5 and gpt-5.6. Administrators can modify this environment variable to enable pixel mode for additional model endpoints, subject to entitlement verification in the proxy layer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →