# How Pixel Mode Converts Skill Text to PNG Pages and When It Fires

> Discover how Caveman CLI's pixel mode converts skill text to PNG images and when it activates. Learn about this lossy compression feature for efficient skill management.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: how-to-guide
- Published: 2026-09-04

---

**Pixel mode is a lossy compression feature in the Caveman CLI that transforms Markdown skill bodies into PNG images named `SKILL.px1.png`, `SKILL.px2.png`, etc., triggered during installation or conversion when specific heuristic and entitlement conditions are met.**

Pixel mode reduces token consumption by rendering skill documentation as images rather than raw text, transmitting base64-encoded PNGs to compatible models. Implemented in the `JuliusBrussee/caveman` repository, this feature integrates with the CLI's installation workflow to automatically compress skills only when the conversion yields smaller token payloads.

## The Pixel Conversion Pipeline

The pixel engine—located at `caveman-engine pixel render`—processes skills through three distinct stages defined in the `engine/pixel` package.

### Text Preparation and Page Segmentation

In [`engine/pixel/textprep.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/pixel/textprep.go), the raw Markdown body is split into discrete pages based on line density limits. Each page is padded or trimmed to meet a target line count before being passed to the rendering layer as a raster canvas.

### PNG Rendering and Encoding

The [`engine/pixel/render.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/pixel/render.go) module rasterizes each text page onto an image canvas and encodes it as a PNG file. The resulting bytes are base64-encoded for transmission to the model, while local copies persist as `SKILL.pxN.png` files (where `N` is the sequential page number) in the skill directory.

### Skill Stub Creation

After successful rendering, the original [`SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/SKILL.md) body is replaced with a JSON stub that tells the model to read the generated PNG pages. The stub preserves the skill name and front-matter intact—ensuring discovery mechanisms remain functional—while containing `image_url` references to the base64-encoded images.

## When Pixel Mode Fires (Activation Conditions)

Pixel mode activation is gated by five specific conditions enforced in [`engine/pixel/gate.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/pixel/gate.go) and the CLI proxy layer. All conditions must be satisfied for conversion to proceed.

**CLI Flag Invocation** – The user must explicitly pass `--pixel` to `caveman install` or `caveman convert` commands. According to the CLI implementation in [`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts) (around line 4963), this flag routes the request to the pixel engine.

**Model Allow-List Validation** – The target model must appear in the `CAVE_PIXEL_MODELS` environment variable (default: `claude-fable-5,gpt-5.6`). The proxy validates this against its documented allow-list before permitting pixel mode entry.

**Size Heuristic Evaluation** – The engine estimates token counts for both the original text and the prospective PNG pages. Conversion proceeds only when the PNG pages are predicted to consume fewer tokens—the "pages beat the text" rule documented in [`README.md`](https://github.com/JuliusBrussee/caveman/blob/main/README.md).

**Successful Render Completion** – If any step in the PNG rendering pipeline fails, the engine falls back to a **byte-identical pass-through**, leaving the skill unchanged and recording the specific gate that aborted the conversion.

**Entitlement Verification** – For gated compression features (e.g., Claude or OpenAI endpoints), a valid account entitlement must be present as verified in `wrap-gate.runtime.mjs`; otherwise, the proxy rejects pixel mode.

## Practical Usage and Examples

### CLI Installation with Pixel Conversion

Invoke pixel mode during skill installation or manual conversion:

```bash

# Install a skill and automatically convert it to PNG pages

caveman install my-skill --pixel

# Convert an existing skill directory in place

caveman convert ./my-skill --pixel

```

### Engine Execution Flow

As implemented in [`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts), the CLI resolves the engine binary and invokes the pixel render command with the skill directory:

```typescript
await execFile(engineBin, [
  "pixel", "render",
  "--dir", skillDir,
  "--density", "max",   // optional density setting
]);

```

### Resulting File Structure

After conversion, the skill directory contains the stub and generated images:

```

my-skill/
├─ SKILL.md          # JSON stub referencing PNG pages

├─ SKILL.px1.png    # First rendered page

├─ SKILL.px2.png    # Second rendered page (if any)

└─ …

```

The stub content follows this structure:

```json
{
  "media_type": "image/png",
  "url": "data:image/png;base64,iVBORw0KGgoAAA…"
}

```

## Summary

- **Pixel mode** converts skill Markdown to PNG sequences via [`engine/pixel/textprep.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/pixel/textprep.go) and [`engine/pixel/render.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/pixel/render.go) to reduce token usage for model interactions.
- Conversion fires only when the `--pixel` flag is present, the model is in the `CAVE_PIXEL_MODELS` allow-list, the "pages beat the text" size heuristic passes, and valid entitlements exist.
- The system replaces original skill bodies with JSON stubs containing base64-encoded image references while preserving metadata for discovery.
- Failed renders trigger a **byte-identical pass-through** that leaves the skill unchanged, ensuring no data loss on conversion errors.

## Frequently Asked Questions

### What file naming convention does pixel mode use for generated images?

Pixel mode generates files following the pattern `SKILL.pxN.png`, where `SKILL` matches the skill identifier and `N` represents the sequential page number (e.g., `SKILL.px1.png`, `SKILL.px2.png`). These files persist locally while their base64-encoded contents transmit to the model.

### Why would pixel mode refuse to convert a skill even with the --pixel flag?

The conversion aborts if the size heuristic determines that PNG pages would exceed the original text's token count, if the target model is not in the `CAVE_PIXEL_MODELS` allow-list, or if the account lacks required entitlements for model-gated compression features. In these cases, the engine falls back to byte-identical pass-through.

### Is the original skill content preserved after pixel conversion?

No, the original Markdown body is replaced with a JSON stub containing image references. However, the stub retains the skill name and front-matter to preserve discovery functionality. Note that conversion only proceeds after successful PNG generation; failures trigger a pass-through that leaves the original content completely unchanged.

### Which models support pixel mode by default?

The default `CAVE_PIXEL_MODELS` allow-list includes `claude-fable-5` and `gpt-5.6`. Administrators can modify this environment variable to enable pixel mode for additional model endpoints, subject to entitlement verification in the proxy layer.