How Progressive Disclosure Architecture Reduces Token Usage in Claude-Skills

Progressive disclosure architecture in claude-skills cuts initial token usage by 40–50% by splitting skills into a lightweight 80‑line SKILL.md front page and deeper reference files that load only when context requires them.

The claude-skills repository implements a strict two‑tier content model that treats token budgets as a finite resource. By separating essential metadata from deep technical detail, the architecture ensures that Claude processes only the information necessary for the current turn, deferring heavy content to subsequent lazy loads.

The Two-Tier Progressive Disclosure Architecture

The architecture is explicitly defined in CLAUDE.md as a tiered loading system designed to minimize context window consumption while maintaining full functional depth.

Tier 1: The Lean SKILL.md Front Page

The front page serves as the sole mandatory payload for every interaction. According to the canonical definition in CLAUDE.md (lines 91–106), this file is constrained to approximately 80–100 lines and contains:

  • Essential metadata and trigger descriptions
  • A five‑step core workflow
  • Hard constraints and guardrails
  • A routing table that maps topics to deeper reference files

By keeping this file deliberately short, the system achieves the stated goal of a 50% token reduction on the initial prompt compared to monolithic skill definitions.

Tier 2: Lazy-Loaded Reference Files

Deep technical content lives in standalone files under the references/ directory. As documented in CLAUDE.md (lines 100–105), these files range from 100–600 lines and contain:

  • Full API specifications and code samples
  • Edge‑case handling procedures
  • Complex tables and configuration schemas

The runtime loads these files only when the routing table determines that the user's query matches the "Load When" conditions, ensuring that 600‑line technical documents do not consume tokens unless specifically required.

How the Routing Table Enables On-Demand Loading

The routing mechanism is the operational core of the progressive disclosure architecture. Implemented as a markdown table within SKILL.md, it defines the lazy‑loading logic that prevents unnecessary token consumption.

The routing table structure, as specified in CONTRIBUTING.md (lines 96–104), contains three columns:

  • Topic: Human‑readable subject name
  • Reference: Relative path to the markdown file (e.g., references/state-management.md)
  • Load When: Trigger keywords or contextual conditions that activate the load

When Claude receives a request, the runtime evaluates the "Load When" column against the current context. Only when a match occurs does the system stream the associated reference file into the context window. This on‑demand pulling is performed by the Claude plugin runtime and does not increase the token count of the initial prompt, as confirmed in the architecture documentation.

Measuring the Token Reduction Impact

The claude-skills repository provides concrete metrics demonstrating the efficiency gains of progressive disclosure architecture.

Configuration Approximate Tokens Reduction
Full skill inline (all docs) 2,000–3,000 Baseline
Split: SKILL.md + refs 1,000–1,500 ~50%
Real‑world measurement 1,300 vs 650 50%

As documented in CONTRIBUTING.md (lines 33–36), the architecture achieves a 40–50% reduction in initial token load by ensuring that only the 80‑line front page is always sent, while deeper references are streamed only when context requires them.

Real-World Implementation Examples

The skills/code-documenter skill demonstrates the practical application of progressive disclosure architecture.

Example SKILL.md with Routing Table


# Code Documenter – Overview

## Core Workflow

1. Discover entry points
2. Diagnose complexity
3. Plan documentation structure
4. Execute docstring generation
5. Review for accuracy

## Reference Guide

| Topic          | Reference                        | Load When                     |
|----------------|----------------------------------|------------------------------|
| Python Docs    | `references/python-docstrings.md` | Python, docstrings, Sphinx   |
| JS/TS Docs     | `references/jsdoc.md`           | JavaScript, TypeScript, JSDoc|
| API Specs      | `references/openapi.md`         | REST, OpenAPI, swagger       |

This routing table ensures that references/python-docstrings.md—which contains 130+ lines of detailed Sphinx and Google style examples—is only loaded when the user query contains "Python" or "docstrings".

Example Reference File Structure

The file skills/code-documenter/references/python-docstrings.md contains deep technical content that remains outside the initial token budget:


# Python Docstrings

> Reference for: Code Documenter  
> Load when: Using Python, docstrings, or Sphinx  

## Google Style Guide Example

```python
def fetch_user(user_id: int, include_meta: bool = False) -> dict:
    """
    Retrieve a user record from the database.

    Args:
        user_id: The unique identifier for the user.
        include_meta: Whether to include metadata fields.

    Returns:
        A dictionary containing user data.

    Raises:
        ValueError: If user_id is negative.
    """
    ...

By segregating this detailed content into `references/`, the progressive disclosure architecture ensures that the initial prompt contains only the 80‑line [`SKILL.md`](https://github.com/Jeffallan/claude-skills/blob/main/SKILL.md), cutting the token payload by half while preserving access to full technical depth when required.

## Summary

- **Progressive disclosure architecture** splits claude-skills into an 80‑100 line [`SKILL.md`](https://github.com/Jeffallan/claude-skills/blob/main/SKILL.md) front page and deeper reference files (100‑600 lines) stored under `references/`.
- A **routing table** in [`SKILL.md`](https://github.com/Jeffallan/claude-skills/blob/main/SKILL.md) maps topics to reference files and defines "Load When" triggers, enabling **lazy loading** of technical content only when context requires it.
- This architecture achieves a **40–50% reduction in initial token usage** (from ~1,300 tokens to ~650 tokens) by ensuring that heavy technical documentation is not included in the initial prompt.
- The design is implemented in [`CLAUDE.md`](https://github.com/Jeffallan/claude-skills/blob/main/CLAUDE.md) (lines 91–106) and demonstrated in production skills such as `skills/code-documenter/`.

## Frequently Asked Questions

### How does the routing table decide which reference files to load?

The routing table evaluates the "Load When" column against the current user query and conversation context. When keywords or semantic conditions match—such as "Python" or "Redux"—the Claude plugin runtime streams the corresponding reference file into the context window. This evaluation happens after the initial prompt, ensuring that reference tokens are only consumed when specifically needed.

### What is the typical file size difference between SKILL.md and reference files?

According to the architecture specification in [`CLAUDE.md`](https://github.com/Jeffallan/claude-skills/blob/main/CLAUDE.md), [`SKILL.md`](https://github.com/Jeffallan/claude-skills/blob/main/SKILL.md) is strictly constrained to approximately 80–100 lines, while reference files range from 100–600 lines each. In practice, this means a reference file can be 6–7 times larger than the front page, making lazy loading essential for token efficiency.

### Can progressive disclosure architecture be applied to other LLM agent systems?

Yes, the pattern is framework-agnostic. Any LLM system that supports dynamic context injection or multi-turn conversation management can implement this two-tier structure. The key requirements are: (1) a lightweight manifest or system prompt that fits within the initial token budget, and (2) a mechanism to conditionally append deeper documentation based on user intent, similar to the routing table implementation in claude-skills.

### Where is the token reduction percentage documented in the repository?

The 40–50% token reduction metric is explicitly documented in [`CONTRIBUTING.md`](https://github.com/Jeffallan/claude-skills/blob/main/CONTRIBUTING.md) at lines 33–36, which states the goal of "50% token reduction through selective loading." Additionally, [`CLAUDE.md`](https://github.com/Jeffallan/claude-skills/blob/main/CLAUDE.md) (lines 91–106) defines the tiered architecture that enables this efficiency, and real-world measurements comparing 1,300 tokens (full) versus 650 tokens (tiered) are referenced in the contribution guidelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →