How book-to-skill Prioritizes Compile-Time Over Runtime Costs: An Architecture Deep Dive
book-to-skill pays the expensive cost of parsing and structuring documents once at compile time, so runtime LLM queries only load ~5 KB of relevant content instead of 200K+ tokens from a full book.
The open-source book-to-skill tool transforms unwieldy technical books into compact, queryable skill files. Its core innovation lies in deliberately front-loading computational work to minimize recurring costs. This article examines the specific mechanisms that enforce this compile-time over runtime design principle, with direct reference to the implementation in virgiliojr94/book-to-skill.
The Design Principle: Compile-Time Over Runtime
The book-to-skill architecture explicitly lists "Compile-time over runtime" as its second foundational design principle. According to docs/architecture.md lines 56-60, the cost of navigation and document structuring is paid once during extraction. At query time, only the relevant chapter is loaded—never the full source material.
This principle addresses a concrete economic problem: LLM token costs scale with input size. A 400-page PDF converted to text costs approximately 200,000 tokens per API call if dumped naively. By restructuring the content, book-to-skill reduces this to roughly 5,000 tokens per targeted query.
The Compile-Time Extraction Pipeline
The heavy lifting occurs in two key files that execute deterministically during the build phase.
scripts/extract.py: The Entry Point
This thin wrapper invokes the extraction logic from the command line:
python -m book_to_skill extract path/to/book.pdf
Under the hood, this delegates to book_to_skill/cli.py, which implements the core transformation.
book_to_skill/cli.py: The Deterministic Transformer
As documented in docs/architecture.md lines 12-23 and 48-51, the CLI performs three critical operations:
- Parses source files (PDF, EPUB, or markdown) into raw text
- Merges content into
full_text.txtfor processing - Generates output structure:
SKILL.mdcore: ~4 KB metadata and navigation- Per-chapter markdown files: ~1 KB each
This extraction is idempotent—running it twice on the same source produces identical outputs. The deterministic nature ensures reproducible builds and cache-friendly deployments.
Runtime Efficiency: Load Only What You Need
The runtime payload is intentionally minimal. As explained in docs/faq.md lines 8-13 and 34-38, query-time behavior follows a strict pattern:
- Load the 4 KB SKILL.md core (cached in memory)
- Load one 1 KB chapter file based on the query
- Leave all other chapters on disk, incurring zero token cost
This contrasts sharply with naive approaches that re-process or re-transmit full source documents on every call.
Runtime Code Example
from pathlib import Path
# Load the compact core (already in memory for every call)
core_path = Path.home() / ".agents" / "skills" / "my_book" / "SKILL.md"
core = core_path.read_text()
def load_chapter(chapter_name: str) -> str:
"""Load a single chapter on demand. This incurs a tiny token cost."""
chapter_path = Path.home() / ".agents" / "skills" / "my_book" / "chapters" / f"{chapter_name}.md"
return chapter_path.read_text()
# Example usage at query time
question = "How does the Observer pattern work?"
chapter = "observer_pattern" # Determined by the agent
text = core + "\n" + load_chapter(chapter)
# `text` is now sent to the LLM – only ~5 KB total.
Cost Comparison: Compile-Time vs. Runtime Strategies
| Approach | Per-Call Token Cost | Build Cost | Total Cost (1000 queries) |
|---|---|---|---|
| Naive full-book dump | ~200,000 tokens | None | 200M tokens |
| book-to-skill | ~5,000 tokens | One-time extraction | 5M tokens |
The 40x reduction comes from amortizing the expensive parsing work. The one-time build cost is negligible compared to cumulative runtime savings at scale.
Key Files Supporting This Architecture
| File | Function |
|---|---|
scripts/extract.py |
CLI entry point for compile-time extraction |
book_to_skill/cli.py |
Core implementation of parsing, merging, and output generation |
docs/architecture.md |
Design documentation including the compile-time principle (lines 56-60) |
docs/faq.md |
Economic justification for the approach (lines 8-13, 34-38) |
Summary
- Explicit principle: "Compile-time over runtime" is codified in
docs/architecture.mdas a core architectural commitment - One-time extraction:
book_to_skill/cli.pyprocesses source documents into a ~4 KB core plus ~1 KB chapter files - Minimal runtime: Only the SKILL.md core and one chapter load per query, reducing token costs by approximately 40x
- Economic leverage: Front-loaded work eliminates recurring charges that would dominate at scale
Frequently Asked Questions
What specific files are generated at compile time?
The extractor produces three artifacts: full_text.txt (complete merged source), SKILL.md (compact 4 KB navigation core), and individual chapter markdown files (~1 KB each) in a chapters/ subdirectory. Only the latter two are used at runtime.
Why not process the book on every query?
Re-parsing a 400-page PDF or re-tokenizing 200,000 characters on each API call would multiply costs proportionally. As stated in docs/faq.md, compile-time extraction "amortizes that cost once and eliminates repeated token-billing at runtime."
Is the extraction process deterministic?
Yes. Given identical source files, book_to_skill/cli.py produces byte-identical outputs. This enables reproducible builds and reliable caching in CI/CD pipelines.
How does an agent know which chapter to load?
The SKILL.md core contains structured metadata mapping concepts to chapter filenames. Agents use this index to resolve queries to specific chapter paths without loading content prematurely.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →