How to Batch Process Multiple Books for Skill Distillation with cangjie-skill
Batch process multiple books for skill distillation by running cangjie-skill's seven-stage pipeline in a loop over isolated book directories, using PIPELINE_STATE.md checkpoints for safe restarts and failure isolation.
The cangjie-skill repository transforms long-form sources—books, transcripts, podcasts, or courses—into atomic, executable Claude skills. When scaling from one book to dozens, the same deterministic pipeline applies repeatedly without interference between books. This guide covers the per-book architecture, practical driver implementations, and operational tips for reliable batch execution.
Understanding the Seven-Stage Pipeline
Each book passes through identical stages defined in SKILL.md and the methodology/ directory. The pipeline is fully deterministic and resumable via checkpoint tracking.
| Stage | Purpose | Output | Reference File |
|---|---|---|---|
| 0 – Adler | High-level book understanding | BOOK_OVERVIEW.md |
methodology/01-stage0-adler.md |
| 1 – Parallel Extraction | Five sub-agents extract frameworks, principles, cases, counter-examples, glossary | candidates/*.md |
SKILL.md § Stage 1 |
| 1.5 – Triple Verification | Human-in-the-loop quality filter | verified.md & rejected/ |
methodology/03-stage1.5-triple-verify.md |
| 2 – RIA++ Construction | Build individual skills from verified units | books/<slug>/<skill-slug>/SKILL.md |
methodology/04-stage2-ria-plus.md |
| 3 – Zettelkasten Linking | Cross-skill knowledge graph | INDEX.md & GLOSSARY.md |
methodology/05-stage3-zettelkasten.md |
| 4 – Pressure Test | Darwin-compatible test suite | test-prompts.json & test-results.md |
methodology/06-stage4-pressure-test.md |
| 5 – Delivery | Final digest and skill installation | DIGEST.md |
methodology/07-stage5-deliver.md |
Each stage updates books/<slug>/PIPELINE_STATE.md with a markdown checklist (e.g., - [x] Stage 2 – RIA++ completed). This file enables idempotent restarts: if a batch job crashes, re-running the same command resumes from the last completed stage rather than restarting from zero.
Batch Processing Architecture
The key insight for batch processing multiple books: isolate each book in its own directory tree and invoke the pipeline sequentially or in parallel slices. Because PIPELINE_STATE.md lives inside each book folder, failures never cascade across the batch.
Required Directory Structure
books/
├── book-slug-1/
│ ├── metadata.yaml # Human-readable reference
│ ├── source.txt # Raw book text
│ └── PIPELINE_STATE.md # Auto-generated checkpoint
├── book-slug-2/
│ └── ...
└── ...
Bash Driver for Sequential Batch Processing
Create a master list file (books_to_distill.txt) with space-separated fields: slug, source_path, title, author, year.
#!/usr/bin/env bash
# batch_distill.sh — Sequential batch processing driver
set -euo pipefail
LIST="books_to_distill.txt"
while read -r slug src title author year; do
echo "=== Distilling: ${title} (${author}) ==="
# Create isolated book directory
mkdir -p "books/${slug}"
# Write minimal metadata for human audit
cat > "books/${slug}/metadata.yaml" <<EOF
title: "${title}"
author: "${author}"
year: "${year}"
source: "${src}"
EOF
# Invoke cangjie-skill pipeline
# Replace 'cangjie-skill' with actual entry point if different
if cangjie-skill "books/${slug}" "${src}"; then
echo "✅ ${title} completed successfully"
else
echo "⚠️ ${title} failed — check books/${slug}/PIPELINE_STATE.md"
# Continue to next book; failure is isolated
fi
done < "${LIST}"
Example books_to_distill.txt:
lean-startup ./raw/lean-startup.txt "The Lean Startup" "Eric Ries" 2011
articulating-decisions ./raw/articulating-decisions.txt "Articulating Design Decisions" "Tom Greever" 2015
Python Driver for Programmatic Control
For complex workflows—database integration, progress tracking, or API wrappers—use a Python driver with the same checkpoint-aware logic.
import subprocess
from pathlib import Path
from dataclasses import dataclass
@dataclass
class BookSpec:
slug: str
source: Path
title: str
author: str
year: str
def distill_one(spec: BookSpec) -> bool:
"""Run cangjie-skill pipeline for a single book. Returns success status."""
book_dir = Path("books") / spec.slug
book_dir.mkdir(parents=True, exist_ok=True)
# Optional metadata for human reference
(book_dir / "metadata.yaml").write_text(
f"title: {spec.title}\n"
f"author: {spec.author}\n"
f"year: {spec.year}\n"
f"source: {spec.source}\n"
)
# Invoke pipeline — checkpoint file handles restarts automatically
cmd = ["cangjie-skill", str(book_dir), str(spec.source)]
result = subprocess.run(cmd, capture_output=True, text=True)
if result.returncode != 0:
print(f"❌ {spec.title}: FAILED")
print(f" See {book_dir / 'PIPELINE_STATE.md'}")
print(f" stderr: {result.stderr[:500]}")
return False
print(f"✅ {spec.title}: COMPLETED")
return True
def batch_from_tsv(list_path: Path) -> dict[str, bool]:
"""Parse TSV/space-delimited list and run batch."""
results = {}
with list_path.open() as f:
for line_num, line in enumerate(f, 1):
line = line.strip()
if not line or line.startswith("#"):
continue
parts = line.split(maxsplit=4)
if len(parts) != 5:
print(f"⚠️ Skipping malformed line {line_num}: {line[:60]}...")
continue
slug, src, title, author, year = parts
spec = BookSpec(slug, Path(src), title, author, year)
results[slug] = distill_one(spec)
return results
if __name__ == "__main__":
summary = batch_from_tsv(Path("books_to_distill.txt"))
completed = sum(summary.values())
print(f"\nBatch complete: {completed}/{len(summary)} books successful")
Parallel Execution Strategies
For large batches, distribute work across multiple processes. The checkpoint system prevents collisions.
GNU Parallel (Simple)
# Split list into chunks, run 4 concurrent jobs
parallel --jobs 4 --colsep ' ' \
'mkdir -p books/{1} && cangjie-skill books/{1} {2}' \
:::: books_to_distill.txt
Python with ProcessPoolExecutor
from concurrent.futures import ProcessPoolExecutor, as_completed
import multiprocessing
def parallel_batch(specs: list[BookSpec], max_workers: int = 4) -> None:
"""Process books in parallel with isolated failures."""
with ProcessPoolExecutor(max_workers=max_workers) as executor:
futures = {
executor.submit(distill_one, spec): spec
for spec in specs
}
for future in as_completed(futures):
spec = futures[future]
try:
success = future.result()
except Exception as e:
print(f"💥 {spec.title}: EXCEPTION — {e}")
Safety guarantee: Each worker operates on a distinct
books/<slug>/directory.PIPELINE_STATE.mdensures that even if two processes accidentally target the same slug (avoidable with proper partitioning), atomic updates prevent corruption.
Operational Best Practices
| Practice | Implementation | Rationale |
|---|---|---|
| Pilot validation | Run one complete book before batch | Catches environment issues (encoding, extractor prompts, model access) early |
| Version control checkpoints | git add books/*/PIPELINE_STATE.md |
Creates auditable history of per-book progress |
| Staging skill directory | Use ~/.claude/skills-staging/ during batch |
Prevents partially-tested skills from polluting production skills |
| Disk monitoring | Alert on books/<slug>/candidates/ size |
Stage 1 generates substantial intermediate files; clean candidates/ post-verification to reclaim space |
| Graceful degradation | Automatic fallback to serial extraction | Per SKILL.md lines 95-96, constrained hosts continue without parallel agents |
Key Source Files and Templates
Reference these files when customizing batch behavior or debugging failures:
| File | Purpose | GitHub Link |
|---|---|---|
SKILL.md |
Master pipeline specification and checkpoint logic | SKILL.md |
methodology/00-overview.md |
Seven-stage conceptual overview | 00-overview.md |
templates/BOOK_OVERVIEW.md.template |
Stage 0 output template | template |
templates/SKILL.md.template |
R-I-A1-A2-E-B skill skeleton | template |
templates/INDEX.md.template |
Cross-skill graph generation | template |
extractors/framework-extractor.md |
Framework extraction agent prompt | extractor |
Handling Pipeline Restarts and Failures
The PIPELINE_STATE.md checkpoint uses a simple markdown checklist format:
# Pipeline State: lean-startup
- [x] Stage 0 – Adler completed (2024-01-15T09:23:00Z)
- [x] Stage 1 – Parallel Extraction completed (2024-01-15T10:45:00Z)
- [x] Stage 1.5 – Triple Verification completed (2024-01-15T11:20:00Z)
- [ ] Stage 2 – RIA++ Construction
If Stage 2 fails or the process terminates, re-running cangjie-skill books/lean-startup <source> detects the partial state and resumes from Stage 2. No manual intervention required.
To force a restart from a specific stage, manually edit the checklist—useful after modifying extractor prompts or code.
Summary
- Batch processing relies on looping over isolated
books/<slug>/directories with a thin driver script PIPELINE_STATE.mdprovides automatic checkpointing for safe restarts and failure isolation- Sequential and parallel drivers work equally well; the checkpoint system prevents race conditions
- Zero code changes to cangjie-skill are required for batch operation—the existing architecture already supports it
- Audit and staging practices ensure reliable production deployments at scale
Frequently Asked Questions
Can I resume a failed batch without reprocessing completed books?
Yes. The driver scripts shown above are idempotent—re-running the same command skips books that reached Stage 5 (Delivery) and resumes interrupted books from their PIPELINE_STATE.md checkpoint. Completed books report success immediately without redundant work.
What happens if one book crashes during a batch?
The bash driver's else clause and Python's try/except isolation ensure the loop continues to the next book. The failing book's directory retains its partial state for debugging. Fix the root cause (corrupt source text, model rate limit), then re-run the batch—only that book resumes processing.
How many books can I process in parallel?
Limited by API rate limits and local compute. Each book spawns up to five parallel extraction agents in Stage 1. For OpenAI or Anthropic APIs, start with 2-4 concurrent books. For local LLMs, scale to CPU/GPU capacity. The checkpoint system allows safe oversubscription—excess processes simply queue.
Do I need to clean up intermediate files after batch completion?
Not required, but recommended for disk efficiency. The candidates/ and rejected/ directories contain raw extraction outputs that can exceed the source text size. Preserve verified.md, test-results.md, and final skill directories; archive or delete candidates/ after successful Stage 5 delivery.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →