How to Batch Process Multiple Books for Skill Distillation with cangjie-skill

Batch process multiple books for skill distillation by running cangjie-skill's seven-stage pipeline in a loop over isolated book directories, using PIPELINE_STATE.md checkpoints for safe restarts and failure isolation.

The cangjie-skill repository transforms long-form sources—books, transcripts, podcasts, or courses—into atomic, executable Claude skills. When scaling from one book to dozens, the same deterministic pipeline applies repeatedly without interference between books. This guide covers the per-book architecture, practical driver implementations, and operational tips for reliable batch execution.

Understanding the Seven-Stage Pipeline

Each book passes through identical stages defined in SKILL.md and the methodology/ directory. The pipeline is fully deterministic and resumable via checkpoint tracking.

Stage Purpose Output Reference File
0 – Adler High-level book understanding BOOK_OVERVIEW.md methodology/01-stage0-adler.md
1 – Parallel Extraction Five sub-agents extract frameworks, principles, cases, counter-examples, glossary candidates/*.md SKILL.md § Stage 1
1.5 – Triple Verification Human-in-the-loop quality filter verified.md & rejected/ methodology/03-stage1.5-triple-verify.md
2 – RIA++ Construction Build individual skills from verified units books/<slug>/<skill-slug>/SKILL.md methodology/04-stage2-ria-plus.md
3 – Zettelkasten Linking Cross-skill knowledge graph INDEX.md & GLOSSARY.md methodology/05-stage3-zettelkasten.md
4 – Pressure Test Darwin-compatible test suite test-prompts.json & test-results.md methodology/06-stage4-pressure-test.md
5 – Delivery Final digest and skill installation DIGEST.md methodology/07-stage5-deliver.md

Each stage updates books/<slug>/PIPELINE_STATE.md with a markdown checklist (e.g., - [x] Stage 2 – RIA++ completed). This file enables idempotent restarts: if a batch job crashes, re-running the same command resumes from the last completed stage rather than restarting from zero.

Batch Processing Architecture

The key insight for batch processing multiple books: isolate each book in its own directory tree and invoke the pipeline sequentially or in parallel slices. Because PIPELINE_STATE.md lives inside each book folder, failures never cascade across the batch.

Required Directory Structure


books/
├── book-slug-1/
│   ├── metadata.yaml      # Human-readable reference

│   ├── source.txt         # Raw book text

│   └── PIPELINE_STATE.md  # Auto-generated checkpoint

├── book-slug-2/
│   └── ...
└── ...

Bash Driver for Sequential Batch Processing

Create a master list file (books_to_distill.txt) with space-separated fields: slug, source_path, title, author, year.

#!/usr/bin/env bash

# batch_distill.sh — Sequential batch processing driver

set -euo pipefail

LIST="books_to_distill.txt"

while read -r slug src title author year; do
  echo "=== Distilling: ${title} (${author}) ==="

  # Create isolated book directory

  mkdir -p "books/${slug}"
  
  # Write minimal metadata for human audit

  cat > "books/${slug}/metadata.yaml" <<EOF
title: "${title}"
author: "${author}"
year: "${year}"
source: "${src}"
EOF

  # Invoke cangjie-skill pipeline

  # Replace 'cangjie-skill' with actual entry point if different

  if cangjie-skill "books/${slug}" "${src}"; then
    echo "✅ ${title} completed successfully"
  else
    echo "⚠️  ${title} failed — check books/${slug}/PIPELINE_STATE.md"
    # Continue to next book; failure is isolated

  fi

done < "${LIST}"

Example books_to_distill.txt:


lean-startup  ./raw/lean-startup.txt  "The Lean Startup"  "Eric Ries"  2011
articulating-decisions ./raw/articulating-decisions.txt "Articulating Design Decisions" "Tom Greever" 2015

Python Driver for Programmatic Control

For complex workflows—database integration, progress tracking, or API wrappers—use a Python driver with the same checkpoint-aware logic.

import subprocess
from pathlib import Path
from dataclasses import dataclass


@dataclass
class BookSpec:
    slug: str
    source: Path
    title: str
    author: str
    year: str


def distill_one(spec: BookSpec) -> bool:
    """Run cangjie-skill pipeline for a single book. Returns success status."""
    book_dir = Path("books") / spec.slug
    book_dir.mkdir(parents=True, exist_ok=True)

    # Optional metadata for human reference

    (book_dir / "metadata.yaml").write_text(
        f"title: {spec.title}\n"
        f"author: {spec.author}\n"
        f"year: {spec.year}\n"
        f"source: {spec.source}\n"
    )

    # Invoke pipeline — checkpoint file handles restarts automatically

    cmd = ["cangjie-skill", str(book_dir), str(spec.source)]
    result = subprocess.run(cmd, capture_output=True, text=True)

    if result.returncode != 0:
        print(f"❌ {spec.title}: FAILED")
        print(f"   See {book_dir / 'PIPELINE_STATE.md'}")
        print(f"   stderr: {result.stderr[:500]}")
        return False
    
    print(f"✅ {spec.title}: COMPLETED")
    return True


def batch_from_tsv(list_path: Path) -> dict[str, bool]:
    """Parse TSV/space-delimited list and run batch."""
    results = {}
    
    with list_path.open() as f:
        for line_num, line in enumerate(f, 1):
            line = line.strip()
            if not line or line.startswith("#"):
                continue
            
            parts = line.split(maxsplit=4)
            if len(parts) != 5:
                print(f"⚠️  Skipping malformed line {line_num}: {line[:60]}...")
                continue
            
            slug, src, title, author, year = parts
            spec = BookSpec(slug, Path(src), title, author, year)
            results[slug] = distill_one(spec)
    
    return results


if __name__ == "__main__":
    summary = batch_from_tsv(Path("books_to_distill.txt"))
    completed = sum(summary.values())
    print(f"\nBatch complete: {completed}/{len(summary)} books successful")

Parallel Execution Strategies

For large batches, distribute work across multiple processes. The checkpoint system prevents collisions.

GNU Parallel (Simple)


# Split list into chunks, run 4 concurrent jobs

parallel --jobs 4 --colsep ' ' \
  'mkdir -p books/{1} && cangjie-skill books/{1} {2}' \
  :::: books_to_distill.txt

Python with ProcessPoolExecutor

from concurrent.futures import ProcessPoolExecutor, as_completed
import multiprocessing

def parallel_batch(specs: list[BookSpec], max_workers: int = 4) -> None:
    """Process books in parallel with isolated failures."""
    with ProcessPoolExecutor(max_workers=max_workers) as executor:
        futures = {
            executor.submit(distill_one, spec): spec 
            for spec in specs
        }
        
        for future in as_completed(futures):
            spec = futures[future]
            try:
                success = future.result()
            except Exception as e:
                print(f"💥 {spec.title}: EXCEPTION — {e}")

Safety guarantee: Each worker operates on a distinct books/<slug>/ directory. PIPELINE_STATE.md ensures that even if two processes accidentally target the same slug (avoidable with proper partitioning), atomic updates prevent corruption.

Operational Best Practices

Practice Implementation Rationale
Pilot validation Run one complete book before batch Catches environment issues (encoding, extractor prompts, model access) early
Version control checkpoints git add books/*/PIPELINE_STATE.md Creates auditable history of per-book progress
Staging skill directory Use ~/.claude/skills-staging/ during batch Prevents partially-tested skills from polluting production skills
Disk monitoring Alert on books/<slug>/candidates/ size Stage 1 generates substantial intermediate files; clean candidates/ post-verification to reclaim space
Graceful degradation Automatic fallback to serial extraction Per SKILL.md lines 95-96, constrained hosts continue without parallel agents

Key Source Files and Templates

Reference these files when customizing batch behavior or debugging failures:

File Purpose GitHub Link
SKILL.md Master pipeline specification and checkpoint logic SKILL.md
methodology/00-overview.md Seven-stage conceptual overview 00-overview.md
templates/BOOK_OVERVIEW.md.template Stage 0 output template template
templates/SKILL.md.template R-I-A1-A2-E-B skill skeleton template
templates/INDEX.md.template Cross-skill graph generation template
extractors/framework-extractor.md Framework extraction agent prompt extractor

Handling Pipeline Restarts and Failures

The PIPELINE_STATE.md checkpoint uses a simple markdown checklist format:


# Pipeline State: lean-startup

- [x] Stage 0 – Adler completed (2024-01-15T09:23:00Z)
- [x] Stage 1 – Parallel Extraction completed (2024-01-15T10:45:00Z)
- [x] Stage 1.5 – Triple Verification completed (2024-01-15T11:20:00Z)
- [ ] Stage 2 – RIA++ Construction

If Stage 2 fails or the process terminates, re-running cangjie-skill books/lean-startup <source> detects the partial state and resumes from Stage 2. No manual intervention required.

To force a restart from a specific stage, manually edit the checklist—useful after modifying extractor prompts or code.

Summary

  • Batch processing relies on looping over isolated books/<slug>/ directories with a thin driver script
  • PIPELINE_STATE.md provides automatic checkpointing for safe restarts and failure isolation
  • Sequential and parallel drivers work equally well; the checkpoint system prevents race conditions
  • Zero code changes to cangjie-skill are required for batch operation—the existing architecture already supports it
  • Audit and staging practices ensure reliable production deployments at scale

Frequently Asked Questions

Can I resume a failed batch without reprocessing completed books?

Yes. The driver scripts shown above are idempotent—re-running the same command skips books that reached Stage 5 (Delivery) and resumes interrupted books from their PIPELINE_STATE.md checkpoint. Completed books report success immediately without redundant work.

What happens if one book crashes during a batch?

The bash driver's else clause and Python's try/except isolation ensure the loop continues to the next book. The failing book's directory retains its partial state for debugging. Fix the root cause (corrupt source text, model rate limit), then re-run the batch—only that book resumes processing.

How many books can I process in parallel?

Limited by API rate limits and local compute. Each book spawns up to five parallel extraction agents in Stage 1. For OpenAI or Anthropic APIs, start with 2-4 concurrent books. For local LLMs, scale to CPU/GPU capacity. The checkpoint system allows safe oversubscription—excess processes simply queue.

Do I need to clean up intermediate files after batch completion?

Not required, but recommended for disk efficiency. The candidates/ and rejected/ directories contain raw extraction outputs that can exceed the source text size. Preserve verified.md, test-results.md, and final skill directories; archive or delete candidates/ after successful Stage 5 delivery.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →