# How to Batch Process Multiple Books for Skill Distillation with cangjie-skill

> Learn to batch process multiple books for skill distillation using cangjie-skill's seven-stage pipeline. Utilize PIPELINE_STATE.md checkpoints for safe restarts and efficient failure isolation.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: how-to-guide
- Published: 2026-08-14

---

**Batch process multiple books for skill distillation by running cangjie-skill's seven-stage pipeline in a loop over isolated book directories, using [`PIPELINE_STATE.md`](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md) checkpoints for safe restarts and failure isolation.**

The **cangjie-skill** repository transforms long-form sources—books, transcripts, podcasts, or courses—into atomic, executable Claude skills. When scaling from one book to dozens, the same deterministic pipeline applies repeatedly without interference between books. This guide covers the per-book architecture, practical driver implementations, and operational tips for reliable batch execution.

## Understanding the Seven-Stage Pipeline

Each book passes through identical stages defined in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) and the `methodology/` directory. The pipeline is **fully deterministic** and **resumable** via checkpoint tracking.

| Stage | Purpose | Output | Reference File |
|-------|---------|--------|----------------|
| **0 – Adler** | High-level book understanding | [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) | [`methodology/01-stage0-adler.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md) |
| **1 – Parallel Extraction** | Five sub-agents extract frameworks, principles, cases, counter-examples, glossary | `candidates/*.md` | [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) § Stage 1 |
| **1.5 – Triple Verification** | Human-in-the-loop quality filter | [`verified.md`](https://github.com/kangarooking/cangjie-skill/blob/main/verified.md) & `rejected/` | [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md) |
| **2 – RIA++ Construction** | Build individual skills from verified units | `books/<slug>/<skill-slug>/SKILL.md` | [`methodology/04-stage2-ria-plus.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/04-stage2-ria-plus.md) |
| **3 – Zettelkasten Linking** | Cross-skill knowledge graph | [`INDEX.md`](https://github.com/kangarooking/cangjie-skill/blob/main/INDEX.md) & [`GLOSSARY.md`](https://github.com/kangarooking/cangjie-skill/blob/main/GLOSSARY.md) | [`methodology/05-stage3-zettelkasten.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/05-stage3-zettelkasten.md) |
| **4 – Pressure Test** | Darwin-compatible test suite | [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) & [`test-results.md`](https://github.com/kangarooking/cangjie-skill/blob/main/test-results.md) | [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) |
| **5 – Delivery** | Final digest and skill installation | [`DIGEST.md`](https://github.com/kangarooking/cangjie-skill/blob/main/DIGEST.md) | [`methodology/07-stage5-deliver.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/07-stage5-deliver.md) |

Each stage updates `books/<slug>/PIPELINE_STATE.md` with a markdown checklist (e.g., `- [x] Stage 2 – RIA++ completed`). This file enables **idempotent restarts**: if a batch job crashes, re-running the same command resumes from the last completed stage rather than restarting from zero.

## Batch Processing Architecture

The key insight for batch processing multiple books: **isolate each book in its own directory tree** and invoke the pipeline sequentially or in parallel slices. Because [`PIPELINE_STATE.md`](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md) lives inside each book folder, failures never cascade across the batch.

### Required Directory Structure

```

books/
├── book-slug-1/
│   ├── metadata.yaml      # Human-readable reference

│   ├── source.txt         # Raw book text

│   └── PIPELINE_STATE.md  # Auto-generated checkpoint

├── book-slug-2/
│   └── ...
└── ...

```

## Bash Driver for Sequential Batch Processing

Create a master list file ([`books_to_distill.txt`](https://github.com/kangarooking/cangjie-skill/blob/main/books_to_distill.txt)) with space-separated fields: `slug`, `source_path`, `title`, `author`, `year`.

```bash
#!/usr/bin/env bash

# batch_distill.sh — Sequential batch processing driver

set -euo pipefail

LIST="books_to_distill.txt"

while read -r slug src title author year; do
  echo "=== Distilling: ${title} (${author}) ==="

  # Create isolated book directory

  mkdir -p "books/${slug}"
  
  # Write minimal metadata for human audit

  cat > "books/${slug}/metadata.yaml" <<EOF
title: "${title}"
author: "${author}"
year: "${year}"
source: "${src}"
EOF

  # Invoke cangjie-skill pipeline

  # Replace 'cangjie-skill' with actual entry point if different

  if cangjie-skill "books/${slug}" "${src}"; then
    echo "✅ ${title} completed successfully"
  else
    echo "⚠️  ${title} failed — check books/${slug}/PIPELINE_STATE.md"
    # Continue to next book; failure is isolated

  fi

done < "${LIST}"

```

Example [`books_to_distill.txt`](https://github.com/kangarooking/cangjie-skill/blob/main/books_to_distill.txt):

```

lean-startup  ./raw/lean-startup.txt  "The Lean Startup"  "Eric Ries"  2011
articulating-decisions ./raw/articulating-decisions.txt "Articulating Design Decisions" "Tom Greever" 2015

```

## Python Driver for Programmatic Control

For complex workflows—database integration, progress tracking, or API wrappers—use a Python driver with the same checkpoint-aware logic.

```python
import subprocess
from pathlib import Path
from dataclasses import dataclass


@dataclass
class BookSpec:
    slug: str
    source: Path
    title: str
    author: str
    year: str


def distill_one(spec: BookSpec) -> bool:
    """Run cangjie-skill pipeline for a single book. Returns success status."""
    book_dir = Path("books") / spec.slug
    book_dir.mkdir(parents=True, exist_ok=True)

    # Optional metadata for human reference

    (book_dir / "metadata.yaml").write_text(
        f"title: {spec.title}\n"
        f"author: {spec.author}\n"
        f"year: {spec.year}\n"
        f"source: {spec.source}\n"
    )

    # Invoke pipeline — checkpoint file handles restarts automatically

    cmd = ["cangjie-skill", str(book_dir), str(spec.source)]
    result = subprocess.run(cmd, capture_output=True, text=True)

    if result.returncode != 0:
        print(f"❌ {spec.title}: FAILED")
        print(f"   See {book_dir / 'PIPELINE_STATE.md'}")
        print(f"   stderr: {result.stderr[:500]}")
        return False
    
    print(f"✅ {spec.title}: COMPLETED")
    return True


def batch_from_tsv(list_path: Path) -> dict[str, bool]:
    """Parse TSV/space-delimited list and run batch."""
    results = {}
    
    with list_path.open() as f:
        for line_num, line in enumerate(f, 1):
            line = line.strip()
            if not line or line.startswith("#"):
                continue
            
            parts = line.split(maxsplit=4)
            if len(parts) != 5:
                print(f"⚠️  Skipping malformed line {line_num}: {line[:60]}...")
                continue
            
            slug, src, title, author, year = parts
            spec = BookSpec(slug, Path(src), title, author, year)
            results[slug] = distill_one(spec)
    
    return results


if __name__ == "__main__":
    summary = batch_from_tsv(Path("books_to_distill.txt"))
    completed = sum(summary.values())
    print(f"\nBatch complete: {completed}/{len(summary)} books successful")

```

## Parallel Execution Strategies

For large batches, distribute work across multiple processes. The checkpoint system prevents collisions.

### GNU Parallel (Simple)

```bash

# Split list into chunks, run 4 concurrent jobs

parallel --jobs 4 --colsep ' ' \
  'mkdir -p books/{1} && cangjie-skill books/{1} {2}' \
  :::: books_to_distill.txt

```

### Python with ProcessPoolExecutor

```python
from concurrent.futures import ProcessPoolExecutor, as_completed
import multiprocessing

def parallel_batch(specs: list[BookSpec], max_workers: int = 4) -> None:
    """Process books in parallel with isolated failures."""
    with ProcessPoolExecutor(max_workers=max_workers) as executor:
        futures = {
            executor.submit(distill_one, spec): spec 
            for spec in specs
        }
        
        for future in as_completed(futures):
            spec = futures[future]
            try:
                success = future.result()
            except Exception as e:
                print(f"💥 {spec.title}: EXCEPTION — {e}")

```

> **Safety guarantee**: Each worker operates on a distinct `books/<slug>/` directory. [`PIPELINE_STATE.md`](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md) ensures that even if two processes accidentally target the same slug (avoidable with proper partitioning), atomic updates prevent corruption.

## Operational Best Practices

| Practice | Implementation | Rationale |
|----------|---------------|-----------|
| **Pilot validation** | Run one complete book before batch | Catches environment issues (encoding, extractor prompts, model access) early |
| **Version control checkpoints** | `git add books/*/PIPELINE_STATE.md` | Creates auditable history of per-book progress |
| **Staging skill directory** | Use `~/.claude/skills-staging/` during batch | Prevents partially-tested skills from polluting production skills |
| **Disk monitoring** | Alert on `books/<slug>/candidates/` size | Stage 1 generates substantial intermediate files; clean `candidates/` post-verification to reclaim space |
| **Graceful degradation** | Automatic fallback to serial extraction | Per [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) lines 95-96, constrained hosts continue without parallel agents |

## Key Source Files and Templates

Reference these files when customizing batch behavior or debugging failures:

| File | Purpose | GitHub Link |
|------|---------|-------------|
| [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) | Master pipeline specification and checkpoint logic | [SKILL.md](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) |
| [`methodology/00-overview.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md) | Seven-stage conceptual overview | [00-overview.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md) |
| `templates/BOOK_OVERVIEW.md.template` | Stage 0 output template | [template](https://github.com/kangarooking/cangjie-skill/blob/main/templates/BOOK_OVERVIEW.md.template) |
| `templates/SKILL.md.template` | R-I-A1-A2-E-B skill skeleton | [template](https://github.com/kangarooking/cangjie-skill/blob/main/templates/SKILL.md.template) |
| `templates/INDEX.md.template` | Cross-skill graph generation | [template](https://github.com/kangarooking/cangjie-skill/blob/main/templates/INDEX.md.template) |
| [`extractors/framework-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md) | Framework extraction agent prompt | [extractor](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md) |

## Handling Pipeline Restarts and Failures

The [`PIPELINE_STATE.md`](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md) checkpoint uses a simple markdown checklist format:

```markdown

# Pipeline State: lean-startup

- [x] Stage 0 – Adler completed (2024-01-15T09:23:00Z)
- [x] Stage 1 – Parallel Extraction completed (2024-01-15T10:45:00Z)
- [x] Stage 1.5 – Triple Verification completed (2024-01-15T11:20:00Z)
- [ ] Stage 2 – RIA++ Construction

```

If Stage 2 fails or the process terminates, re-running `cangjie-skill books/lean-startup <source>` detects the partial state and resumes from Stage 2. No manual intervention required.

To force a restart from a specific stage, manually edit the checklist—useful after modifying extractor prompts or code.

## Summary

- **Batch processing** relies on looping over isolated `books/<slug>/` directories with a thin driver script
- **[`PIPELINE_STATE.md`](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md)** provides automatic checkpointing for safe restarts and failure isolation
- **Sequential and parallel drivers** work equally well; the checkpoint system prevents race conditions
- **Zero code changes** to cangjie-skill are required for batch operation—the existing architecture already supports it
- **Audit and staging practices** ensure reliable production deployments at scale

## Frequently Asked Questions

### Can I resume a failed batch without reprocessing completed books?

Yes. The driver scripts shown above are idempotent—re-running the same command skips books that reached Stage 5 (Delivery) and resumes interrupted books from their [`PIPELINE_STATE.md`](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md) checkpoint. Completed books report success immediately without redundant work.

### What happens if one book crashes during a batch?

The bash driver's `else` clause and Python's `try/except` isolation ensure the loop continues to the next book. The failing book's directory retains its partial state for debugging. Fix the root cause (corrupt source text, model rate limit), then re-run the batch—only that book resumes processing.

### How many books can I process in parallel?

Limited by API rate limits and local compute. Each book spawns up to five parallel extraction agents in Stage 1. For OpenAI or Anthropic APIs, start with 2-4 concurrent books. For local LLMs, scale to CPU/GPU capacity. The checkpoint system allows safe oversubscription—excess processes simply queue.

### Do I need to clean up intermediate files after batch completion?

Not required, but recommended for disk efficiency. The `candidates/` and `rejected/` directories contain raw extraction outputs that can exceed the source text size. Preserve [`verified.md`](https://github.com/kangarooking/cangjie-skill/blob/main/verified.md), [`test-results.md`](https://github.com/kangarooking/cangjie-skill/blob/main/test-results.md), and final skill directories; archive or delete `candidates/` after successful Stage 5 delivery.