How to Batch Process Multiple Books with Quality Control in Cangjie-Skill
Batch processing books in Cangjie-Skill relies on the RIA-TV++ pipeline, which automatically enforces quality control through triple verification and pressure testing for every book in the batch.
Cangjie-Skill transforms long-form content—books, videos, podcasts, and courses—into structured, reusable AI Skills. To batch process multiple books with quality control in cangjie-skill, you iterate the seven-stage RIA-TV++ pipeline across your source library, letting embedded validation gates filter raw extractions into verified, testable skills ready for deployment.
The RIA-TV++ Pipeline Architecture
The batch processing workflow runs the following seven stages for every source file. Quality control is baked into stages 3 and 6, ensuring each book meets the same rigorous standards before delivery.
Stage 1: Adler Analysis – Overall Comprehension
The pipeline begins by parsing the entire work into structure, explanation, critique, and application. This generates BOOK_OVERVIEW.md, providing a high-level anchor that ensures every batch entry has a consistent baseline for downstream processing.
Stage 2: Parallel Extraction
Five specialized extractors run simultaneously against the raw text: framework, principle, case, counter-example, and glossary extractors. Located in the extractors/ directory (e.g., extractors/framework-extractor.md), these components speed up batch processing by generating candidate methodology units in parallel. Details on invocation logic are documented in methodology/02-stage1-parallel-extract.md.
Stage 3: Triple Verification
This quality gate applies a three-fold screening criteria to every candidate unit:
- Citation requirement: Must have ≥2 independent citations in the source text
- Novelty requirement: Must answer a new predictive question
- Uniqueness requirement: Must be non-trivial (not common knowledge)
As defined in methodology/03-stage1.5-triple-verify.md, this triage typically cuts raw output to 25–50%, providing a strong quality filter for each book in the batch.
Stage 4: RIA++ Construction
For each verified unit, the pipeline constructs a structured skill in R-I-A-A-E-B format:
- R: Raw quote
- I: Paraphrase
- A1: Example
- A2: Future scenario
- E: Executable steps
- B: Boundaries
This format makes each skill self-contained and testable.
Stage 5: Zettelkasten Linking
The system detects dependencies and complementary relationships between skills, generating INDEX.md and a visual skill map. This enables cross-skill reuse and identifies redundant units across the batch.
Stage 6: Pressure Test
Automated bait-question testing generates test cases that mix skills to verify they work in isolation and together. Failures trigger an automatic re-run of earlier stages. According to methodology/06-stage4-pressure-test.md, this stage catches subtle execution errors before release.
Stage 7: Delivery
The pipeline emits DIGEST.md (a human-readable summary) and installs verified skills into Claude Code or Cursor directories, providing polished artifacts ready for deployment.
Implementing Batch Processing
To process dozens of books, wrap the pipeline in a simple orchestration loop. The following Python helper walks a directory of source books and invokes the Cangjie-Skill CLI (referenced here as cangjie-cli) for each file, collecting quality metrics in a consolidated report.
import pathlib
import subprocess
import json
# ----------------------------------------------------------------------
# 1️⃣ Locate all book files (e.g., *.txt, *.md) in the source folder
# ----------------------------------------------------------------------
BOOK_ROOT = pathlib.Path("/path/to/books")
OUTPUT_ROOT = pathlib.Path("/path/to/output")
OUTPUT_ROOT.mkdir(parents=True, exist_ok=True)
def run_pipeline(book_path: pathlib.Path) -> dict:
"""
Executes the Cangjie‑Skill pipeline for a single book.
Returns a JSON summary of the run (generated by the CLI).
"""
cmd = [
"cangjie-cli", # <-- replace with the real entrypoint
"--input", str(book_path),
"--output", str(OUTPUT_ROOT / book_path.stem),
"--format", "json" # ask the tool to emit a machine‑readable report
]
result = subprocess.run(
cmd, capture_output=True, text=True, check=False
)
if result.returncode != 0:
raise RuntimeError(f"Pipeline failed for {book_path.name}:\n{result.stderr}")
return json.loads(result.stdout)
# ----------------------------------------------------------------------
# 2️⃣ Process every book, collecting quality metrics
# ----------------------------------------------------------------------
batch_report = []
for book_file in BOOK_ROOT.rglob("*.[tT][xX][tT]"):
try:
report = run_pipeline(book_file)
batch_report.append(report)
print(f"[✅] {book_file.name} processed – {report['skill_count']} skills")
except Exception as e:
print(f"[❌] {book_file.name} error: {e}")
# ----------------------------------------------------------------------
# 3️⃣ Summarise batch‑level quality control
# ----------------------------------------------------------------------
summary_path = OUTPUT_ROOT / "batch_quality_summary.json"
with summary_path.open("w", encoding="utf-8") as f:
json.dump(batch_report, f, indent=2, ensure_ascii=False)
print(f"\nBatch quality summary written to {summary_path}")
How the Script Ensures Quality
| Step | Function | Quality Impact |
|---|---|---|
| Locate source files | Recursively finds every .txt or .md file |
Guarantees no book is skipped in the batch |
| Run pipeline | Calls the CLI with JSON reporting | Internally executes all seven stages, including triple verification and pressure tests |
| Collect reports | Captures skill counts and failure states | Flags books producing few skills (possible quality issues) or pipeline errors |
| Summarize | Writes batch_quality_summary.json |
Provides a single reviewable artifact for the entire batch before publication |
Integrate this script into a CI workflow (GitHub Actions, GitLab CI) to automatically trigger batch runs and fail builds when books do not meet quality thresholds.
Key Files and Templates
The pipeline relies on specific files in the kangarooking/cangjie-skill repository:
SKILL.md– Master execution specification defining the pipeline logicmethodology/02-stage1-parallel-extract.md– Describes the five parallel extractors and their invocation patternsmethodology/03-stage1.5-triple-verify.md– Details the three-fold verification criteria (citation, novelty, uniqueness)methodology/06-stage4-pressure-test.md– Defines the automated bait-question testing protocoltemplates/BOOK_OVERVIEW.md.template– Skeleton for the high-level overview generated per booktemplates/INDEX.md.template– Template for the skill map (INDEX.md)templates/DIGEST.md.template– Template for the reader-friendly digest (DIGEST.md)extractors/*-extractor.md– Prompt definitions for each specialized extractor
Summary
- Batch processing in Cangjie-Skill repeats the RIA-TV++ pipeline for each source book, ensuring consistent output structure.
- Quality control is automated through triple verification (stage 3), which filters candidates requiring dual citations and novel insights, and pressure testing (stage 6), which validates skills against bait questions.
- Parallel extraction (stage 2) accelerates throughput by running five extractors simultaneously per book.
- Implementation requires only a simple loop invoking the CLI; the repository's methodology documents in
methodology/define the quality gates, whiletemplates/provide the output structure. - Monitoring batch health relies on parsing JSON reports for skill counts and pipeline failures, enabling systematic review before deployment.
Frequently Asked Questions
What is the RIA-TV++ pipeline in Cangjie-Skill?
The RIA-TV++ pipeline is a seven-stage workflow that converts raw text into verified AI Skills. It stands for Reading, Interpretation, Application, plus Triple verification, Visual linking, and ++ (pressure testing and delivery). As implemented in kangarooking/cangjie-skill, it ensures every processed book undergoes the same rigorous extraction, verification, and testing sequence.
How does triple verification filter low-quality extractions?
Triple verification requires every candidate skill to meet three criteria documented in methodology/03-stage1.5-triple-verify.md: at least two independent citations from the source text, the ability to answer a new predictive question, and non-triviality (excluding common knowledge). This typically reduces raw extractor output by 50–75%, ensuring only substantive, well-supported concepts become skills.
What happens when a book fails the pressure test?
If a skill fails the automated bait-question testing in stage 6 (defined in methodology/06-stage4-pressure-test.md), the pipeline automatically triggers a re-run of earlier stages—typically returning to parallel extraction or triple verification—to regenerate the problematic units. This feedback loop prevents broken or inconsistent skills from reaching the final DIGEST.md output.
Can I customize extractors for specific domains or book types?
Yes. The extractor prompts in extractors/ (such as framework-extractor.md and principle-extractor.md) are modular templates. You can modify these Markdown files to adjust extraction logic for technical manuals, fiction, or academic papers without altering the core pipeline architecture, allowing domain-specific batch processing while maintaining the standard quality control gates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →