How to Audit and Review Rejected Candidates in the Cangjie-Skill Pipeline
Rejected candidates are stored as individual markdown files in books/<slug>/rejected/<id>.md with YAML front matter documenting which of the three verification checks failed and why, enabling systematic human review.
The Cangjie-Skill project converts high-value books, videos, and podcasts into reusable AI Skills through a structured extraction pipeline. During Stage 1.5 – Triple Verification, candidates undergo three independent quality checks, and those that fail are written to dedicated rejection folders with full audit trails. Learning how to audit and review rejected candidates ensures transparent quality control and allows teams to rescue valuable content that may have failed due to overly strict heuristics.
Understanding the Triple Verification Gate
The pipeline’s quality control centers on three verification checks defined in methodology/03-stage1.5-triple-verify.md:
- V1 Cross-Domain: Validates whether the concept applies beyond its immediate context
- V2 Predictive Power: Determines if the concept explains outcomes across different scenarios
- V3 Exclusivity: Ensures the concept is non-trivial and not already covered by existing skills
According to the methodology documentation, a candidate must pass all three checks to proceed to Stage 2. If any check fails, the candidate is immediately routed to the rejection workflow documented in methodology/00-overview.md, which shows the "through + rejected" branching logic.
Locating Rejected Candidate Files
Rejected candidates follow a strict file hierarchy documented in SKILL.md. Each rejection creates a standalone markdown file at:
books/<slug>/rejected/<id>.md
The <slug> represents the source content identifier, while <id> is the unique candidate identifier. These files remain in the repository permanently, creating an immutable audit log that reviewers can revisit to understand pipeline decisions or retrieve incorrectly filtered content.
Parsing the Rejection Metadata
Each rejected file contains YAML front matter that records the failure details. The structure includes:
id: Internal candidate identifiertitle: Human-readable candidate nametype: Classification (e.g.,framework,principle)V1_cross_domain,V2_predictive_power,V3_exclusivity: Objects containingpassed: falseand areasonstring
This metadata design allows automated tools to parse exactly which verification stage rejected the candidate and why, as implemented in the verification logic shown in the methodology files.
Running the Audit Review Process
Step 1 – Collect Rejected Files
Begin the audit by gathering all rejection records from the repository. The system writes every rejection to the rejected/ subdirectory within the corresponding book folder, making batch collection straightforward using standard file system tools or path globbing.
Step 2 – Parse YAML Headers
Extract the front matter from each markdown file to identify the failure mode. The YAML block appears at the start of each file, delimited by --- markers, containing the verification results and human-readable explanations for each failure.
Step 3 – Analyze Failure Reasons
Human reviewers examine the reason fields in the YAML to determine if the rejection is justified. Common scenarios include candidates that narrowly missed the predictive power threshold or cross-domain applicability checks but contain valuable underlying concepts suitable for rescue.
Step 4 – Confirm or Rescue Candidates
Before proceeding to Stage 2, the system presents a concise summary showing passed N and rejected M counts to the user for final "light-confirmation". During this step, reviewers can flag recoverable candidates by modifying the rejection file or moving it to the verification queue, triggering reprocessing through the expensive Stage 2-4 steps only for confirmed high-quality content.
Automating the Audit with Python
The following script demonstrates how to programmatically audit rejected candidates by parsing the YAML metadata and generating a summary report:
import pathlib
import yaml
import json
# 1️⃣ Locate all rejected markdown files
REJECTED_ROOT = pathlib.Path("books")
rejected_files = list(REJECTED_ROOT.glob("*/rejected/*.md"))
def load_yaml(md_path: pathlib.Path) -> dict:
"""Extract the leading YAML block from a markdown file."""
with md_path.open(encoding="utf-8") as f:
lines = []
for line in f:
if line.strip() == "---" and lines:
break # stop after second delimiter
lines.append(line)
return yaml.safe_load("\n".join(lines))
# 2️⃣ Build a summary table
summary = []
for file in rejected_files:
data = load_yaml(file)
# Which verification failed?
failed = [v for v, info in data.items()
if v.startswith("V") and not info.get("passed", True)]
summary.append({
"id": data.get("id"),
"title": data.get("title"),
"type": data.get("type"),
"failed_checks": failed,
"reasons": {v: data[v].get("reason") for v in failed}
})
# 3️⃣ Pretty‑print the audit report (JSON for easy downstream consumption)
print(json.dumps(summary, indent=2, ensure_ascii=False))
This automation walks the books/*/rejected/ hierarchy, extracts the verification status from the YAML front matter, and outputs a structured JSON array suitable for import into review tools, spreadsheets, or custom dashboards.
Why Audit Trails Matter
Maintaining detailed rejection records serves three critical functions in the Cangjie-Skill pipeline:
- Quality Assurance: Only candidates surviving all three verification checks become independent skills, ensuring the final set is both non-trivial and genuinely explanatory
- Review Transparency: Storing rejection rationales enables downstream reviewers to understand exclusion decisions and retrieve candidates if evaluation criteria evolve
- Pipeline Improvement: Analyzing failure patterns across rejected files highlights gaps in extraction prompts or verification heuristics, driving continuous refinement of the Stage 1.5 logic
Summary
- Rejected candidates are written to
books/<slug>/rejected/<id>.mdwith full YAML metadata explaining the failure - Each rejection file documents which of the three verification checks (V1 Cross-Domain, V2 Predictive Power, V3 Exclusivity) failed and why
- The audit process involves collecting these files, parsing their YAML front matter, and reviewing failure reasons for potential rescue
- Python scripts can automate the extraction of rejection data for batch review and reporting
- All rejections remain in the repository as permanent audit logs, supporting iterative pipeline improvement and quality verification
Frequently Asked Questions
What causes a candidate to be rejected in the Cangjie-Skill pipeline?
A candidate is rejected during Stage 1.5 Triple Verification if it fails any of the three required checks: V1 Cross-Domain applicability, V2 Predictive Power, or V3 Exclusivity. According to methodology/03-stage1.5-triple-verify.md, the verification process evaluates whether concepts are generalizable, explanatory across scenarios, and non-redundant with existing skills.
How is rejection data structured within the markdown files?
Each rejected candidate follows a standard markdown format with YAML front matter containing id, title, and type fields, plus three verification objects (V1_cross_domain, V2_predictive_power, V3_exclusivity). Each verification object includes a boolean passed field and a reason string explaining the failure, as documented in the SKILL.md file layout specification.
Can rejected candidates be recovered and reprocessed?
Yes, rejected candidates can be rescued during the user confirmation step before proceeding to Stage 2. Reviewers examine the reason fields in the rejection files and may update the candidate's status or move the file to trigger reprocessing through the expensive Stage 2-4 pipeline steps, ensuring valuable content isn't permanently lost due to initial strict filtering.
Where are the verification criteria defined?
The three verification checks are defined in methodology/03-stage1.5-triple-verify.md, while the overall workflow including the rejection branch is shown in methodology/00-overview.md. The repository layout and audit requirements are documented in SKILL.md, specifically noting the rejected/ directory structure and its role in maintaining quality gates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →