# How to Audit and Review Rejected Candidates in the Cangjie-Skill Pipeline

> Learn to audit and review rejected candidates in the Cangjie-Skill pipeline. Examine rejection reasons in markdown files for systematic human review and process improvement.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: how-to-guide
- Published: 2026-07-22

---

**Rejected candidates are stored as individual markdown files in `books/<slug>/rejected/<id>.md` with YAML front matter documenting which of the three verification checks failed and why, enabling systematic human review.**

The Cangjie-Skill project converts high-value books, videos, and podcasts into reusable AI Skills through a structured extraction pipeline. During **Stage 1.5 – Triple Verification**, candidates undergo three independent quality checks, and those that fail are written to dedicated rejection folders with full audit trails. Learning how to audit and review rejected candidates ensures transparent quality control and allows teams to rescue valuable content that may have failed due to overly strict heuristics.

## Understanding the Triple Verification Gate

The pipeline’s quality control centers on three verification checks defined in [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md):

- **V1 Cross-Domain**: Validates whether the concept applies beyond its immediate context
- **V2 Predictive Power**: Determines if the concept explains outcomes across different scenarios
- **V3 Exclusivity**: Ensures the concept is non-trivial and not already covered by existing skills

According to the methodology documentation, a candidate must pass all three checks to proceed to Stage 2. If any check fails, the candidate is immediately routed to the rejection workflow documented in [`methodology/00-overview.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md), which shows the "through + rejected" branching logic.

## Locating Rejected Candidate Files

Rejected candidates follow a strict file hierarchy documented in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md). Each rejection creates a standalone markdown file at:

```text
books/<slug>/rejected/<id>.md

```

The `<slug>` represents the source content identifier, while `<id>` is the unique candidate identifier. These files remain in the repository permanently, creating an immutable audit log that reviewers can revisit to understand pipeline decisions or retrieve incorrectly filtered content.

## Parsing the Rejection Metadata

Each rejected file contains YAML front matter that records the failure details. The structure includes:

- `id`: Internal candidate identifier
- `title`: Human-readable candidate name
- `type`: Classification (e.g., `framework`, `principle`)
- `V1_cross_domain`, `V2_predictive_power`, `V3_exclusivity`: Objects containing `passed: false` and a `reason` string

This metadata design allows automated tools to parse exactly which verification stage rejected the candidate and why, as implemented in the verification logic shown in the methodology files.

## Running the Audit Review Process

### Step 1 – Collect Rejected Files

Begin the audit by gathering all rejection records from the repository. The system writes every rejection to the `rejected/` subdirectory within the corresponding book folder, making batch collection straightforward using standard file system tools or path globbing.

### Step 2 – Parse YAML Headers

Extract the front matter from each markdown file to identify the failure mode. The YAML block appears at the start of each file, delimited by `---` markers, containing the verification results and human-readable explanations for each failure.

### Step 3 – Analyze Failure Reasons

Human reviewers examine the `reason` fields in the YAML to determine if the rejection is justified. Common scenarios include candidates that narrowly missed the predictive power threshold or cross-domain applicability checks but contain valuable underlying concepts suitable for rescue.

### Step 4 – Confirm or Rescue Candidates

Before proceeding to Stage 2, the system presents a concise summary showing **passed N** and **rejected M** counts to the user for final "light-confirmation". During this step, reviewers can flag recoverable candidates by modifying the rejection file or moving it to the verification queue, triggering reprocessing through the expensive Stage 2-4 steps only for confirmed high-quality content.

## Automating the Audit with Python

The following script demonstrates how to programmatically audit rejected candidates by parsing the YAML metadata and generating a summary report:

```python
import pathlib
import yaml
import json

# 1️⃣ Locate all rejected markdown files

REJECTED_ROOT = pathlib.Path("books")
rejected_files = list(REJECTED_ROOT.glob("*/rejected/*.md"))

def load_yaml(md_path: pathlib.Path) -> dict:
    """Extract the leading YAML block from a markdown file."""
    with md_path.open(encoding="utf-8") as f:
        lines = []
        for line in f:
            if line.strip() == "---" and lines:
                break          # stop after second delimiter

            lines.append(line)
        return yaml.safe_load("\n".join(lines))

# 2️⃣ Build a summary table

summary = []
for file in rejected_files:
    data = load_yaml(file)
    # Which verification failed?

    failed = [v for v, info in data.items()
              if v.startswith("V") and not info.get("passed", True)]
    summary.append({
        "id": data.get("id"),
        "title": data.get("title"),
        "type": data.get("type"),
        "failed_checks": failed,
        "reasons": {v: data[v].get("reason") for v in failed}
    })

# 3️⃣ Pretty‑print the audit report (JSON for easy downstream consumption)

print(json.dumps(summary, indent=2, ensure_ascii=False))

```

This automation walks the `books/*/rejected/` hierarchy, extracts the verification status from the YAML front matter, and outputs a structured JSON array suitable for import into review tools, spreadsheets, or custom dashboards.

## Why Audit Trails Matter

Maintaining detailed rejection records serves three critical functions in the Cangjie-Skill pipeline:

- **Quality Assurance**: Only candidates surviving all three verification checks become independent skills, ensuring the final set is both non-trivial and genuinely explanatory
- **Review Transparency**: Storing rejection rationales enables downstream reviewers to understand exclusion decisions and retrieve candidates if evaluation criteria evolve
- **Pipeline Improvement**: Analyzing failure patterns across rejected files highlights gaps in extraction prompts or verification heuristics, driving continuous refinement of the Stage 1.5 logic

## Summary

- Rejected candidates are written to `books/<slug>/rejected/<id>.md` with full YAML metadata explaining the failure
- Each rejection file documents which of the three verification checks (V1 Cross-Domain, V2 Predictive Power, V3 Exclusivity) failed and why
- The audit process involves collecting these files, parsing their YAML front matter, and reviewing failure reasons for potential rescue
- Python scripts can automate the extraction of rejection data for batch review and reporting
- All rejections remain in the repository as permanent audit logs, supporting iterative pipeline improvement and quality verification

## Frequently Asked Questions

### What causes a candidate to be rejected in the Cangjie-Skill pipeline?

A candidate is rejected during Stage 1.5 Triple Verification if it fails any of the three required checks: V1 Cross-Domain applicability, V2 Predictive Power, or V3 Exclusivity. According to [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md), the verification process evaluates whether concepts are generalizable, explanatory across scenarios, and non-redundant with existing skills.

### How is rejection data structured within the markdown files?

Each rejected candidate follows a standard markdown format with YAML front matter containing `id`, `title`, and `type` fields, plus three verification objects (`V1_cross_domain`, `V2_predictive_power`, `V3_exclusivity`). Each verification object includes a boolean `passed` field and a `reason` string explaining the failure, as documented in the [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) file layout specification.

### Can rejected candidates be recovered and reprocessed?

Yes, rejected candidates can be rescued during the user confirmation step before proceeding to Stage 2. Reviewers examine the `reason` fields in the rejection files and may update the candidate's status or move the file to trigger reprocessing through the expensive Stage 2-4 pipeline steps, ensuring valuable content isn't permanently lost due to initial strict filtering.

### Where are the verification criteria defined?

The three verification checks are defined in [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md), while the overall workflow including the rejection branch is shown in [`methodology/00-overview.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md). The repository layout and audit requirements are documented in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md), specifically noting the `rejected/` directory structure and its role in maintaining quality gates.