How to Audit and Recover Rejected Candidates from the `rejected/` Directory

The rejected/ directory in the kangarooking/cangjie-skill repository stores every candidate that fails the triple-verify stage as individual markdown files, enabling you to audit failure reasons and recover valuable items by moving them back to the candidate pool or directly to the verified list.

The kangarooking/cangjie-skill repository implements a rigorous triple-verify pipeline (Stage 1.5) that filters skill candidates through cross-domain, predictive power, and exclusivity checks. When a candidate fails verification, it is not deleted but instead archived in books/<book-slug>/rejected/<candidate-id>.md, creating a permanent audit trail. This design allows maintainers to audit and recover rejected candidates long after the initial verification process, ensuring no valuable content is lost due to temporary rule misconfigurations or evolving quality standards.

Understanding the Rejected/ Directory Architecture

The rejected/ folder serves as a built-in audit trail that captures the complete context of every rejection. According to the repository's SKILL.md, this directory holds "淘汰的单元 + 原因 (审计用)"—eliminated units plus reasons for auditing purposes.

The Triple-Verify Pipeline Flow

  1. Candidate Pool: All raw extracts are first placed in candidates/ before verification begins.
  2. Triple-Verify Stage: Each candidate undergoes three checks:
    • V1: Cross-domain validation
    • V2: Predictive power assessment
    • V3: Exclusivity verification
  3. Outcome Handling:
    • Pass: Written to books/<slug>/verified.md
    • Fail: Written to books/<slug>/rejected/<id>.md with structured failure metadata

As defined in methodology/03-stage1.5-triple-verify.md, the system deliberately preserves which verification step failed and why, making the rejection reason machine-readable and auditable.

How to Audit Rejected Candidates

Auditing requires inspecting the markdown files in the rejected directory to understand why specific candidates failed verification. Each file contains YAML front-matter with the candidate title and a V*_* block indicating the specific failure.

Step 1: Locate and List Rejected Items

Navigate to the rejected folder for your target book and list all candidates with their failure reasons:

cd books/<book-slug>/rejected

for f in *.md; do
  echo "=== $f ==="
  grep -E "title:|passed: false|reason:" "$f"
  echo
done

This command extracts the title and failure metadata from each rejected candidate's front-matter, giving you a quick overview of what was eliminated and why.

Step 2: Inspect Detailed Failure Reasons

To examine the full context of a specific rejection, view the verification block that records which stage failed:

sed -n '/V1_/,${/V1_/p;q}' <candidate-id>.md

Alternatively, open the file in your editor. The rejection record includes the specific verification step (V1, V2, or V3) that returned passed: false and the accompanying reason string.

Step 3: Generate Audit Reports

For systematic analysis, generate a summary report of all rejections:

> audit-report.md && for f in *.md; do 
  echo "## $f" >> audit-report.md

  grep -E "title:|V[123]_|passed:|reason:" "$f" >> audit-report.md
  echo "" >> audit-report.md
done

This creates audit-report.md containing a structured summary of every rejected candidate and its failure mode, useful for identifying patterns in the verification process.

How to Recover Rejected Candidates

The kangarooking/cangjie-skill workflow supports three recovery paths depending on your confidence level in the candidate's validity. All operations are performed as simple file manipulations since the rejected files are never deleted automatically.

Recover to Candidates Pool for Re-verification

Use this method when you want the candidate to undergo the full triple-verify process again—ideal after updating verification rules or fixing data quality issues in the source:

mv books/<book-slug>/rejected/<candidate-id>.md \
   books/<book-slug>/candidates/<candidate-id>.md

After moving the file, re-run the triple-verify script on the candidates directory to reprocess the item through V1, V2, and V3 checks.

Direct Recovery to Verified List

When manual review confirms the candidate is valid and should bypass re-verification, append it directly to the verified catalogue:

REJ="books/<book-slug>/rejected/<candidate-id>.md"
VERIFIED="books/<book-slug>/verified.md"

# Append content (skipping the front-matter separator lines)

sed -n '/---/,${/---/d;p}' "$REJ" >> "$VERIFIED"

# Remove the rejected copy

rm "$REJ"

Note that this merges the content into verified.md; ensure the YAML front-matter from the rejected file is properly handled if your verified catalogue requires specific metadata formatting.

Batch Recovery Operations

For recovering multiple candidates after a bulk rule review:

REJ_DIR="books/<book-slug>/rejected"
RECOVER_DIR="books/<book-slug>/candidates/recovered"

mkdir -p "$RECOVER_DIR"
mv "$REJ_DIR"/*.md "$RECOVER_DIR"/

# Optional: trigger re-verification on the recovered batch

# ./scripts/triple_verify.sh "$RECOVER_DIR"

After any recovery operation, regenerate the book index by updating SKILL.md or processing templates/INDEX.md.template to ensure recovered skills appear in the public catalogue.

Why the Audit Trail Matters

The rejected/ directory provides three critical functions for the cangjie-skill workflow:

  • Transparency: Every rejection is justified with specific verification step failures, making it easy for reviewers to understand past decisions without guessing why content was excluded.
  • Safety Net: Verification mistakes or overly aggressive thresholds can be corrected without re-extracting content from original sources—simply recover the rejected file.
  • Metrics: By counting files in rejected/ versus verified.md, you can monitor pass-rate trends and adjust the "数量预期" (quantity expectations) defined in Stage 1.5 of the methodology.

Summary

  • The rejected/ directory stores failed candidates as individual markdown files at books/<slug>/rejected/<id>.md with structured failure reasons from the triple-verify stage.
  • Audit by listing files with grep to extract titles and failure reasons, or generate comprehensive reports using shell loops.
  • Recover individual candidates by moving them back to candidates/ for re-verification, or append directly to verified.md for immediate publication.
  • Batch recover using standard mv commands to transfer multiple rejected files to a recovery subdirectory.
  • Always regenerate the book index after recovery to update SKILL.md and public-facing catalogues.

Frequently Asked Questions

Where exactly are rejected candidates stored in the repository?

Rejected candidates are stored in books/<book-slug>/rejected/<candidate-id>.md according to the architecture defined in SKILL.md and methodology/03-stage1.5-triple-verify.md. Each rejection is saved as a separate markdown file containing the candidate's content and a structured record of which verification step (V1, V2, or V3) failed.

Can I recover a candidate that failed the exclusivity check (V3)?

Yes. Any rejected candidate can be recovered regardless of which verification stage failed. Move the file from books/<slug>/rejected/ back to books/<slug>/candidates/ to subject it to re-verification, or manually append it to verified.md if you have determined the exclusivity conflict was erroneous or acceptable.

How do I know why a specific candidate was rejected?

Each rejected markdown file contains a verification block indicating the specific stage failure. Use grep -E "V[123]_|reason:" books/<slug>/rejected/<id>.md to quickly identify which check (cross-domain, predictive power, or exclusivity) returned passed: false and the accompanying explanation.

Will recovered candidates automatically appear in the skill index?

No. After moving files from rejected/ to either candidates/ or verified.md, you must regenerate the book index by processing templates/INDEX.md.template or updating SKILL.md to ensure the recovered skills appear in the public-facing catalogue and audit trail sections.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →