How to Audit and Recover Rejected Candidates from the `rejected/` Directory
The rejected/ directory in the kangarooking/cangjie-skill repository stores every candidate that fails the triple-verify stage as individual markdown files, enabling you to audit failure reasons and recover valuable items by moving them back to the candidate pool or directly to the verified list.
The kangarooking/cangjie-skill repository implements a rigorous triple-verify pipeline (Stage 1.5) that filters skill candidates through cross-domain, predictive power, and exclusivity checks. When a candidate fails verification, it is not deleted but instead archived in books/<book-slug>/rejected/<candidate-id>.md, creating a permanent audit trail. This design allows maintainers to audit and recover rejected candidates long after the initial verification process, ensuring no valuable content is lost due to temporary rule misconfigurations or evolving quality standards.
Understanding the Rejected/ Directory Architecture
The rejected/ folder serves as a built-in audit trail that captures the complete context of every rejection. According to the repository's SKILL.md, this directory holds "淘汰的单元 + 原因 (审计用)"—eliminated units plus reasons for auditing purposes.
The Triple-Verify Pipeline Flow
- Candidate Pool: All raw extracts are first placed in
candidates/before verification begins. - Triple-Verify Stage: Each candidate undergoes three checks:
- V1: Cross-domain validation
- V2: Predictive power assessment
- V3: Exclusivity verification
- Outcome Handling:
- Pass: Written to
books/<slug>/verified.md - Fail: Written to
books/<slug>/rejected/<id>.mdwith structured failure metadata
- Pass: Written to
As defined in methodology/03-stage1.5-triple-verify.md, the system deliberately preserves which verification step failed and why, making the rejection reason machine-readable and auditable.
How to Audit Rejected Candidates
Auditing requires inspecting the markdown files in the rejected directory to understand why specific candidates failed verification. Each file contains YAML front-matter with the candidate title and a V*_* block indicating the specific failure.
Step 1: Locate and List Rejected Items
Navigate to the rejected folder for your target book and list all candidates with their failure reasons:
cd books/<book-slug>/rejected
for f in *.md; do
echo "=== $f ==="
grep -E "title:|passed: false|reason:" "$f"
echo
done
This command extracts the title and failure metadata from each rejected candidate's front-matter, giving you a quick overview of what was eliminated and why.
Step 2: Inspect Detailed Failure Reasons
To examine the full context of a specific rejection, view the verification block that records which stage failed:
sed -n '/V1_/,${/V1_/p;q}' <candidate-id>.md
Alternatively, open the file in your editor. The rejection record includes the specific verification step (V1, V2, or V3) that returned passed: false and the accompanying reason string.
Step 3: Generate Audit Reports
For systematic analysis, generate a summary report of all rejections:
> audit-report.md && for f in *.md; do
echo "## $f" >> audit-report.md
grep -E "title:|V[123]_|passed:|reason:" "$f" >> audit-report.md
echo "" >> audit-report.md
done
This creates audit-report.md containing a structured summary of every rejected candidate and its failure mode, useful for identifying patterns in the verification process.
How to Recover Rejected Candidates
The kangarooking/cangjie-skill workflow supports three recovery paths depending on your confidence level in the candidate's validity. All operations are performed as simple file manipulations since the rejected files are never deleted automatically.
Recover to Candidates Pool for Re-verification
Use this method when you want the candidate to undergo the full triple-verify process again—ideal after updating verification rules or fixing data quality issues in the source:
mv books/<book-slug>/rejected/<candidate-id>.md \
books/<book-slug>/candidates/<candidate-id>.md
After moving the file, re-run the triple-verify script on the candidates directory to reprocess the item through V1, V2, and V3 checks.
Direct Recovery to Verified List
When manual review confirms the candidate is valid and should bypass re-verification, append it directly to the verified catalogue:
REJ="books/<book-slug>/rejected/<candidate-id>.md"
VERIFIED="books/<book-slug>/verified.md"
# Append content (skipping the front-matter separator lines)
sed -n '/---/,${/---/d;p}' "$REJ" >> "$VERIFIED"
# Remove the rejected copy
rm "$REJ"
Note that this merges the content into verified.md; ensure the YAML front-matter from the rejected file is properly handled if your verified catalogue requires specific metadata formatting.
Batch Recovery Operations
For recovering multiple candidates after a bulk rule review:
REJ_DIR="books/<book-slug>/rejected"
RECOVER_DIR="books/<book-slug>/candidates/recovered"
mkdir -p "$RECOVER_DIR"
mv "$REJ_DIR"/*.md "$RECOVER_DIR"/
# Optional: trigger re-verification on the recovered batch
# ./scripts/triple_verify.sh "$RECOVER_DIR"
After any recovery operation, regenerate the book index by updating SKILL.md or processing templates/INDEX.md.template to ensure recovered skills appear in the public catalogue.
Why the Audit Trail Matters
The rejected/ directory provides three critical functions for the cangjie-skill workflow:
- Transparency: Every rejection is justified with specific verification step failures, making it easy for reviewers to understand past decisions without guessing why content was excluded.
- Safety Net: Verification mistakes or overly aggressive thresholds can be corrected without re-extracting content from original sources—simply recover the rejected file.
- Metrics: By counting files in
rejected/versusverified.md, you can monitor pass-rate trends and adjust the "数量预期" (quantity expectations) defined in Stage 1.5 of the methodology.
Summary
- The
rejected/directory stores failed candidates as individual markdown files atbooks/<slug>/rejected/<id>.mdwith structured failure reasons from the triple-verify stage. - Audit by listing files with
grepto extract titles and failure reasons, or generate comprehensive reports using shell loops. - Recover individual candidates by moving them back to
candidates/for re-verification, or append directly toverified.mdfor immediate publication. - Batch recover using standard
mvcommands to transfer multiple rejected files to a recovery subdirectory. - Always regenerate the book index after recovery to update
SKILL.mdand public-facing catalogues.
Frequently Asked Questions
Where exactly are rejected candidates stored in the repository?
Rejected candidates are stored in books/<book-slug>/rejected/<candidate-id>.md according to the architecture defined in SKILL.md and methodology/03-stage1.5-triple-verify.md. Each rejection is saved as a separate markdown file containing the candidate's content and a structured record of which verification step (V1, V2, or V3) failed.
Can I recover a candidate that failed the exclusivity check (V3)?
Yes. Any rejected candidate can be recovered regardless of which verification stage failed. Move the file from books/<slug>/rejected/ back to books/<slug>/candidates/ to subject it to re-verification, or manually append it to verified.md if you have determined the exclusivity conflict was erroneous or acceptable.
How do I know why a specific candidate was rejected?
Each rejected markdown file contains a verification block indicating the specific stage failure. Use grep -E "V[123]_|reason:" books/<slug>/rejected/<id>.md to quickly identify which check (cross-domain, predictive power, or exclusivity) returned passed: false and the accompanying explanation.
Will recovered candidates automatically appear in the skill index?
No. After moving files from rejected/ to either candidates/ or verified.md, you must regenerate the book index by processing templates/INDEX.md.template or updating SKILL.md to ensure the recovered skills appear in the public-facing catalogue and audit trail sections.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →