Difference Between candidates/ and rejected/ Directories in the Cangjie-Skill Audit Trail

The candidates/ directory contains raw knowledge units produced by stage 1 parallel extraction, while rejected/ holds units that failed stage 1.5 triple-verification checks with documented failure reasons.

The cangjie-skill repository implements a rigorous knowledge extraction pipeline that maintains complete auditability through structured directories. Understanding the distinction between these two folders is essential for debugging extraction quality and ensuring compliance with the methodology defined in the project source code.

What the candidates/ Directory Contains

The candidates/ folder is populated immediately after stage 1 – parallel extraction completes. According to SKILL.md lines 60-62, this directory serves as the immutable record of everything the five sub-agents initially produce.

Each extractor outputs specific markdown files:

These files represent the unfiltered, raw extraction results before any quality verification occurs. The directory provides a "show-your-work" audit trail that enables reproducibility and debugging of the extraction process, as detailed in methodology/02-stage1-parallel-extract.md.

What the rejected/ Directory Contains

The rejected/ directory is created during stage 1.5 – triple-verification filtering, as implemented in methodology/03-stage1.5-triple-verify.md lines 45-50. This folder archives candidates that fail any of the three verification checks:

  1. V1 (Cross-domain) – Fails to generalize beyond the source material
  2. V2 (Predictive power) – Lacks novel use cases or actionable insights
  3. V3 (Exclusivity) – Represents common wisdom rather than specialized knowledge

Unlike the aggregated markdown files in candidates/, the rejected/ directory contains individual files named <id>.md (e.g., f01.md, f07.md). Each file includes YAML frontmatter documenting which check failed and why:

failed_check: V2_predictive_power
reason: no novel use case

This structure preserves the audit trail of discard decisions, allowing later review or potential "rescue" of incorrectly filtered items.

Key Differences at a Glance

Aspect candidates/ rejected/
Creation Timing After stage 1 extraction After stage 1.5 verification
Content State Raw, unfiltered outputs Failed verification units
File Structure Five aggregated markdown files per book Individual markdown files per rejected unit
Metadata None (raw content) YAML frontmatter with failed_check and reason
Primary Purpose Reproducibility and debugging Compliance and quality audit trail

How to Access and Parse Audit Files

Downstream tools can programmatically analyze both directories to generate quality reports or re-evaluate extraction decisions. The templates/INDEX.md.template file generates the final audit view linking to both directories.

Here is a Python snippet demonstrating how to read these audit trails:

import pathlib
import yaml

# Path to the book slug (replace <slug> with actual folder name)

book_root = pathlib.Path("books/<slug>")

# 1️⃣ List all raw candidates

candidates_dir = book_root / "candidates"
for md_file in candidates_dir.glob("*.md"):
    print("Candidate:", md_file.name)

# 2️⃣ List all rejected items with reasons

rejected_dir = book_root / "rejected"
for md_file in rejected_dir.glob("*.md"):
    data = yaml.safe_load(md_file.read_text())
    print(f"Rejected {md_file.name}: failed {data.get('failed_check')} – {data.get('reason')}")

Typical output demonstrates the lifecycle distinction:


Candidate: frameworks.md
Candidate: principles.md
...
Rejected f01.md: failed V2_predictive_power – no novel use case
Rejected f07.md: failed V3_exclusivity – common wisdom

Summary

  • candidates/ stores the complete, unfiltered output from stage 1 parallel extraction across five specialized sub-agents, providing reproducibility and debugging capabilities.
  • rejected/ archives individual units that failed the three verification checks (cross-domain, predictive power, exclusivity) with explicit YAML-documented reasons.
  • Both directories are essential for the audit trail architecture defined in SKILL.md and processed through templates in templates/INDEX.md.template.
  • Programmatic access requires standard markdown parsing for candidates/ and YAML frontmatter extraction for rejected/.

Frequently Asked Questions

When are the candidates/ and rejected/ directories created?

The candidates/ directory is created immediately after stage 1 completes, containing raw outputs from the five parallel extractors. The rejected/ directory is populated during stage 1.5 after the triple-verification filtering process identifies units that fail quality checks.

What specific information is stored in rejected/ files?

Each file in rejected/ contains YAML frontmatter with two key fields: failed_check (indicating which of the three verification checks failed: V1, V2, or V3) and reason (providing a human-readable explanation of why the unit was discarded). This metadata enables precise audit trails and potential re-evaluation.

Can a rejected candidate be recovered or reinstated?

Yes. Because the rejected/ directory preserves the original content alongside the failure metadata, downstream processes can implement "rescue" logic to re-evaluate discarded items. The explicit reason field allows developers to identify false positives or adjust verification thresholds without re-running the entire extraction pipeline.

How does the audit trail structure ensure reproducibility?

The separation of candidates/ (raw extraction) and rejected/ (filtered results) creates an immutable record of the pipeline's decision-making process. By maintaining these directories as specified in SKILL.md and methodology/03-stage1.5-triple-verify.md, the system enables complete reconstruction of how final knowledge units were selected, satisfying compliance requirements for "show-your-work" auditing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →