# Difference Between candidates/ and rejected/ Directories in the Cangjie-Skill Audit Trail

> Understand the difference between candidates/ and rejected/ directories in the Cangjie-Skill audit trail. Learn what knowledge units are extracted and why some are rejected.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: documentation
- Published: 2026-07-21

---

**The `candidates/` directory contains raw knowledge units produced by stage 1 parallel extraction, while `rejected/` holds units that failed stage 1.5 triple-verification checks with documented failure reasons.**

The cangjie-skill repository implements a rigorous knowledge extraction pipeline that maintains complete auditability through structured directories. Understanding the distinction between these two folders is essential for debugging extraction quality and ensuring compliance with the methodology defined in the project source code.

## What the candidates/ Directory Contains

The `candidates/` folder is populated immediately after **stage 1 – parallel extraction** completes. According to [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) lines 60-62, this directory serves as the immutable record of everything the five sub-agents initially produce.

Each extractor outputs specific markdown files:

- [`frameworks.md`](https://github.com/kangarooking/cangjie-skill/blob/main/frameworks.md) – Architectural patterns and mental models
- [`principles.md`](https://github.com/kangarooking/cangjie-skill/blob/main/principles.md) – Core rules and heuristics
- [`cases.md`](https://github.com/kangarooking/cangjie-skill/blob/main/cases.md) – Concrete implementation examples
- [`counter-examples.md`](https://github.com/kangarooking/cangjie-skill/blob/main/counter-examples.md) – Anti-patterns and pitfalls
- [`glossary.md`](https://github.com/kangarooking/cangjie-skill/blob/main/glossary.md) – Domain terminology definitions

These files represent the **unfiltered, raw extraction results** before any quality verification occurs. The directory provides a "show-your-work" audit trail that enables reproducibility and debugging of the extraction process, as detailed in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md).

## What the rejected/ Directory Contains

The `rejected/` directory is created during **stage 1.5 – triple-verification filtering**, as implemented in [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md) lines 45-50. This folder archives candidates that fail any of the three verification checks:

1. **V1 (Cross-domain)** – Fails to generalize beyond the source material
2. **V2 (Predictive power)** – Lacks novel use cases or actionable insights
3. **V3 (Exclusivity)** – Represents common wisdom rather than specialized knowledge

Unlike the aggregated markdown files in `candidates/`, the `rejected/` directory contains individual files named `<id>.md` (e.g., [`f01.md`](https://github.com/kangarooking/cangjie-skill/blob/main/f01.md), [`f07.md`](https://github.com/kangarooking/cangjie-skill/blob/main/f07.md)). Each file includes YAML frontmatter documenting which check failed and why:

```yaml
failed_check: V2_predictive_power
reason: no novel use case

```

This structure preserves the audit trail of discard decisions, allowing later review or potential "rescue" of incorrectly filtered items.

## Key Differences at a Glance

| Aspect | candidates/ | rejected/ |
|--------|-------------|-----------|
| **Creation Timing** | After stage 1 extraction | After stage 1.5 verification |
| **Content State** | Raw, unfiltered outputs | Failed verification units |
| **File Structure** | Five aggregated markdown files per book | Individual markdown files per rejected unit |
| **Metadata** | None (raw content) | YAML frontmatter with `failed_check` and `reason` |
| **Primary Purpose** | Reproducibility and debugging | Compliance and quality audit trail |

## How to Access and Parse Audit Files

Downstream tools can programmatically analyze both directories to generate quality reports or re-evaluate extraction decisions. The `templates/INDEX.md.template` file generates the final audit view linking to both directories.

Here is a Python snippet demonstrating how to read these audit trails:

```python
import pathlib
import yaml

# Path to the book slug (replace <slug> with actual folder name)

book_root = pathlib.Path("books/<slug>")

# 1️⃣ List all raw candidates

candidates_dir = book_root / "candidates"
for md_file in candidates_dir.glob("*.md"):
    print("Candidate:", md_file.name)

# 2️⃣ List all rejected items with reasons

rejected_dir = book_root / "rejected"
for md_file in rejected_dir.glob("*.md"):
    data = yaml.safe_load(md_file.read_text())
    print(f"Rejected {md_file.name}: failed {data.get('failed_check')} – {data.get('reason')}")

```

Typical output demonstrates the lifecycle distinction:

```

Candidate: frameworks.md
Candidate: principles.md
...
Rejected f01.md: failed V2_predictive_power – no novel use case
Rejected f07.md: failed V3_exclusivity – common wisdom

```

## Summary

- **`candidates/`** stores the complete, unfiltered output from stage 1 parallel extraction across five specialized sub-agents, providing reproducibility and debugging capabilities.
- **`rejected/`** archives individual units that failed the three verification checks (cross-domain, predictive power, exclusivity) with explicit YAML-documented reasons.
- Both directories are essential for the audit trail architecture defined in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) and processed through templates in `templates/INDEX.md.template`.
- Programmatic access requires standard markdown parsing for `candidates/` and YAML frontmatter extraction for `rejected/`.

## Frequently Asked Questions

### When are the candidates/ and rejected/ directories created?

The `candidates/` directory is created immediately after **stage 1** completes, containing raw outputs from the five parallel extractors. The `rejected/` directory is populated during **stage 1.5** after the triple-verification filtering process identifies units that fail quality checks.

### What specific information is stored in rejected/ files?

Each file in `rejected/` contains YAML frontmatter with two key fields: `failed_check` (indicating which of the three verification checks failed: V1, V2, or V3) and `reason` (providing a human-readable explanation of why the unit was discarded). This metadata enables precise audit trails and potential re-evaluation.

### Can a rejected candidate be recovered or reinstated?

Yes. Because the `rejected/` directory preserves the original content alongside the failure metadata, downstream processes can implement "rescue" logic to re-evaluate discarded items. The explicit `reason` field allows developers to identify false positives or adjust verification thresholds without re-running the entire extraction pipeline.

### How does the audit trail structure ensure reproducibility?

The separation of `candidates/` (raw extraction) and `rejected/` (filtered results) creates an immutable record of the pipeline's decision-making process. By maintaining these directories as specified in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) and [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md), the system enables complete reconstruction of how final knowledge units were selected, satisfying compliance requirements for "show-your-work" auditing.