# Understanding candidates/, rejected/, and verified.md in cangjie-skill

> Explore kangarooking/cangjie-skill and understand the purpose of candidates, rejected, and verified.md directories. Discover their role in the skill pipeline and audit trail.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: internals
- Published: 2026-07-19

---

**The `candidates/`, `rejected/`, and [`verified.md`](https://github.com/kangarooking/cangjie-skill/blob/main/verified.md) directories form a three-stage audit trail in cangjie-skill's seven-stage pipeline, separating raw extracted insights from failed verifications and final validated skills.**

The kangarooking/cangjie-skill repository implements a rigorous methodology for converting raw textual material—books, videos, and podcasts—into reusable AI skills. Central to this workflow are three distinct storage locations that track ideas from initial extraction through triple verification to final validation. These directories provide the transparency and traceability required for high-quality knowledge engineering.

## The candidates/ Directory: Raw Extraction Pool

During **Stage 1 (Parallel Extraction)**, five specialized extractors process source material simultaneously according to [`02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/02-stage1-parallel-extract.md). These extractors target specific knowledge types:

- **Frameworks**: Mental models and structural thinking tools
- **Principles**: Core rules and governing concepts  
- **Cases**: Specific examples and success stories
- **Counter-examples**: Failures and anti-patterns
- **Glossary**: Domain-specific terminology definitions

Each extractor writes its raw output to the **`candidates/`** folder, creating an initial "candidate pool" of unvetted insights. As documented in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md), files like [`candidates/frameworks.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/frameworks.md), [`candidates/principles.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/principles.md), and [`candidates/cases.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/cases.md) contain YAML-formatted extractions representing the unfiltered raw material awaiting verification.

## The rejected/ Directory: Failed Verification Archive

Not every candidate survives scrutiny. During **Stage 1.5 (Triple Verification)**, each candidate must satisfy three rigorous criteria defined in [`03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/03-stage1.5-triple-verify.md):

1. **Multiple Evidence**: Corroboration from independent sources
2. **Predictive Power**: Ability to forecast outcomes in new contexts  
3. **Novelty**: Distinct contribution beyond existing knowledge

When a candidate fails any verification check, it is immediately written to the **`rejected/`** subdirectory within the specific book's folder (e.g., `books/<slug>/rejected/<id>.md`). Each rejection file includes a clear reason for the failure, creating a **transparent audit trail** that preserves decision rationale. This architecture allows teams to later resurrect rejected items if new evidence emerges or verification criteria evolve.

## The verified.md File: Validated Skill Repository

Candidates that pass all three verification checks graduate to final status. The **[`verified.md`](https://github.com/kangarooking/cangjie-skill/blob/main/verified.md)** file, located at `books/<slug>/verified.md`, contains the **fully-formed, vetted skill definitions** ready for downstream processing. Unlike the fragmented candidate files, this single document aggregates all validated frameworks, principles, cases, and glossary entries into a coherent skill package. This file serves as the handoff point for Stage 2 (Structural Compression) and subsequent pipeline phases.

## Working with the Directory Structure

Understanding the file layouts enables effective auditing and debugging of the extraction pipeline.

```python

# Example: reading the candidate pool for a book

import pathlib, yaml

candidates_path = pathlib.Path("candidates/frameworks.md")
candidates = yaml.safe_load_all(candidates_path.read_text())
for cand in candidates:
    print(cand["title"], cand["source"])

```

```python

# Example: listing rejected items with reasons

rejected_dir = pathlib.Path("books/my-book/rejected")
for md in rejected_dir.glob("*.md"):
    print(md.name, "→", md.read_text().splitlines()[0])  # first line often holds the reason

```

## Summary

- **candidates/** stores raw, unvetted output from the five parallel extractors during Stage 1, representing the initial knowledge pool.
- **rejected/** archives candidates that failed triple verification (multiple evidence, predictive power, novelty) with explicit failure reasons, preserving full auditability.
- **verified.md** contains the final, validated skill definitions that have passed all checks and are ready for structural compression and delivery.
- Together, these locations create a **transparent, traceable workflow** from raw extraction to validated skills, as implemented in kangarooking/cangjie-skill.

## Frequently Asked Questions

### What triggers a candidate to move from candidates/ to rejected/?

A candidate moves to the `rejected/` directory when it fails any of the three verification checks during Stage 1.5: lacking multiple independent evidence sources, demonstrating insufficient predictive power for new contexts, or failing the novelty test by duplicating existing knowledge. Each rejection is documented with the specific reason in the rejection file.

### Can rejected candidates be recovered later in the cangjie-skill pipeline?

Yes. The `rejected/` directory serves as an archive rather than a deletion mechanism. Because rejections include detailed reasoning and are stored as discrete files (e.g., `books/<slug>/rejected/<id>.md`), teams can re-evaluate previously rejected items if new evidence emerges or if verification criteria are updated during later iterations.

### How does verified.md differ from the individual candidate files?

While `candidates/` contains fragmented, single-type extractions (frameworks separate from principles), [`verified.md`](https://github.com/kangarooking/cangjie-skill/blob/main/verified.md) consolidates all validated knowledge units into a single, coherent document. This file represents the final output of Stage 1.5, containing only skills that have passed triple verification and are ready for Stage 2 structural compression.

### Where are these directories located in the cangjie-skill repository structure?

The `candidates/` directory typically resides at the processing workspace root, containing files like [`frameworks.md`](https://github.com/kangarooking/cangjie-skill/blob/main/frameworks.md) and [`principles.md`](https://github.com/kangarooking/cangjie-skill/blob/main/principles.md). The `rejected/` directories are nested within individual book folders under `books/<slug>/rejected/`. The [`verified.md`](https://github.com/kangarooking/cangjie-skill/blob/main/verified.md) files appear alongside them at `books/<slug>/verified.md`, as specified in the pipeline documentation within [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md).