# `candidates/` vs `rejected/` Directories in cangjie-skill: Audit Pipeline Explained

> Understand cangjie-skill's candidates/ and rejected/ directories. Learn how these directories store raw data and failed verifications for complete pipeline traceability.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: internals
- Published: 2026-08-14

---

**The `candidates/` directory stores raw extracted units from Stage 1 without judgement, while `rejected/` preserves units that failed Stage 1.5 verification with explicit exclusion reasons—together they provide full traceability through the skill-building pipeline.**

The **cangjie-skill** repository implements a structured methodology for converting source books into reusable skill units. Understanding the difference between these two directories is essential for anyone auditing, contributing to, or debugging the extraction pipeline. Both folders serve complementary **audit-tracking purposes**, ensuring every decision in the workflow can be traced and reviewed.

## What the `candidates/` Directory Contains

The **`candidates/`** directory functions as the **raw extraction pool**—the output of *Stage 1* where multiple sub-agents extract information from a source book.

### When and How Candidates Are Created

Each specialized extractor writes its findings here immediately upon completion:

- `framework-extractor` → [`candidates/frameworks.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/frameworks.md)
- `principle-extractor` → [`candidates/principles.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/principles.md)
- `case-extractor` → [`candidates/cases.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/cases.md)
- `counter-example-extractor` → [`candidates/counter-examples.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/counter-examples.md)
- `glossary-extractor` → `glossary/md`

As documented in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md), these files contain **every candidate unit without judgement**. This is the **audit trail** of everything the system considered, described in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) as the "原始候选池 (审计用)" (raw candidate pool for audit)【SKILL.md†L60-L61】.

### Key Characteristics

- **No filtering applied**: All extracted content appears here, regardless of quality
- **Grouped by type**: One markdown file per extractor type, not per individual unit
- **Preserved intact**: Never modified after creation to maintain traceability

## What the `rejected/` Directory Contains

The **`rejected/`** directory serves as the **elimination pool** (淘汰池)—the collection of units that failed the *Stage 1.5* **triple-verify stage**.

### When Items Move to `rejected/`

When a candidate is inspected during verification and found unsuitable, it is moved here with explicit documentation. According to [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md), the requirement is strict: "未通过的,写入 `books/<slug>/rejected/<id>.md` … **必须写明不通过的是哪一项、原因是什么**" (for those that don't pass, write to `rejected/<id>.md` ... must specify which item failed and why)【03-stage1.5-triple-verify.md†L49-L50】.

### Structure and Content

Unlike `candidates/`, the `rejected/` directory contains:

- **One file per rejected unit**: `rejected/<id>.md` with unique identifiers
- **Mandatory reason field**: YAML frontmatter explaining the exclusion rationale
- **Original content preserved**: Full text retained for potential reinstatement

This design records **what was rejected and why**, preserving the rationale for exclusion and enabling later review or reinstatement【SKILL.md†L85-L86】.

## Practical Differences at a Glance

| Aspect | `candidates/` | `rejected/` |
|--------|---------------|-------------|
| **Pipeline stage** | Stage 1 (extraction) | Stage 1.5 (verification) |
| **Content state** | Raw, unjudged | Filtered, with rejection reasons |
| **File organization** | By extractor type ([`frameworks.md`](https://github.com/kangarooking/cangjie-skill/blob/main/frameworks.md), [`principles.md`](https://github.com/kangarooking/cangjie-skill/blob/main/principles.md), etc.) | By individual unit (`<id>.md`) |
| **Purpose** | Complete audit trail of extraction | Audit trail of exclusion decisions |
| **Mutability** | Write-once, read-only | Appended during verification |

## Working with These Directories

### Listing All Candidate Units

```bash

# From repository root

ls candidates/*.md

```

Typical output: [`candidates/cases.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/cases.md), [`candidates/frameworks.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/frameworks.md), [`candidates/principles.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/principles.md), etc.

### Adding a Rejected Entry (Manual Audit)

```bash
cat <<EOF > books/my-book/rejected/001.md
---
id: 001
type: principle
reason: "Too generic, not actionable for the target audience"
---

# Principle 001 – … (original text)

EOF

```

The YAML frontmatter is **required** per the methodology specification.

### Generating Rejection Summary for Review

```bash
grep -h "reason:" books/my-book/rejected/*.md | sort

```

This produces a quick list of all exclusion reasons, useful for identifying systematic issues in the extraction stage.

### Converting Candidates to Final Skill Units

```python
import pathlib
import yaml

def load_candidates():
    cand_dir = pathlib.Path('candidates')
    for md in cand_dir.glob('*.md'):
        # Parse YAML entries within each candidate file

        # Transform into final skill artifacts after verification

        content = md.read_text()
        # Processing logic here...

        ...

load_candidates()

```

The script reads from `candidates/` but **must cross-reference `rejected/`** to exclude failed units, leaving both source directories untouched for audit.

## Key Source Files

| File | Relevance |
|------|-----------|
| [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) | Defines directory layout and audit roles for both pools【SKILL.md†L60-L61】【SKILL.md†L85-L86】 |
| [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md) | Documents how extractors populate `candidates/` |
| [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md) | Specifies rejection workflow and reason requirement【03-stage1.5-triple-verify.md†L49-L50】 |
| [`methodology/07-stage5-deliver.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/07-stage5-deliver.md) | Confirms both directories are included in final deliverables |

## Design Philosophy: Raw → Filtered → Reasoned

The dual-directory structure mirrors a **standard data-pipeline audit pattern**. By maintaining both the complete raw material and the complete record of exclusions, **cangjie-skill** enables:

- **Forensic traceability**: Any final skill unit can be traced to its original source
- **Decision transparency**: Every exclusion has a documented justification
- **Process improvement**: Rejection patterns reveal extractor weaknesses
- **Reversible decisions**: Rejected items can be re-evaluated and reinstated

## Summary

- **`candidates/`** contains unfiltered, type-grouped extraction output from Stage 1—everything the system considered
- **`rejected/`** holds individually documented, reasoned exclusions from Stage 1.5 verification
- Both directories are **preserved intact** through delivery to maintain full audit capability
- The methodology enforces **mandatory rejection reasons** via [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md)
- Contributors interact with these directories through standard shell commands and parsing scripts

## Frequently Asked Questions

### Can items in `rejected/` be moved back to active consideration?

Yes. The `rejected/` directory preserves original content and rejection reasons precisely to enable reinstatement. If verification criteria change or initial assessments are overturned, files can be removed from `rejected/` and reprocessed.

### Why doesn't `candidates/` use individual files like `rejected/` does?

The extractor-type grouping in `candidates/` reflects the **parallel extraction architecture** where each sub-agent produces consolidated output. This batch format is efficient for initial generation, while `rejected/` uses granular files because rejection decisions happen individually during verification.

### How do I know if a final skill unit came from `candidates/` or passed through verification?

Final skill artifacts in later pipeline stages should reference their source via IDs. The audit trail requires that you can reconstruct: `candidates/` entry → verification decision → either final skill unit or `rejected/<id>.md` entry.

### Are these directories required in the final delivered skill package?

Yes. According to [`methodology/07-stage5-deliver.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/07-stage5-deliver.md), both `candidates/` and `rejected/` are included in Stage 5 delivery to preserve complete provenance for downstream users.