cangjie-skill Directory Structure Explained: candidates/, rejected/, and verified.md

The candidates/, rejected/, and verified.md directories in cangjie-skill serve as a three-stage audit trail: raw extractor outputs, failed verifications with reasons, and finalized skills ready for delivery.

The cangjie-skill repository implements a rigorous seven-stage pipeline for transforming raw source material into reusable AI skills. Each directory plays a specific role in ensuring transparency, traceability, and quality control throughout this process. This article examines how these three locations function within the pipeline architecture defined in SKILL.md.

The cangjie-skill Pipeline Context

Before diving into directory specifics, understand that cangjie-skill processes books, videos, and podcasts through systematic extraction and verification. The pipeline separates raw candidate generation from quality assurance, creating clear decision points that later stages can audit and reproduce.

candidates/: The Raw Extraction Pool

The candidates/ directory holds initial outputs from Stage 1 — Parallel Extraction. Five specialized extractors run simultaneously against source material:

  • frameworks.md — mental models and structural approaches
  • principles.md — core rules and guidelines
  • cases.md — illustrative examples
  • counter-examples.md — boundary conditions and exceptions
  • glossary.md — domain vocabulary definitions

Each extractor writes its findings as YAML-fronted markdown files to this directory. These represent unverified claims that require subsequent validation.

From 02-stage1-parallel-extract.md, the parallel extraction design maximizes coverage while keeping outputs decoupled for independent verification.


# Example: reading the candidate pool for a book

import pathlib, yaml

candidates_path = pathlib.Path("candidates/frameworks.md")
candidates = yaml.safe_load_all(candidates_path.read_text())
for cand in candidates:
    print(cand["title"], cand["source"])

rejected/: The Audit Trail for Failed Verifications

The rejected/ subdirectory appears within individual book folders and stores candidates that failed Stage 1.5 — Triple Verification. According to 03-stage1.5-triple-verify.md, every candidate must satisfy three criteria:

  1. Multiple evidence — corroboration across sources
  2. Predictive power — demonstrated utility for forecasting outcomes
  3. Novelty — distinct from existing skills in the knowledge base

When a candidate fails any check, it moves to rejected/ with a documented reason preserved in the file header. This design enables several downstream benefits:

  • Re-evaluation — new evidence may resurrect previously rejected items
  • Pattern analysis — identify systematic extraction weaknesses
  • Compliance — maintain complete decision records for audit purposes

# Example: listing rejected items with reasons

rejected_dir = pathlib.Path("books/my-book/rejected")
for md in rejected_dir.glob("*.md"):
    print(md.name, "→", md.read_text().splitlines()[0])  # first line often holds the reason

Typical rejection locations follow the pattern books/<slug>/rejected/<id>.md.

verified.md: The Final Skill Deliverable

The verified.md file contains validated skills ready for downstream consumption. Located at books/<slug>/verified.md, this single file aggregates all candidates that passed triple verification.

Key characteristics of verified.md:

  • Structured format — consistent YAML frontmatter with skill metadata
  • Link-ready — formatted for cross-referencing and composition
  • Pressure-test eligible — qualified for Stage 2 validation

Unlike the distributed candidates/ directory, verified.md represents a curated, consolidated output. It serves as the handoff point between extraction/verification phases and subsequent pipeline stages involving skill refinement and deployment.

Directory Comparison: Full Reference

Location Pipeline Stage Content State Persistence Model
candidates/ Stage 1 Raw, unverified Multiple files (frameworks.md, principles.md, etc.)
rejected/ Stage 1.5 Failed verification Individual .md files with rejection reasons
verified.md Stage 1.5 complete Validated and approved Single consolidated file per book

Workflow Integration

The three directories create a linear progression with full traceability:

  1. Extractors populate candidates/ with potential skills
  2. Verification moves failures to rejected/ and successes toward verified.md
  3. Downstream stages consume only verified.md contents

This architecture prevents contamination of the final skill set while preserving reversibility — any decision can be revisited with full context.

Summary

  • candidates/ stores raw, parallel-extracted outputs from five specialized extractors before any quality checks
  • rejected/ maintains failed candidates with explicit rejection reasons, enabling re-evaluation and audit trails
  • verified.md delivers consolidated, validated skills ready for downstream pipeline stages

The cangjie-skill directory structure reflects defensive design: maximize extraction coverage, rigorously filter through transparent criteria, and preserve all intermediate states for reproducibility and improvement.

Frequently Asked Questions

Why does cangjie-skill keep rejected candidates instead of deleting them?

Rejected candidates are preserved to enable future re-evaluation when new evidence emerges, support pattern analysis to improve extractors, and maintain complete audit trails for compliance and debugging. The SKILL.md documentation emphasizes that knowledge cutoff dates and evolving source material may render previously rejected items valid later.

Can I manually move items from rejected/ back to candidates/?

Yes. The pipeline treats rejected/ as an organizational convenience, not a permanent exile. Moving a file back to candidates/ and re-running verification will re-evaluate it against current criteria. This manual override supports edge cases where automated checks incorrectly filtered valid skills.

How does verified.md differ from the individual candidates/ files?

verified.md contains only items that passed triple verification, formatted as a single consolidated document with enriched metadata. The candidates/ files remain unaltered from extractor output and may contain overlapping, contradictory, or unverified claims. Think of candidates/ as a working draft and verified.md as the published edition.

What happens to verified.md in later pipeline stages?

verified.md feeds into Stage 2 — Pressure Testing, where skills undergo simulated application scenarios. Surviving skills may then proceed to linking, composition, and eventual inclusion in the master skill registry. The verified file thus serves as the gateway artifact between raw extraction and refined delivery.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →