Understanding candidates/, rejected/, and verified.md in cangjie-skill
The candidates/, rejected/, and verified.md directories form a three-stage audit trail in cangjie-skill's seven-stage pipeline, separating raw extracted insights from failed verifications and final validated skills.
The kangarooking/cangjie-skill repository implements a rigorous methodology for converting raw textual material—books, videos, and podcasts—into reusable AI skills. Central to this workflow are three distinct storage locations that track ideas from initial extraction through triple verification to final validation. These directories provide the transparency and traceability required for high-quality knowledge engineering.
The candidates/ Directory: Raw Extraction Pool
During Stage 1 (Parallel Extraction), five specialized extractors process source material simultaneously according to 02-stage1-parallel-extract.md. These extractors target specific knowledge types:
- Frameworks: Mental models and structural thinking tools
- Principles: Core rules and governing concepts
- Cases: Specific examples and success stories
- Counter-examples: Failures and anti-patterns
- Glossary: Domain-specific terminology definitions
Each extractor writes its raw output to the candidates/ folder, creating an initial "candidate pool" of unvetted insights. As documented in SKILL.md, files like candidates/frameworks.md, candidates/principles.md, and candidates/cases.md contain YAML-formatted extractions representing the unfiltered raw material awaiting verification.
The rejected/ Directory: Failed Verification Archive
Not every candidate survives scrutiny. During Stage 1.5 (Triple Verification), each candidate must satisfy three rigorous criteria defined in 03-stage1.5-triple-verify.md:
- Multiple Evidence: Corroboration from independent sources
- Predictive Power: Ability to forecast outcomes in new contexts
- Novelty: Distinct contribution beyond existing knowledge
When a candidate fails any verification check, it is immediately written to the rejected/ subdirectory within the specific book's folder (e.g., books/<slug>/rejected/<id>.md). Each rejection file includes a clear reason for the failure, creating a transparent audit trail that preserves decision rationale. This architecture allows teams to later resurrect rejected items if new evidence emerges or verification criteria evolve.
The verified.md File: Validated Skill Repository
Candidates that pass all three verification checks graduate to final status. The verified.md file, located at books/<slug>/verified.md, contains the fully-formed, vetted skill definitions ready for downstream processing. Unlike the fragmented candidate files, this single document aggregates all validated frameworks, principles, cases, and glossary entries into a coherent skill package. This file serves as the handoff point for Stage 2 (Structural Compression) and subsequent pipeline phases.
Working with the Directory Structure
Understanding the file layouts enables effective auditing and debugging of the extraction pipeline.
# Example: reading the candidate pool for a book
import pathlib, yaml
candidates_path = pathlib.Path("candidates/frameworks.md")
candidates = yaml.safe_load_all(candidates_path.read_text())
for cand in candidates:
print(cand["title"], cand["source"])
# Example: listing rejected items with reasons
rejected_dir = pathlib.Path("books/my-book/rejected")
for md in rejected_dir.glob("*.md"):
print(md.name, "→", md.read_text().splitlines()[0]) # first line often holds the reason
Summary
- candidates/ stores raw, unvetted output from the five parallel extractors during Stage 1, representing the initial knowledge pool.
- rejected/ archives candidates that failed triple verification (multiple evidence, predictive power, novelty) with explicit failure reasons, preserving full auditability.
- verified.md contains the final, validated skill definitions that have passed all checks and are ready for structural compression and delivery.
- Together, these locations create a transparent, traceable workflow from raw extraction to validated skills, as implemented in kangarooking/cangjie-skill.
Frequently Asked Questions
What triggers a candidate to move from candidates/ to rejected/?
A candidate moves to the rejected/ directory when it fails any of the three verification checks during Stage 1.5: lacking multiple independent evidence sources, demonstrating insufficient predictive power for new contexts, or failing the novelty test by duplicating existing knowledge. Each rejection is documented with the specific reason in the rejection file.
Can rejected candidates be recovered later in the cangjie-skill pipeline?
Yes. The rejected/ directory serves as an archive rather than a deletion mechanism. Because rejections include detailed reasoning and are stored as discrete files (e.g., books/<slug>/rejected/<id>.md), teams can re-evaluate previously rejected items if new evidence emerges or if verification criteria are updated during later iterations.
How does verified.md differ from the individual candidate files?
While candidates/ contains fragmented, single-type extractions (frameworks separate from principles), verified.md consolidates all validated knowledge units into a single, coherent document. This file represents the final output of Stage 1.5, containing only skills that have passed triple verification and are ready for Stage 2 structural compression.
Where are these directories located in the cangjie-skill repository structure?
The candidates/ directory typically resides at the processing workspace root, containing files like frameworks.md and principles.md. The rejected/ directories are nested within individual book folders under books/<slug>/rejected/. The verified.md files appear alongside them at books/<slug>/verified.md, as specified in the pipeline documentation within SKILL.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →