What Audit Trails Are Preserved in the `rejected/` Directory of cangjie-skill? A Complete Guide

The rejected/ directory in cangjie-skill preserves a per-candidate audit trail for every unit that fails Stage 1.5 triple-verification, storing both the original extracted content and a documented rejection reason in markdown format.

The cangjie-skill repository implements a rigorous knowledge extraction pipeline that filters candidate units through multiple verification stages. Understanding what audit trails are preserved in the rejected/ directory helps developers trace elimination decisions, recover discarded material, or debug extraction quality. This article examines the complete audit structure based on the source code and methodology documentation.

How the rejected/ Directory Functions as an Audit Log

The rejected/ directory serves as a dedicated audit-log location within each book's output structure. According to SKILL.md (lines 60–62), the path books/<slug>/rejected/ stores "阶段 1.5 淘汰的单元 + 原因 (审计用)" — units eliminated at Stage 1.5 plus their reasons, for audit purposes.

The methodology explicitly states (lines 106–107): "不通过的写入 books/<slug>/rejected/ 并附原因 — 保留审计轨迹,也允许用户事后捞回" — failures are written to rejected/ with reasons attached, preserving the audit trail and allowing users to recover items later.

What Each Audit File Contains

Every rejected candidate generates an individual markdown file at books/<slug>/rejected/<id>.md containing two essential components:

  • Original candidate content — the raw extract from one of the five parallel extractors (framework, principle, case, counter-example, or glossary)
  • Clear rejection reason — which verification criterion failed, such as "跨域验证失败" (cross-domain verification failure), "预测力不足" (insufficient predictive power), or "独特性不足" (insufficient uniqueness)

This dual-record structure ensures auditors can reconstruct what was rejected and why.

Triple-Verification Failures That Trigger Rejection

The methodology/03-stage1.5-triple-verify.md file defines the three checks whose failures populate the rejected/ directory:

  1. Cross-domain verification — validating the unit applies beyond its origin context
  2. Predictive power assessment — evaluating whether the unit enables meaningful predictions
  3. Uniqueness verification — confirming the unit adds distinct value not covered by existing entries

Any candidate failing one or more checks receives a permanent record in the audit trail.

Accessing and Reading Audit Trails Programmatically

The following Python utility demonstrates how to inspect the rejected audit trail for any processed book:

import os
from pathlib import Path

def list_rejected_audit(slug: str) -> None:
    """Print the audit files stored in the rejected directory for a given book."""
    rejected_dir = Path("books") / slug / "rejected"
    if not rejected_dir.is_dir():
        print("No rejected directory for this slug.")
        return

    for md_file in sorted(rejected_dir.glob("*.md")):
        print(f"--- {md_file.name} ---")
        print(md_file.read_text(encoding="utf-8"))
        print("\n")

# Example usage (replace <slug> with your actual book slug):

list_rejected_audit("<slug>")

This script outputs human-readable audit records combining rejected content with elimination rationale.

Key Source Files Supporting the Audit System

File Role
SKILL.md Documents directory layout and rejected/ purpose (lines 60–62, 106–107)
methodology/03-stage1.5-triple-verify.md Defines verification criteria triggering rejection
books/<slug>/rejected/<id>.md Generated per-candidate audit records

Summary

  • The rejected/ directory preserves complete audit trails for Stage 1.5 eliminations
  • Each audit record combines original content with documented rejection reasons
  • The system supports post-hoc recovery of discarded material
  • Five extractor types feed into this audit pipeline: framework, principle, case, counter-example, and glossary
  • Source documentation in SKILL.md and methodology files explicitly codify this audit purpose

Frequently Asked Questions

What does "triple-verification" mean in cangjie-skill?

Triple-verification refers to the three-stage validation process defined in methodology/03-stage1.5-triple-verify.md: cross-domain verification, predictive power assessment, and uniqueness verification. Each candidate unit must pass all three checks to avoid the rejected/ audit trail.

Can rejected units be recovered after elimination?

Yes. The SKILL.md methodology explicitly states that the audit trail design "允许用户事后捞回" — allows users to recover items afterward. The preserved markdown files contain complete original content, enabling reprocessing or manual review.

How are rejection reasons formatted in audit files?

Rejection reasons appear as documented explanations within each rejected/<id>.md file, typically citing which verification criterion failed (跨域验证失败, 预测力不足, or 独特性不足). The exact format depends on the extractor that produced the candidate.

Why does cangjie-skill use markdown for audit trails?

Markdown provides human-readable, version-control-friendly storage that preserves structure without proprietary formatting. This aligns with the repository's open-source philosophy and enables straightforward programmatic access using standard text processing tools.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →