# What Audit Trails Are Preserved in the `rejected/` Directory of cangjie-skill? A Complete Guide

> Discover what audit trails are preserved in the rejected/ directory of cangjie-skill. Learn how it stores unit failures and rejection reasons for troubleshooting.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: how-to-guide
- Published: 2026-08-15

---

**The `rejected/` directory in cangjie-skill preserves a per-candidate audit trail for every unit that fails Stage 1.5 triple-verification, storing both the original extracted content and a documented rejection reason in markdown format.**

The `cangjie-skill` repository implements a rigorous knowledge extraction pipeline that filters candidate units through multiple verification stages. Understanding what audit trails are preserved in the `rejected/` directory helps developers trace elimination decisions, recover discarded material, or debug extraction quality. This article examines the complete audit structure based on the source code and methodology documentation.

## How the `rejected/` Directory Functions as an Audit Log

The `rejected/` directory serves as a dedicated audit-log location within each book's output structure. According to [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) (lines 60–62), the path `books/<slug>/rejected/` stores "阶段 1.5 淘汰的单元 + 原因 (审计用)" — units eliminated at Stage 1.5 plus their reasons, for audit purposes.

The methodology explicitly states (lines 106–107): "不通过的写入 `books/<slug>/rejected/` 并附原因 — 保留审计轨迹,也允许用户事后捞回" — failures are written to `rejected/` with reasons attached, preserving the audit trail and allowing users to recover items later.

## What Each Audit File Contains

Every rejected candidate generates an individual markdown file at `books/<slug>/rejected/<id>.md` containing two essential components:

- **Original candidate content** — the raw extract from one of the five parallel extractors (framework, principle, case, counter-example, or glossary)
- **Clear rejection reason** — which verification criterion failed, such as "跨域验证失败" (cross-domain verification failure), "预测力不足" (insufficient predictive power), or "独特性不足" (insufficient uniqueness)

This dual-record structure ensures auditors can reconstruct *what* was rejected and *why*.

## Triple-Verification Failures That Trigger Rejection

The [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md) file defines the three checks whose failures populate the `rejected/` directory:

1. **Cross-domain verification** — validating the unit applies beyond its origin context
2. **Predictive power assessment** — evaluating whether the unit enables meaningful predictions
3. **Uniqueness verification** — confirming the unit adds distinct value not covered by existing entries

Any candidate failing one or more checks receives a permanent record in the audit trail.

## Accessing and Reading Audit Trails Programmatically

The following Python utility demonstrates how to inspect the rejected audit trail for any processed book:

```python
import os
from pathlib import Path

def list_rejected_audit(slug: str) -> None:
    """Print the audit files stored in the rejected directory for a given book."""
    rejected_dir = Path("books") / slug / "rejected"
    if not rejected_dir.is_dir():
        print("No rejected directory for this slug.")
        return

    for md_file in sorted(rejected_dir.glob("*.md")):
        print(f"--- {md_file.name} ---")
        print(md_file.read_text(encoding="utf-8"))
        print("\n")

# Example usage (replace <slug> with your actual book slug):

list_rejected_audit("<slug>")

```

This script outputs human-readable audit records combining rejected content with elimination rationale.

## Key Source Files Supporting the Audit System

| File | Role |
|------|------|
| [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) | Documents directory layout and `rejected/` purpose (lines 60–62, 106–107) |
| [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md) | Defines verification criteria triggering rejection |
| `books/<slug>/rejected/<id>.md` | Generated per-candidate audit records |

## Summary

- The `rejected/` directory preserves **complete audit trails** for Stage 1.5 eliminations
- Each audit record combines **original content** with **documented rejection reasons**
- The system supports **post-hoc recovery** of discarded material
- Five extractor types feed into this audit pipeline: framework, principle, case, counter-example, and glossary
- Source documentation in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) and methodology files explicitly codify this audit purpose

## Frequently Asked Questions

### What does "triple-verification" mean in cangjie-skill?

Triple-verification refers to the three-stage validation process defined in [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md): cross-domain verification, predictive power assessment, and uniqueness verification. Each candidate unit must pass all three checks to avoid the `rejected/` audit trail.

### Can rejected units be recovered after elimination?

Yes. The [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) methodology explicitly states that the audit trail design "允许用户事后捞回" — allows users to recover items afterward. The preserved markdown files contain complete original content, enabling reprocessing or manual review.

### How are rejection reasons formatted in audit files?

Rejection reasons appear as documented explanations within each `rejected/<id>.md` file, typically citing which verification criterion failed (跨域验证失败, 预测力不足, or 独特性不足). The exact format depends on the extractor that produced the candidate.

### Why does cangjie-skill use markdown for audit trails?

Markdown provides human-readable, version-control-friendly storage that preserves structure without proprietary formatting. This aligns with the repository's open-source philosophy and enables straightforward programmatic access using standard text processing tools.