Triple Verification (V1, V2, V3) in cangjie-skill: How It Filters Candidates

Triple Verification (V1, V2, V3) is a three-stage quality gate that filters extracted methodological units from a book, allowing only candidates that demonstrate cross-domain applicability, predictive power, and author exclusivity to become independent skill packs.

The cangjie-skill repository implements this system in its RIA-TV++ pipeline to solve a critical problem: most book excerpts are quotes, examples, or terminology—not true methodologies worth packaging as reusable skills. Located in methodology/03-stage1.5-triple-verify.md, Triple Verification sits between Parallel Extraction (Stage 1) and RIA++ Construction (Stage 2), enforcing strict quality control before any candidate advances to verified.md.


What Each Verification Level Checks

Triple Verification operates as a cumulative filter. A candidate must pass all three checks to proceed. The typical pass rate is 25–50%, reflecting the system's strictness.

V1 – Cross-Domain (Contextual Independence)

Goal: Confirm the unit appears in at least 2 independent contexts within the source book.

The find_distinct_contexts() logic scans the book for distinct chapters, stories, or objects that convey the same underlying principle. For example, "逆向思维" (reverse thinking) might appear in investment decisions, disaster avoidance, and teaching methodology.

Why it matters: Repeated appearance across domains signals a stable, author-intended methodology rather than a one-off remark or decorative example.

V2 – Predictive Power (Extrapolation Ability)

Goal: Test whether the unit yields meaningful, non-trivial conclusions when applied to novel scenarios.

The pipeline designs a novel_scenario unrelated to the book's examples, then applies the candidate methodology via apply_methodology(). The result must be both meaningful and non-obvious.

Why it matters: Genuine methods solve problems the author never explicitly addressed. Pure descriptions cannot extrapolate.

V3 – Exclusivity (Author Distinctiveness)

Goal: Exclude common-sense knowledge that any informed reader already possesses.

The is_common_sense() check flags whether the insight requires the author's unique perspective. Valid candidates must present anti-intuitive findings, specialized terminology, or distinctive reasoning frameworks that only the author would articulate.

Why it matters: Skills should encapsulate the author's contribution. Trivial commonsense does not need packaging.


How the Filtering Pipeline Executes

According to methodology/03-stage1.5-triple-verify.md (lines 11–85), the execution flow follows these steps:

  1. Aggregate outputs from 5 parallel extractors (candidates/*.md) into a single pool
  2. Deduplicate overlapping units extracted by multiple extractors
  3. Run V1 → V2 → V3 sequentially, recording pass/fail status with justification
  4. Pass → Write YAML entry to books/<slug>/verified.md
  5. Fail → Write rejection record to rejected/<id>.md documenting which verification failed
  6. User confirmation → Present counts; user approves or vetoes before Stage 2

This human-in-the-loop design ensures quality without sacrificing authorial judgment.


Verified Candidate Format

Candidates passing all three verifications are written to verified.md using a structured YAML format. The template (derived from methodology/03-stage1.5-triple-verify.md, lines 54–71) captures evidence for each check:


# books/<slug>/verified.md

id: f01
title: 逆向思维
type: framework
V1_cross_domain:
  passed: true
  evidence:
    - 第 3 讲: 投资决策场景
    - 第 7 讲: 工程设计场景
    - 第 11 讲: 教学方法场景
V2_predictive_power:
  passed: true
  novel_question: "如果面试官问我一个不知道答案的问题该怎么办?"
  derived_answer: "逆问‘我最不希望他认为我是什么样的人’,从这个反面倒推应展现的特质"
V3_exclusivity:
  passed: true
  why_not_common: "常识是‘要多想’,逆向思维是‘优先反着想’——这是反直觉的排序"

Implementation Logic (Pseudocode)

The core verification logic mirrors this Python structure:

def verify_candidate(candidate, book):
    # V1 – Cross-domain

    contexts = find_distinct_contexts(candidate, book)
    v1_pass = len(contexts) >= 2

    # V2 – Predictive Power

    novel_scenario = design_novel_scenario(candidate)
    result = apply_methodology(candidate, novel_scenario)
    v2_pass = is_meaningful(result) and not_trivial(result)

    # V3 – Exclusivity

    v3_pass = not is_common_sense(candidate) and has_unique_perspective(candidate)

    return {
        "V1_cross_domain": {"passed": v1_pass, "evidence": contexts},
        "V2_predictive_power": {
            "passed": v2_pass,
            "novel_question": novel_scenario,
            "derived_answer": result
        },
        "V3_exclusivity": {
            "passed": v3_pass,
            "why_not_common": explain_uniqueness(candidate)
        },
    }

The production pipeline uses markdown files and human verification, but the decision logic remains identical.


Key Source Files

File Purpose
README.en.md High-level Triple Verification summary (lines 34–35)
methodology/03-stage1.5-triple-verify.md Complete V1–V3 specification, workflow, and failure modes
extractors/*-extractor.md Candidate sources feeding into verification
verified.md / rejected/<id>.md Output destinations for passed and failed candidates

As documented in README.en.md, the 25–50% pass rate guarantees that only executable knowledge—methods with external validity and authorial distinctiveness—proceeds to skill pack construction.


Summary

  • Triple Verification (V1, V2, V3) filters book excerpts through three cumulative quality gates: cross-domain appearance, predictive extrapolation, and author exclusivity
  • Only candidates passing all three checks are promoted to verified.md and become independent skill packs
  • The system is implemented in methodology/03-stage1.5-triple-verify.md with 25–50% typical pass rates
  • Each verification produces structured evidence (context list, novel scenario test, uniqueness justification)
  • Human confirmation provides final approval before Stage 2 construction

Frequently Asked Questions

What happens if a candidate fails only V3 (Exclusivity)?

The candidate is written to rejected/<id>.md with documentation explaining why it was deemed common-sense. This record allows future review—if the user's domain knowledge suggests the insight is actually non-obvious, they can override the rejection during the confirmation step.

Why does V2 require "novel scenarios" not in the original book?

Predictive power distinguishes methodologies from descriptions. A true method generates valid conclusions in situations the author never considered. Without this test, the pipeline would package contextual examples that cannot transfer to new problems.

How does the 25–50% pass rate compare to other extraction systems?

Most book summarization tools retain 70–90% of extracted content. cangjie-skill's deliberately low pass rate reflects its goal: producing executable skill packs rather than comprehensive notes. The strictness ensures each output skill has genuine reuse value.

Can users modify the verification thresholds (e.g., require 3 contexts for V1)?

The current pipeline uses fixed thresholds defined in methodology/03-stage1.5-triple-verify.md. However, the human confirmation step allows case-by-case overrides. Future versions may expose configurable parameters in a config.yaml file.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →