Triple Verification in Cangjie-Skill: The 3 Criteria for High‑Quality AI Skills
Triple verification in Cangjie-Skill requires three independent criteria: ≥2 cross-domain evidences, predictive power, and uniqueness, with only 25–50% of candidates typically passing.
The kangarooking/cangjie-skill repository implements a rigorous RIA‑TV++ pipeline that transforms unstructured content—books, long‑form videos, podcasts—into structured AI skills. At the heart of this pipeline sits Stage 3: Triple Verification, a strict gating mechanism that filters out trivial or unsupported knowledge. The "TV" in RIA‑TV++ stands for Triple Verification, ensuring every skill unit is well‑substantiated, actionable, and novel.
The Three Verification Criteria
Each candidate method‑unit undergoes three independent checks. According to the README.md documentation, a unit must satisfy all three to advance.
≥ 2 Independent Evidences (跨域)
The same core idea must appear in at least two separate passages within the source material, ideally from different sections or contexts. This cross‑domain requirement guards against:
- Isolated or cherry‑picked statements
- Outlier interpretations that lack corroboration
- Fragile knowledge tied to a single mention
Evidence locations are tracked and counted. A candidate with only one supporting passage fails immediately, regardless of its apparent quality.
Predictive Power (预测力)
A valid skill unit must demonstrate extrapolation capability—it should answer questions that the source text does not state explicitly. This criterion transforms static information into actionable reasoning patterns. The unit functions as a compressed model that enables inference beyond literal retrieval.
Uniqueness (独特性)
The unit must represent distinct, non‑obvious knowledge. Common sense, generic principles, or widely known techniques are filtered out. Acceptable units include:
- Novel frameworks or mental models
- Specific methodological principles
- Counter‑intuitive insights with demonstrated utility
Pipeline Implementation Context
Triple verification operates within the broader seven‑stage RIA‑TV++ architecture:
- Raw ingestion
- Ideation and chunking
- Annotation and candidate generation
- Triple Verification (the gate described here)
- Validation and refinement
- + Integration
- + Packaging
The pass‑rate of 25–50% reflects the selectivity of this stage. Most candidates fail at least one criterion, ensuring only refined knowledge enters downstream stages.
Code‑Level Workflow Example
The following pseudo‑workflow illustrates how triple verification is applied within the extraction pipeline. While the actual implementation lives in methodology docs and extractor prompts, this structure captures the decision logic:
def triple_verify(candidate, source_text):
# 1. Check for at least two independent citations
evidence = find_evidence_locations(candidate, source_text)
if len(evidence) < 2:
return False, "insufficient_evidence"
# 2. Assess predictive power through extrapolation test
test_questions = generate_extrapolation_queries(candidate)
predictions = [candidate.apply(q) for q in test_questions]
if not any(predictions):
return False, "no_predictive_power"
# 3. Evaluate uniqueness against common‑knowledge base
similarity_score = compare_to_baseline(candidate, common_knowledge_db)
if similarity_score > UNIQUENESS_THRESHOLD:
return False, "not_unique"
return True, "verified"
The find_evidence_locations() function implements the 跨域 check, generate_extrapolation_queries() tests 预测力, and compare_to_baseline() enforces 独特性.
Design Rationale
The three criteria are orthogonal by design. A candidate could have strong evidence but lack predictive power, or be unique yet unsupported. This independence prevents gaming—no single strong signal compensates for weakness in another dimension.
The strict pass‑rate ensures that final skill packs contain only reliably invokable knowledge. For downstream AI agents, this translates to higher success rates when retrieving and applying skills during task execution.
Summary
- Triple verification is Stage 3 of the RIA‑TV++ pipeline in
kangarooking/cangjie-skill - Three mandatory criteria: ≥2 independent evidences, predictive power, and uniqueness
- Pass‑rate of 25–50% maintains high quality in exported skill sets
- Implementation combines evidence location tracking, extrapolation testing, and baseline comparison
- Source documentation located at README.md#L47
Frequently Asked Questions
What happens to candidates that fail triple verification?
Failed candidates are discarded from the skill pipeline. They do not advance to later stages and are not included in the final skill pack. The system does not currently implement fallback or recycling mechanisms for borderline cases.
Why does Cangjie‑Skill require two independent evidences instead of one?
Single‑source knowledge is fragile and context‑dependent. Cross‑domain corroboration (跨域) ensures the insight appears consistently across the material, reducing noise from outlier statements or authorial emphasis that may not generalize.
How is "predictive power" measured in practice?
The system generates novel queries that the source text does not explicitly answer. If the candidate unit can produce valid inferences or predictions for these queries, it demonstrates extrapolation capability. No human annotation of the test questions is required.
Can common techniques from specialized domains pass the uniqueness filter?
Only if they are non‑obvious within the target domain's baseline. A technique widely known among professional practitioners but absent from general knowledge bases may still qualify. The comparison runs against a curated common‑knowledge database rather than universal knowledge.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →