How to Determine if Content is Suitable for Cangjie-Skill Distillation: 5 Validation Criteria

Content is suitable for Cangjie-Skill distillation only when it contains discrete, actionable methodological units—such as frameworks, checklists, or decision procedures—that can be expressed with clear trigger signals and verified through the Triple Verification process.

The kangarooking/cangjie-skill repository transforms high-value material like books, long videos, and courses into agent-callable skills. To determine if content is suitable for cangjie-skill distillation, you must evaluate it against five strict criteria defined in Stage 0 – Adler analysis. Only material that maps cleanly to the R-I-A1-A2-E-B structure and survives the full validation pipeline—including Triple Verification and pressure testing—qualifies for skill-ization.

The 5 Suitability Criteria for Cangjie-Skill Distillation

According to methodology/01-stage0-adler.md, suitable content must satisfy all five invariants to proceed through the RIA-TV++ pipeline.

Extractable Methodology

The content must contain frameworks, checklists, principles, or decision procedures that can be decomposed into executable steps. In methodology/01-stage0-adler.md (lines 40-42), the authors specify that only units describing how to do something—such as a "5-step product launch" or "SWOT analysis"—can be broken into the six-field SKILL.md template. Pure theoretical discussion without actionable structure fails this criterion.

Clear Trigger Signals

Suitable content explicitly defines when the method applies and when it does not, including key activation phrases. This enables population of the A2 (Trigger) field in the skill template. Without clear conditional logic—such as "use this framework when market saturation exceeds 80%"—the resulting skill cannot be reliably invoked by an AI agent.

Atomic Scope

The content must address a single, well-bounded problem rather than a broad narrative. This satisfies Invariant 1 – Atomicity documented in methodology/00-overview.md (lines 76-79). If a section covers multiple disconnected topics or requires more than a discrete methodological unit to execute, it violates atomicity and must be split before distillation.

Verifiable via Triple Verification

Candidates must appear at least twice independently in the source material, demonstrate predictive power, and remain non-common-sense. As described in methodology/00-overview.md, the Triple Verification stage (Stage 1.5) filters out coincidental patterns and generic advice like "do your best" or "be honest." Only novel, reproducible methodologies survive this gate.

Traceable Source

Each methodological unit must tie to a specific book chapter, video timestamp, or podcast episode. This satisfies Invariant 2 – Traceability (methodology/00-overview.md, lines 78-80) and populates the Boundary (B) field in the final skill. Skills without verifiable provenance cannot be audited or updated when source material changes.

Content Types to Exclude from Distillation

The Stage 0 analysis explicitly disqualifies four categories of content that lack the structural prerequisites for agent execution.

  • Pure historical facts: Lists of dates, events, or raw data contain no actionable procedure or decision logic to expose as a skill (methodology/01-stage0-adler.md, lines 41-42).

  • Pure narrative or storytelling: Anecdotes without reusable patterns lack the framework/principle structure required for the R-I-A1-A2-E-B mapping.

  • Pure emotional content: Motivational prose and personal feelings cannot be reduced to deterministic trigger-action pairs suitable for agent invocation.

  • Overly broad concepts: Abstract themes like "leadership" or "success" without concrete steps violate the atomicity invariant and produce skills too vague to trigger reliably.

The Validation Pipeline: From Adler Analysis to Pressure Testing

Understanding why these criteria matter requires following the architectural flow from raw content to deployed skill.

  1. Stage 0 – Adler analysis: Extracts the source material's structure and identifies skill-worthy units (框架/清单/原则/决策程序) while discarding unsuitable material per the criteria above.

  2. Stage 1 – Parallel extraction: Five specialized extractors—defined in the extractors/ directory—process framework, principle, case, counter-example, and glossary units simultaneously.

  3. Stage 1.5 – Triple Verification: Cross-validates candidates against the three verification rules (independent occurrence, predictive power, uniqueness).

  4. Stage 2 – RIA++ construction: Maps each validated unit into templates/SKILL.md.template, filling the six fields: Role, Instruction, A1ction, A2 Trigger, Example, and Boundary.

  5. Stage 4 – Pressure Test: Generates test-prompts.json (Darwin-compatible) to verify that the A2 trigger fires correctly and the skill does not activate on unrelated prompts, as detailed in methodology/06-stage4-pressure-test.md.

Only material surviving all five stages becomes a callable skill.

Implementing the Suitability Check in Code

While the production pipeline uses the five extractors and Triple Verification rules, you can implement a heuristic check using the following logic:

def is_suitable(candidate_text: str, source_ref: str) -> bool:
    # 1️⃣ Does it describe a framework / checklist / principle / decision procedure?

    if not any(keyword in candidate_text.lower()
               for keyword in ["step", "process", "framework", "principle", "checklist"]):
        return False

    # 2️⃣ Is there a clear trigger description?

    if "when to use" not in candidate_text.lower():
        return False

    # 3️⃣ Is the unit atomic (single problem)?

    if len(candidate_text.split("\n")) > 8:  # heuristic: too long = likely multi-topic

        return False

    # 4️⃣ Does it appear at least twice in the source?

    occurrences = source_ref.count(candidate_text[:30])  # rough match

    if occurrences < 2:
        return False

    # 5️⃣ Is it non-trivial (not common sense)?

    if candidate_text.lower().strip() in {"do your best", "be honest", "stay focused"}:
        return False

    return True

This function mirrors the suitability gates described in methodology/01-stage0-adler.md, checking for methodological structure, trigger clarity, atomicity, source redundancy, and non-triviality before passing content to the distillation pipeline.

Summary

To determine if content is suitable for cangjie-skill distillation, evaluate it against these core principles:

  • Actionable structure: Content must contain frameworks, checklists, or decision procedures, not just facts or stories.
  • Trigger clarity: The material must define explicit activation conditions for the A2 field.
  • Atomic scope: Each skill addresses one bounded problem to satisfy Invariant 1.
  • Triple Verification: Content must appear twice independently, demonstrate predictive power, and avoid common-sense platitudes.
  • Full pipeline survival: Material must pass Stage 0 Adler analysis, Stage 1 extraction, Triple Verification, RIA++ construction, and Stage 4 pressure testing.

Frequently Asked Questions

What does the R-I-A1-A2-E-B structure represent in Cangjie-Skill?

The R-I-A1-A2-E-B structure is the canonical schema defined in templates/SKILL.md.template. Role defines the skill's persona, Instruction provides context, A1 specifies the action to take, A2 defines the trigger conditions, E provides an example, and B establishes the boundary/source traceability. This six-field format ensures the skill is atomic, traceable, and callable by AI agents.

Why does pure narrative content fail the suitability test?

Pure narrative or storytelling lacks the reusable methodological framework required for agent execution. While anecdotes may illustrate a point, they cannot be reduced to deterministic trigger-action pairs or populated into the A1/A2 fields of the SKILL.md template. The Stage 0 analysis explicitly excludes 纯故事 (pure stories) from distillation.

How does Triple Verification ensure content quality?

Triple Verification (Stage 1.5) requires candidates to appear at least twice independently in the source material, demonstrate predictive power (the method predicts outcomes better than chance), and remain non-common-sense. This filters out generic advice and coincidental patterns, ensuring only novel, reproducible methodologies proceed to the RIA++ construction phase.

What happens during the Stage 4 Pressure Test?

During Stage 4 – Pressure Testing (methodology/06-stage4-pressure-test.md), the system generates adversarial test prompts to verify the skill's A2 trigger fires correctly under expected conditions and remains silent on unrelated queries. This produces test-prompts.json files compatible with Darwin evaluation frameworks, ensuring the skill does not hallucinate or activate inappropriately when deployed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →