How Cross-Skill Trap Questions Are Designed for cangjie-skill Pressure Testing
Cross-skill trap questions are mandatory bait prompts designed to test whether a skill's trigger description is precise enough to avoid accidental activation when faced with queries intended for sibling skills from the same source material.
In the kangarooking/cangjie-skill repository, these trap questions form a critical component of Stage 4 pressure testing. They specifically target the most common real-world deployment failure: confusion between multiple skills extracted from the same source material. By forcing a sub-agent to evaluate realistic user queries without metadata hints, the methodology ensures that each skill's trigger logic maintains robust isolation from its peers.
The Purpose of Cross-Skill Trap Questions
Cross-skill trap questions (also called 诱饵 or bait tests) expose situations where multiple skills might be confused with one another. When a user submits a query that should activate Skill A, the system must ensure that Skill B (from the same book or domain) does not incorrectly trigger.
According to methodology/06-stage4-pressure-test.md, these traps verify that a skill's trigger description is discriminative enough to distinguish its intended use case from similar, adjacent capabilities. This prevents the "over-eager skill" problem where an LLM-powered agent activates multiple skills for a single intent.
Design Rules and Requirements
The methodology enforces four strict rules for designing cross-skill trap questions in the pressure-testing stage.
Mandatory Inclusion Requirement
Every skill's test suite must contain at least one cross-skill trap case. This is a hard requirement documented in the Stage 4 methodology file. No skill graduates from pressure testing without demonstrating its ability to remain silent when presented with prompts belonging to sibling skills.
Bait Prompt Construction
The prompt must be a realistic user query that should activate a different skill from the same book, not the skill under test. For example, if testing a decision-making skill like inversion-thinking, the bait prompt might be a pure information lookup that belongs to an api-lookup skill.
{
"skill": "inversion-thinking",
"version": "0.1.0",
"test_cases": [
{
"id": "cross-skill-trap-01",
"type": "should_not_trigger",
"prompt": "帮我查一下这个 API 的参数",
"expected_behavior": "纯信息查询, 不应调用任何决策 skill",
"notes": "诱饵:实际应触发另一本书的 “api‑lookup” skill"
}
]
}
In this example from the test-prompts.json schema, the query is a benign information request that should naturally trigger a lookup skill, not a thinking framework.
Blind Evaluation Protocol
When the sub-agent evaluates the prompt, it operates under strict information hiding. The agent receives only:
- The skill's description
- The user prompt
It does not see the type, expected_behavior, or notes fields. This forces the agent to decide whether to trigger the skill purely from the textual description, mimicking real-world deployment conditions where trigger metadata is unavailable.
Outcome Recording and Validation
The sub-agent must return three specific fields:
would_trigger: Boolean indicating if the skill would activatereason: Justification for the decisionif_triggered_action: The action the agent would take if triggered
The main test runner compares this output against the expected behavior defined in the trap case. A correct response shows would_trigger: false with a reason indicating the query belongs to a different functional domain.
Implementation in the Test Suite
The cross-skill trap mechanism is implemented across three key files in the repository.
JSON Schema for Trap Cases
Trap cases follow the schema defined in templates/test-prompts.json.template. Each case requires:
- A unique
idprefixed withcross-skill-trap typeset toshould_not_trigger- A
promptthat realistically belongs to a sibling skill - Detailed
notesexplaining which skill should actually handle this query
The SKILL.md meta-skill definition specifies where these test-prompts.json files are consumed during the CI pipeline.
Test Runner Execution Flow
The test runner processes traps using a blind evaluation loop. Here is the simplified execution logic:
# Pseudocode excerpt from the evaluation pipeline
for case in test_cases:
# Hide metadata from sub-agent to prevent bias
sub_agent_input = {
"skill_description": skill.description,
"user_prompt": case["prompt"]
}
result = sub_agent.evaluate(sub_agent_input)
# Compare actual trigger decision against expected behavior
evaluate_cross_skill(result, case)
The evaluate_cross_skill function checks if the agent correctly identified the prompt as belonging to a different skill. If the agent incorrectly triggers the skill under test, the trap exposes a flaw in the trigger description's specificity.
Handling Test Failures
When a cross-skill trap fails (meaning the skill incorrectly triggers on a sibling's query), the methodology in methodology/06-stage4-pressure-test.md mandates immediate remediation. The development team must either:
- Refine the trigger description to add discriminative constraints that exclude the sibling skill's domain
- Adjust the skill itself to narrow its activation criteria
This ensures robust isolation before the skill is delivered to production, preventing the costly error of multiple skills competing for the same user intent.
Summary
- Cross-skill trap questions are mandatory bait tests in Stage 4 pressure testing that verify skill isolation.
- Each skill must include at least one trap case targeting a sibling skill from the same source material.
- The blind evaluation protocol hides
typeandexpected_behaviormetadata from the sub-agent to simulate real deployment conditions. - Trap cases use the
test-prompts.jsonschema and are processed by the runner defined in theSKILL.mdpipeline. - Failures require immediate refinement of the skill's trigger description or activation logic to prevent cross-skill confusion.
Frequently Asked Questions
What makes a cross-skill trap question different from a regular negative test case?
A regular negative test case checks if a skill avoids triggering on completely irrelevant queries, while a cross-skill trap specifically uses realistic prompts that should activate a different skill from the same book. The trap tests discriminative precision between similar capabilities, not just broad relevance filtering.
How does the blind evaluation protocol prevent bias in cross-skill testing?
The protocol strips the type, expected_behavior, and notes fields from the sub-agent's input, as implemented in the test runner logic. This forces the agent to rely solely on the skill's textual description to make trigger decisions, eliminating the possibility of the agent gaming the test by matching against expected outcomes rather than evaluating semantic fit.
Where are cross-skill trap questions defined in the cangjie-skill repository?
Trap questions are defined in individual test-prompts.json files located within each skill's directory, following the template in templates/test-prompts.json.template. The methodology specification in methodology/06-stage4-pressure-test.md (lines 66-71) defines the requirements, while SKILL.md specifies how the pipeline consumes these files during automated testing.
What should developers do when a cross-skill trap test fails?
Developers must refine the skill's trigger description to add specific exclusion criteria that distinguish it from the sibling skill, or modify the skill's activation logic to narrow its scope. The methodology treats trap failures as blocking issues that must be resolved before the skill can pass Stage 4 pressure testing and proceed to deployment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →