How to Add Custom Extractors Beyond the Default Five in Cangjie-Skill
To add custom extractors to the cangjie-skill repository, create a new *-extractor.md prompt file in the extractors/ directory, register it in the SKILL.md table, and the parallel extraction pipeline will automatically execute it alongside the five built-in extractors.
The cangjie-skill repository (kangaroking/cangjie-skill) ships with five built-in parallel extractors—framework, principle, case, counter-example, and glossary—that analyze source material through Markdown prompt templates. To extend the system with custom extractors, you only need to define a new prompt file and update the skill documentation, as the extraction stage automatically discovers and executes every file matching the *-extractor.md pattern without requiring changes to the core extraction logic.
Where the Default Extractors Live
The five default extractors are defined as Markdown prompt templates inside the extractors/ directory. Each file follows the strict naming convention <type>-extractor.md, such as extractors/framework-extractor.md or extractors/glossary-extractor.md. These templates contain the LLM prompts that extract specific types of information from the source text, and they serve as the reference structure for new custom extractors.
Step-by-Step: Create and Register a Custom Extractor
1. Create the Extractor Prompt File
Add a new file named <your-type>-extractor.md to the extractors/ directory. Follow the structure of existing files like extractors/framework-extractor.md, providing a clear description, example inputs/outputs, and the actual prompt template.
Use the {{source_text}} placeholder variable to reference the content being analyzed:
# Algorithm Extractor
> Purpose: Pull out algorithmic procedures or step‑by‑step methods described in the source material.
## Prompt
You are an expert who extracts concise algorithmic steps from the following text. Return each step as a numbered list. If no algorithm is present, respond with "N/A".
{{source_text}}
Save this as extractors/algorithm-extractor.md.
2. Register in SKILL.md
Open SKILL.md and locate the Extractors table (approximately lines 80-90). Add a new row following the existing format to make your extractor discoverable:
| 算法提取器 | `extractors/algorithm-extractor.md` | 算法/步骤 |
This registration step documents the extractor's purpose and file path for users reviewing the skill definition.
3. Verify Parallel Execution Behavior
The extraction stage defined in methodology/02-stage1-parallel-extract.md runs all *-extractor.md files in parallel. The pipeline automatically builds a list of prompts by reading the extractors/ folder, so no manual registration in the execution logic is required. If you need to exclude specific extractors or enforce a specific execution order, modify the runner script that collects these prompts before the parallel stage executes.
Testing Your Custom Extractor
Run the skill locally using the appropriate entry point (e.g., npm run skill). Verify that your extractor's output appears alongside the default five results and that the overall skill output produces a coherent summary. Check that the LLM correctly handles the {{source_text}} substitution and returns data in your expected format.
Summary
- Custom extractors are Markdown prompt templates stored as
extractors/<type>-extractor.md. - The pipeline in
methodology/02-stage1-parallel-extract.mdauto-discovers all*-extractor.mdfiles and executes them in parallel. - Registration requires only adding a row to the Extractors table in
SKILL.md(around lines 80-90). - No code changes are required to extend the system beyond the five default extractors.
Frequently Asked Questions
Do I need to modify the extraction pipeline code to add a custom extractor?
No. According to the cangjie-skill source code, the parallel extraction stage automatically loads every file matching extractors/*-extractor.md from the directory. You only need to create the Markdown file and document it in SKILL.md.
What file naming convention should I use for custom extractors?
Name your file <descriptive-type>-extractor.md (for example, algorithm-extractor.md or pattern-extractor.md). The pipeline specifically looks for files ending in -extractor.md within the extractors/ directory to build its execution list.
Can I control the execution order of extractors?
Yes. While the default behavior in methodology/02-stage1-parallel-extract.md runs all extractors simultaneously, you can modify the runner script that collects the prompt files to specify a sequential order or exclude specific extractors from the list before the parallel stage executes.
How do I reference the source text inside my custom extractor prompt?
Use the {{source_text}} placeholder variable in your Markdown template. The extraction pipeline substitutes this with the actual content being analyzed when the extractor runs, allowing the LLM to process the specific input text for your custom extraction logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →