How to Contribute to the Development of cangjie-skill: A Complete Guide
You can contribute to cangjie-skill by adding extractors in extractors/, extending the RIA++ schema in SKILL.md, improving methodology documentation, creating new skill packs, or enhancing CI automation—the pipeline is modular by design to support parallel contributions across all seven stages.
The kangarooking/cangjie-skill repository transforms books, videos, and podcasts into AI-callable "skills" through a modular pipeline. If you want to contribute to the development of cangjie-skill, the architecture is deliberately designed to let you plug into any stage of the RIA-TV++ workflow without disrupting other components. This guide covers the specific file paths, contribution patterns, and submission workflows used by the project.
Understanding the Modular RIA-TV++ Architecture
The cangjie-skill pipeline processes raw content through seven distinct stages, from initial overview to final delivery. Each stage is defined in the methodology/ directory and operates on specific artifacts.
The Seven-Stage Pipeline
-
Stage 0 – Overview: Defines project motivation and ecosystem in
methodology/00-overview.mdand the rootREADME.md. -
Stage 1 – Adler Analysis: Performs high-level reading of source text to produce
BOOK_OVERVIEW.md. -
Stage 2 – Parallel Extraction: Fires five specialized extractors defined in
extractors/framework-extractor.md,extractors/principle-extractor.md,extractors/case-extractor.md,extractors/counter-example-extractor.md, andextractors/glossary-extractor.md. -
Stage 3 – Triple Verification: Validates candidates for citation strength, predictive power, and uniqueness as documented in
methodology/03-stage1.5-triple-verify.md. -
Stage 4 – RIA++ Construction: Formats verified items into the six-field RIA++ schema defined in
SKILL.md. -
Stage 5 – Zettelkasten Linking: Builds dependency indexes following
methodology/05-stage3-zettelkasten.md. -
Stage 6 – Pressure Testing: Generates bait-question test cases according to
methodology/06-stage4-pressure-test.md. -
Stage 7 – Delivery: Emits
DIGEST.mdand installs packs into Claude Code or Cursor permethodology/07-stage5-deliver.md.
All stages rely on template files in templates/ and prompt files in extractors/ to transform data into markdown artifacts.
Six Ways to Contribute to cangjie-skill
1. Add or Improve Extractors
Create new prompt markdown files in extractors/ or refine existing ones. The five core extractors (framework, principle, case, counter-example, glossary) run in parallel during Stage 2. New extractors must output JSON conforming to the schema in SKILL.md.
2. Extend the RIA++ Schema
Modify SKILL.md to add new fields to the six-field RIA++ schema (R/I/A1/A2/E/B). Update corresponding templates in templates/ to render the new fields into BOOK_OVERVIEW.md, INDEX.md, and DIGEST.md.
3. Enhance Methodology Documentation
Clarify existing stages by editing files in methodology/. For example, update the triple verification criteria in methodology/03-stage1.5-triple-verify.md or add visual workflow diagrams to methodology/00-overview.md.
4. Create New Skill Packs
Generate complete skill directories containing BOOK_OVERVIEW.md, INDEX.md, and DIGEST.md, then submit them as new top-level directories. This requires no code changes—only properly formatted markdown artifacts.
5. Automate CI and Testing
The repository currently ships a GitHub Action for star-history generation. You can add unit-style checks for extractor prompt syntax validation or schema compliance tests in scripts/.
6. Documentation and Localization
Improve existing translations in README.en.md or README.ja.md, or add new language versions following the structure of the root README.md.
Step-by-Step Contribution Workflows
Running the Pipeline Locally
Test your changes end-to-end using the pipeline entry point. You need an OpenAI-compatible LLM endpoint configured in your environment.
# Install dependencies
pip install -r requirements.txt
# Prepare input directory
mkdir -p input && cp path/to/book.txt input/
# Execute pipeline
python scripts/run_pipeline.py input/book.txt output/
Adding a New Extractor
Create a new extractor by adding a prompt file and registering it in the configuration.
# Create the prompt file
cat > extractors/strategy-extractor.md <<'EOF'
# Strategy Extractor Prompt
You are a knowledge engineer. Scan the provided text and output every distinct
strategy, its purpose, and a concrete example. Use the following JSON schema:
{
"strategy": "<name>",
"purpose": "<short description>",
"example": "<quote or scenario>"
}
EOF
# Register in pipeline_config.yaml under the extractors: section
Ensure your extractor references the central SKILL.md definition so downstream tools can consume the output.
Submitting a New Skill Pack
Fork the repository, then use standard Git workflow to submit your generated skills.
# Create feature branch
git checkout -b new-skill-pack
# Copy generated artifacts
cp -r output/my-book my-book-skill/
# Commit and push
git add my-book-skill/
git commit -m "Add skill pack for My Book"
git push origin new-skill-pack
# Open Pull Request on kangarooking/cangjie-skill
Key Files Every Contributor Must Know
These files constitute the "contract" that contributors extend:
-
README.md– Overall project introduction, contribution guide, and ecosystem overview. -
SKILL.md– Formal definition of the RIA-TV++ skill schema used by downstream agents like Claude Code and Cursor. -
methodology/*.md– Detailed documentation for each pipeline stage, from00-overview.mdto07-stage5-deliver.md. -
extractors/*.md– LLM prompt templates for the five core extraction types. -
templates/*.md.template– Jinja-style templates that render final artifacts.
Updating any of these files automatically propagates through the pipeline and improves all subsequently generated skill packs.
Summary
- cangjie-skill uses a seven-stage RIA-TV++ pipeline defined in
methodology/and driven byextractors/andtemplates/. - Extractors are the primary extension point—add new ones by creating markdown files in
extractors/and registering them inpipeline_config.yaml. - Schema changes require updating
SKILL.mdand corresponding templates intemplates/. - Skill packs can be contributed as standalone directories without modifying core pipeline code.
- All contributions must reference the central
SKILL.mddefinition to ensure compatibility with downstream agents.
Frequently Asked Questions
What programming languages are required to contribute to cangjie-skill?
The pipeline primarily uses Python for automation scripts and Markdown for prompt engineering and documentation. Extractors are written as Markdown prompt files rather than code, making contributions accessible to domain experts without deep programming experience.
Do I need an LLM API key to test my changes?
Yes. Running the full pipeline locally requires an OpenAI-compatible LLM endpoint configured in your environment. The scripts/run_pipeline.py entry point makes API calls to process raw transcripts through the extraction stages.
How do I validate my new extractor before submitting?
Test your extractor against sample texts using the local pipeline command. Verify that output conforms to the JSON schema defined in SKILL.md and that the markdown renders correctly through the templates in templates/. Check that your prompt file syntax matches the style of existing files like extractors/framework-extractor.md.
Can I contribute without modifying the core pipeline code?
Yes. Creating new skill packs is a zero-code contribution path. Simply generate the required artifacts (BOOK_OVERVIEW.md, INDEX.md, DIGEST.md) using the existing pipeline, place them in a new directory, and submit via Pull Request. This approach is ideal for subject matter experts adding content rather than changing tooling.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →