Why the Cangjie-Skill Pipeline Limits Original-Quote Length (原文引用)

The cangjie-skill pipeline caps original quotes at ≤150 Chinese characters (or ≤100 English words) to maintain signal-to-noise ratio for Claude Code agents, ensure fair-use compliance, and guarantee consistent Markdown formatting across all skill files.

The kangariking/cangjie-skill repository implements a rigorous RIA++ (Reading, Interpretation, Application) extraction pipeline that transforms book insights into actionable AI skills. One of its most specific architectural constraints is the hard limit placed on the original-quote (原文引用) field, which governs how candidate passages graduate from initial extraction to finalized skill artifacts.

Core Architectural Purposes

Signal-to-Noise Ratio for AI Recognition

The primary driver for the character limit is optimizing the Reading (R) segment for Claude Code agent consumption. Long passages dilute the distinctive signal required to match user queries against specific skills. According to methodology/04-stage2-ria-plus.md, the specification mandates "直接引用 ≤150字 (英文原文 ≤100词)"—ensuring the excerpt remains instantly recognizable without noise from surrounding text.

By restricting excerpts to short fragments, the pipeline stays safely within fair-use boundaries and avoids reproducing large copyrighted passages. The methodology explicitly directs contributors to avoid using existing translations and instead translate content themselves, further minimizing exposure to third-party rights violations as documented in the same stage-2 methodology file.

Consistent Skill Formatting

All skills must conform to the uniform front-matter template defined in templates/SKILL.md.template. Keeping the Reading field compact ensures every skill fits cleanly into the Markdown-based structure and that downstream automation—such as test-prompts.json generation and Zettelkasten linking—can parse the source_quote field reliably without encountering oversized blocks that break the schema.

Where the Limit is Defined in Source

The constraint appears at multiple pipeline stages:

  • Candidate Schema: In methodology/02-stage1-parallel-extract.md, the candidate YAML schema defines source_quote: | … with the explicit requirement that passages must not exceed the 150-character (Chinese) or 100-word (English) threshold.
  • RIA++ Construction: The stage-2 methodology (methodology/04-stage2-ria-plus.md) reiterates this limit when defining how the R (Reading) section is constructed from validated candidates.
  • Template Structure: The templates/SKILL.md.template enforces this constraint structurally, ensuring that the source_quote field in the final skill front matter remains within bounds.

Implementation Examples

Below are concrete examples demonstrating how the limit manifests in the repository's data structures and validation logic.

Candidate Definition (Stage 1 Extraction)

In methodology/02-stage1-parallel-extract.md, candidates are defined with strict length constraints:

id: f01
title: 逆向思维
type: framework
source_chapter: 第三讲
source_quote: |
  反过来想, 总是反过来想…   # ≤150字 / English ≤100 words

summary: |
  …
tags: [decision, mental-model]

Skill Template Structure

The templates/SKILL.md.template shows how the constraint carries into the final skill file:

---
name: reverse-thinking
description: |
  当用户在纠结决策、列举正面理由却理不出头绪时触发…
source_book: 《穷查理宝典》 查理·芒格
source_chapter: 第三讲
tags: [decision, mental-model]
related_skills: []
---

## Reading (原文引用)

> 反过来想, 总是反过来想…

Validation Script Pattern

A typical validation check implemented in the pipeline uses the following logic:

MAX_CHINESE = 150          # characters

MAX_ENGLISH_WORDS = 100

def valid_quote(text: str) -> bool:
    """Validate original quote length according to cangjie-skill pipeline rules."""
    # Check for Chinese characters

    if any('\u4e00' <= ch <= '\u9fff' for ch in text):
        return len(text) <= MAX_CHINESE
    else:
        return len(text.split()) <= MAX_ENGLISH_WORDS

Summary

  • Signal clarity: The ≤150 character/≤100 word limit ensures Claude Code agents can reliably match skills to user queries without dilution from extraneous text.
  • Legal safety: Short quotes stay within fair-use doctrine and minimize copyright exposure according to the methodology in kangariking/cangjie-skill.
  • Technical consistency: The constraint guarantees uniform parsing across methodology/02-stage1-parallel-extract.md schemas and templates/SKILL.md.template outputs.
  • Pipeline integrity: Enforcement at both candidate extraction and RIA++ construction stages prevents oversized quotes from entering the skill library.

Frequently Asked Questions

What happens if a quote exceeds the 150-character limit in the cangjie-skill pipeline?

Candidates exceeding the limit are filtered out during the parallel extraction stage defined in methodology/02-stage1-parallel-extract.md. The validation logic checks source_quote length before allowing the unit to proceed to RIA++ construction, ensuring only compliant excerpts reach the final skill template.

Does the limit apply to English source material differently than Chinese text?

Yes. According to methodology/04-stage2-ria-plus.md, English passages are capped at ≤100 words while Chinese text is limited to ≤150 characters. This distinction accounts for information density differences between logographic and alphabetic writing systems while maintaining equivalent signal strength for the AI agent.

Why does the pipeline use 'words' for English but 'characters' for Chinese?

Chinese characters carry higher semantic density per unit than English words. The 150-character limit for Chinese approximates the information content of 100 English words, creating parity in signal-to-noise ratio for the Claude Code agent's matching algorithms while adhering to fair-use standards for both languages.

Is the original-quote limit configurable in the skill templates?

No. The limit is hardcoded in the pipeline methodology and enforced by the schema in templates/SKILL.md.template. This standardization ensures interoperability across the entire cangjie-skill ecosystem and prevents parsing errors in downstream automation tools that expect uniform field lengths.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →