# Why the Cangjie-Skill Pipeline Limits Original-Quote Length (原文引用)

> Discover why the cangjie-skill pipeline limits original quotes to 150 Chinese characters. Learn how this improves Claude Code agents signal-to-noise ratio fair use and formatting.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: deep-dive
- Published: 2026-07-19

---

**The cangjie-skill pipeline caps original quotes at ≤150 Chinese characters (or ≤100 English words) to maintain signal-to-noise ratio for Claude Code agents, ensure fair-use compliance, and guarantee consistent Markdown formatting across all skill files.**

The `kangariking/cangjie-skill` repository implements a rigorous RIA++ (Reading, Interpretation, Application) extraction pipeline that transforms book insights into actionable AI skills. One of its most specific architectural constraints is the hard limit placed on the *original-quote* (原文引用) field, which governs how candidate passages graduate from initial extraction to finalized skill artifacts.

## Core Architectural Purposes

### Signal-to-Noise Ratio for AI Recognition

The primary driver for the character limit is optimizing the **Reading (R)** segment for Claude Code agent consumption. Long passages dilute the distinctive signal required to match user queries against specific skills. According to [`methodology/04-stage2-ria-plus.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/04-stage2-ria-plus.md), the specification mandates "直接引用 ≤150字 (英文原文 ≤100词)"—ensuring the excerpt remains instantly recognizable without noise from surrounding text.

### Legal and Licensing Safeguards

By restricting excerpts to short fragments, the pipeline stays safely within *fair-use* boundaries and avoids reproducing large copyrighted passages. The methodology explicitly directs contributors to avoid using existing translations and instead translate content themselves, further minimizing exposure to third-party rights violations as documented in the same stage-2 methodology file.

### Consistent Skill Formatting

All skills must conform to the uniform front-matter template defined in `templates/SKILL.md.template`. Keeping the *Reading* field compact ensures every skill fits cleanly into the Markdown-based structure and that downstream automation—such as [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) generation and Zettelkasten linking—can parse the `source_quote` field reliably without encountering oversized blocks that break the schema.

## Where the Limit is Defined in Source

The constraint appears at multiple pipeline stages:

- **Candidate Schema**: In [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md), the candidate YAML schema defines `source_quote: | …` with the explicit requirement that passages must not exceed the 150-character (Chinese) or 100-word (English) threshold.
- **RIA++ Construction**: The stage-2 methodology ([`methodology/04-stage2-ria-plus.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/04-stage2-ria-plus.md)) reiterates this limit when defining how the **R** (Reading) section is constructed from validated candidates.
- **Template Structure**: The `templates/SKILL.md.template` enforces this constraint structurally, ensuring that the `source_quote` field in the final skill front matter remains within bounds.

## Implementation Examples

Below are concrete examples demonstrating how the limit manifests in the repository's data structures and validation logic.

### Candidate Definition (Stage 1 Extraction)

In [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md), candidates are defined with strict length constraints:

```yaml
id: f01
title: 逆向思维
type: framework
source_chapter: 第三讲
source_quote: |
  反过来想, 总是反过来想…   # ≤150字 / English ≤100 words

summary: |
  …
tags: [decision, mental-model]

```

### Skill Template Structure

The `templates/SKILL.md.template` shows how the constraint carries into the final skill file:

```markdown
---
name: reverse-thinking
description: |
  当用户在纠结决策、列举正面理由却理不出头绪时触发…
source_book: 《穷查理宝典》 查理·芒格
source_chapter: 第三讲
tags: [decision, mental-model]
related_skills: []
---

## Reading (原文引用)

> 反过来想, 总是反过来想…

```

### Validation Script Pattern

A typical validation check implemented in the pipeline uses the following logic:

```python
MAX_CHINESE = 150          # characters

MAX_ENGLISH_WORDS = 100

def valid_quote(text: str) -> bool:
    """Validate original quote length according to cangjie-skill pipeline rules."""
    # Check for Chinese characters

    if any('\u4e00' <= ch <= '\u9fff' for ch in text):
        return len(text) <= MAX_CHINESE
    else:
        return len(text.split()) <= MAX_ENGLISH_WORDS

```

## Summary

- **Signal clarity**: The ≤150 character/≤100 word limit ensures Claude Code agents can reliably match skills to user queries without dilution from extraneous text.
- **Legal safety**: Short quotes stay within fair-use doctrine and minimize copyright exposure according to the methodology in `kangariking/cangjie-skill`.
- **Technical consistency**: The constraint guarantees uniform parsing across [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md) schemas and `templates/SKILL.md.template` outputs.
- **Pipeline integrity**: Enforcement at both candidate extraction and RIA++ construction stages prevents oversized quotes from entering the skill library.

## Frequently Asked Questions

### What happens if a quote exceeds the 150-character limit in the cangjie-skill pipeline?

Candidates exceeding the limit are filtered out during the parallel extraction stage defined in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md). The validation logic checks `source_quote` length before allowing the unit to proceed to RIA++ construction, ensuring only compliant excerpts reach the final skill template.

### Does the limit apply to English source material differently than Chinese text?

Yes. According to [`methodology/04-stage2-ria-plus.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/04-stage2-ria-plus.md), English passages are capped at **≤100 words** while Chinese text is limited to **≤150 characters**. This distinction accounts for information density differences between logographic and alphabetic writing systems while maintaining equivalent signal strength for the AI agent.

### Why does the pipeline use 'words' for English but 'characters' for Chinese?

Chinese characters carry higher semantic density per unit than English words. The 150-character limit for Chinese approximates the information content of 100 English words, creating parity in signal-to-noise ratio for the Claude Code agent's matching algorithms while adhering to fair-use standards for both languages.

### Is the original-quote limit configurable in the skill templates?

No. The limit is hardcoded in the pipeline methodology and enforced by the schema in `templates/SKILL.md.template`. This standardization ensures interoperability across the entire cangjie-skill ecosystem and prevents parsing errors in downstream automation tools that expect uniform field lengths.