# How to Handle Long Texts Exceeding Agent Context Limits in cangjie-skill

> Learn how cangjie-skill handles long texts exceeding agent context limits using parallel chunking and specialized extractors. Process large documents efficiently without context issues.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: how-to-guide
- Published: 2026-08-15

---

**cangjie-skill uses a parallel chunking pipeline that splits large documents into ~2,000-token segments and processes each through specialized extractors, ensuring no single LLM call ever exceeds context limits.**

The `kangarooking/cangjie-skill` repository implements the **RIA-TV++** pipeline, a methodology specifically designed to transform books, video transcripts, and podcast scripts into structured skill packs without overwhelming language model context windows. This approach fundamentally avoids the common failure mode of dumping entire documents into a single prompt.

## The Core Strategy: Never Load Full Text Into One Request

Rather than attempting to compress or truncate source material, cangjie-skill **architecturally prevents** context overflow through six staged operations. Each stage operates on progressively smaller data volumes:

| Stage | Operation | Context Safety Mechanism |
|-------|-----------|--------------------------|
| 1 | **Overall-content understanding** (Adler analysis) | Generates [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) from skimming; no full text in prompt |
| 2 | **Parallel extraction** with 5 extractors | Source chunked into ~2k-token blocks; extractors run concurrently |
| 3 | **Triple verification** | Cross-checks compact candidate items, not raw source |
| 4 | **RIA++ construction** | Assembles filtered items into schema (`R/I/A1/A2/E/B`) |
| 5 | **Zettelkasten linking** | Operates on lightweight skill graph |
| 6 | **Pressure testing** | Tests final skill definitions only |

By **never loading the full source into a single prompt**, the system remains compatible with any model's token ceiling regardless of input length.

## Stage-by-Stage Context Management

### Stage 1: Adler Analysis for High-Level Overview

The pipeline first generates [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) through a focused summarization prompt. This step intentionally avoids ingesting the complete document, instead using targeted extraction to produce a concise summary that fits easily within standard context limits.

### Stage 2: Parallel Chunked Extraction

The critical implementation resides in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md). This stage:

- **Splits** source text into approximately 2,000-token chunks
- **Dispatches** each chunk to five specialized extractors simultaneously:
  - [`principle-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/principle-extractor.md)
  - [`framework-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/framework-extractor.md)
  - [`case-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/case-extractor.md)
  - [`counter-example-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/counter-example-extractor.md)
  - [`glossary-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/glossary-extractor.md)

Each extractor prompt in `extractors/*.md` is designed to accept a text fragment and return structured candidates. The parallel execution ensures wall-clock time remains reasonable even for 100k+ token inputs.

### Stage 3: Triple Verification on Compact Data

Documented in [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md), this stage validates extractor outputs by requiring cross-confirmation across at least two independent passages. Crucially, verification operates on **already-extracted candidates**—a dramatically smaller dataset than the original source.

### Stages 4-6: Progressive Reduction

The remaining stages manipulate only filtered, structured data:

- **RIA++ construction** follows [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) schema definitions
- **Zettelkasten linking** uses the generated skill graph
- **Pressure testing** validates final skill definitions

None of these stages touch the original long-form text.

## Key Implementation Files

| File Path | Purpose |
|-----------|---------|
| [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md) | Chunking strategy and parallel dispatch specification |
| [`extractors/principle-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/principle-extractor.md) | Template for principle extraction from chunks |
| [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md) | Cross-validation rules for candidate quality |
| [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) | RIA++ schema definition (`R/I/A1/A2/E/B` structure) |
| `templates/INDEX.md.template` | Skill map generation (minimal, LLM-friendly output) |

## Practical Pipeline Execution

The command-line interface implements this architecture through chained operations:

```bash

# Process a long transcript without context overflow

cangjie-skill run-pipeline \
  --input lengthy-transcript.txt \
  --output-dir ./skill-pack

```

Internally, `run-pipeline` executes:

```python

# Conceptual flow based on cangjie_skill package structure

from cangjie_skill import utils, extractor, verifier, writer

# 1. Chunk the source

chunks = utils.chunk_text(source="lengthy-transcript.txt", max_tokens=2000)

# 2. Parallel extraction across five extractors

candidates = extractor.run_parallel(
    chunks=chunks,
    extractors=["principle", "framework", "case", "counter-example", "glossary"]
)

# 3. Triple verification on compact candidates

verified = verifier.triple_verify(candidates)

# 4. Assemble final skill pack

writer.write_skill_pack(items=verified, output_dir="./skill-pack")

```

Each function operates on bounded data: `chunk_text` enforces token limits, `run_parallel` processes slices independently, and subsequent stages handle only filtered results.

## Performance Characteristics

| Input Size | Chunk Count | Parallel Extractor Calls | Wall-Clock Impact |
|------------|-------------|--------------------------|-------------------|
| 10k tokens | 5 | 25 concurrent | ~1x single-chunk time |
| 50k tokens | 25 | 125 concurrent | ~5x (amortized via parallelism) |
| 100k tokens | 50 | 250 concurrent | ~10x (maintained through parallel extraction) |

The constant factor per chunk remains stable because extractors never see more than their designated ~2k-token window.

## Summary

- **Chunk-first architecture** guarantees no single request exceeds model limits
- **Five parallel extractors** process segments concurrently for efficiency
- **Triple verification** prunes low-quality candidates early, reducing downstream data volume
- **RIA++ schema** produces compact, agent-ready output from arbitrarily large inputs

## Frequently Asked Questions

### What happens if a single chunk still exceeds the context limit?

The `utils.chunk_text` function in the cangjie-skill package enforces hard token ceilings. If segmenting at natural boundaries produces a chunk near the limit, it further subdivides at sentence or paragraph breaks. The extractor prompts are also optimized for ~1,500-token inputs, leaving headroom even with 2k-token chunks.

### Can I adjust the chunk size for different models?

Yes. The chunking parameters are configurable in the pipeline configuration. Smaller models (4k context) benefit from ~1k-token chunks, while larger models (32k+ context) can use larger segments. The extractor prompts in `extractors/*.md` remain compatible across these ranges.

### Does chunking lose cross-chunk relationships?

The triple-verification stage ([`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md)) specifically addresses this by requiring independent confirmation across multiple passages. Items appearing only once are filtered out, ensuring captured principles and frameworks have sufficient support across the source material.

### How does cangjie-skill compare to simple truncation or summarization?

Truncation discards content arbitrarily. Single-pass summarization loses granular detail. The RIA-TV++ pipeline preserves extractable knowledge through **dedicated extractors** per concept type, **cross-validation** for quality, and **structured output** that maintains relationships—all while respecting context constraints through architectural chunking rather than content loss.