# How to Handle Long Texts Exceeding Single Sub-Agent Context Limits in Cangjie-Skill

> Easily handle long texts exceeding sub-agent context limits in Cangjie Skill. Chunk oversized inputs, maintain context with BOOK_OVERVIEW.md, and keep slices under 50,000 characters.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: how-to-guide
- Published: 2026-08-13

---

**Chunk oversized inputs at natural boundaries (chapters, volumes, parts), keep each slice under ≈50,000 characters, and inject the global [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) into every chunk to maintain context consistency.**

The `kangarooking/cangjie-skill` repository processes large documents through specialized sub-agents, each with finite context windows. When source material surpasses what a single sub-agent can ingest alongside the mandatory global overview, the system implements a **deliberate chunking strategy** rather than truncation or summarization. This preserves semantic integrity while keeping every agent operation within token limits.

## The Core Problem: Context Constraints in Sub-Agent Architecture

Sub-agents in Cangjie-Skill operate as isolated workers. Each invocation must accommodate both the global [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) (shared project context) and the actual content being processed. For multi-volume collections, long transcripts, or extensive source books, this creates a hard capacity boundary.

The repository addresses this through six coordinated steps documented in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md):

## Step-by-Step Chunking Strategy

### 1. Detect Oversized Input

Before processing begins, the system evaluates whether the combined size of [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) plus the full source text exceeds safe operational limits. If so, the pipeline flags the input for segmentation rather than attempting single-pass processing.

### 2. Chunk by Natural Boundaries

The text splits at **semantic break points** rather than arbitrary character positions:

- Chapter headings (第X章)
- Volume markers (卷X)
- Part divisions (Part, Section, or "P" segments in video transcripts)

This preserves narrative and logical coherence within each chunk, ensuring extractor sub-agents operate on self-contained units.

### 3. Enforce Size Constraints

Each chunk must satisfy:

```

chunk_size + BOOK_OVERVIEW_size ≤ sub_agent_context_limit

```

The repository establishes a **practical upper bound of ≈50,000 characters (~5万字)** per chunk as a conservative safeguard. This accounts for varying tokenization rates across different encodings and content types.

### 4. Inject Consistent Context

Every chunk receives identical header content: the complete [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) file. This guarantees:

- No sub-agent loses awareness of the broader project scope
- Cross-chunk references remain interpretable
- Downstream merging produces coherent, non-contradictory results

### 5. Execute Parallel Extraction

Five extractor sub-agents process each chunk independently:

- **Principle extractor**: Core concepts and rules
- **Glossary extractor**: Terminology definitions
- **Framework extractor**: Structural models and relationships
- **Counter-example extractor**: Exception cases and edge conditions
- **Case extractor**: Concrete illustrations and applications

Parallel execution accelerates throughput for large documents.

### 6. Maintain Serial Fallback

If the runtime environment cannot spawn parallel processes, the identical chunk list processes **sequentially** without format changes. This ensures deterministic, reproducible behavior across different deployment contexts.

## Implementation Example

Below is a reference Python implementation mirroring the methodology from [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md). It demonstrates boundary-aware chunking with the repository's size constraints.

```python
import os
from pathlib import Path

MAX_CHUNK_SIZE = 50_000  # characters, roughly 5万字

def load_overview(book_dir: Path) -> str:
    """Read the global BOOK_OVERVIEW.md that all sub-agents need."""
    return (book_dir / "BOOK_OVERVIEW.md").read_text(encoding="utf-8")

def chunk_by_boundary(text: str) -> list[str]:
    """
    Split a large text into natural chunks.
    Heuristics target chapter headings (e.g. '第X章') or
    volume markers (e.g. '卷X'), falling back to line length limits.
    """
    chunks = []
    current = []

    for line in text.splitlines(keepends=True):
        # Size guard: start new chunk if limit would be exceeded

        if sum(len(l) for l in current) + len(line) > MAX_CHUNK_SIZE:
            chunks.append("".join(current))
            current = []

        current.append(line)

        # Natural boundary detection

        if line.lstrip().startswith(("第", "卷", "Part", "Section")):
            if len(current) > 1:  # Avoid empty chunks

                chunks.append("".join(current))
                current = []

    if current:
        chunks.append("".join(current))

    return chunks

def prepare_subagent_inputs(book_dir: Path):
    overview = load_overview(book_dir)
    full_text = (book_dir / "source.txt").read_text(encoding="utf-8")
    chunks = chunk_by_boundary(full_text)

    input_dir = book_dir / "subagent_inputs"
    input_dir.mkdir(exist_ok=True)

    for i, chunk in enumerate(chunks, start=1):
        payload = f"{overview}\n\n---\n\n{chunk}"
        (input_dir / f"chunk_{i:03}.md").write_text(payload, encoding="utf-8")
    
    return input_dir

# Example invocation

book_path = Path("/path/to/books/example_book")
prepare_subagent_inputs(book_path)

```

**Key implementation details aligned with source methodology:**

| Aspect | Implementation | Source Reference |
|--------|---------------|----------------|
| Size limit | `MAX_CHUNK_SIZE = 50_000` | [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md) "单块 ≤5 万字" |
| Boundary detection | Chapter/volume prefix matching | "按章节/卷/分P等自然边界切块" |
| Context injection | Overview prepended to every chunk | "每块配同一份 [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md)" |
| Output format | Numbered `.md` files in dedicated directory | Stage 1 input specification |

## Critical Design Considerations

### Keeping BOOK_OVERVIEW.md Concise

As noted in [`methodology/07-stage5-deliver.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/07-stage5-deliver.md), the global overview must itself remain compact because it replicates into every sub-agent invocation. An bloated overview directly reduces available capacity for actual content chunks.

### Testing Chunked Processing

The [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) specification defines validation criteria for sub-agent outputs. When testing chunked workflows, verify that:

- Cross-chunk entity references resolve consistently
- Merged partial results contain no duplication or contradiction
- Edge chunks (first and last) receive identical processing quality as middle chunks

### Merging Partial Results

Downstream stages combine extractor outputs from all chunks into unified deliverables. The repository assumes chunk boundaries align with logical content divisions, making automated merging feasible without deep semantic reconciliation.

## Summary

- **Detect** oversized inputs early by measuring against the 50,000-character practical limit
- **Split** at natural boundaries (chapters, volumes, parts) rather than arbitrary positions
- **Inject** identical [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) context into every chunk to maintain coherence
- **Process** chunks through all five extractors in parallel when possible, serially when necessary
- **Validate** that partial results merge cleanly into final deliverables

This methodology, formalized in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md), enables Cangjie-Skill to handle works of arbitrary length without modifying sub-agent architecture or exceeding context limits.

## Frequently Asked Questions

### What happens if a single chapter exceeds 50,000 characters?

The size constraint takes precedence: the chapter splits at the nearest natural sub-boundary (section, paragraph break) before the limit. If no sub-boundary exists, the split occurs at the character boundary with a line-break insertion. The [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) injection ensures the sub-agent still understands chapter-wide context.

### Does chunking affect the five extractor sub-agents differently?

No. All five extractors (principle, glossary, framework, counter-example, case) receive identically chunked inputs. Each runs independently per chunk, producing partial outputs that merge downstream. The methodology intentionally keeps extraction logic agnostic to chunking.

### Can I adjust the 50,000-character limit for different models?

Yes, though the repository documents 50,000 as a conservative default. Adjust `MAX_CHUNK_SIZE` based on your sub-agent's actual context window, accounting for [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) size and tokenization overhead. Verify through [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) testing that outputs remain complete.

### How does the system handle dependencies between distant chapters?

The [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) file must explicitly capture cross-cutting relationships, character arcs, or conceptual dependencies. Since sub-agents only see their chunk plus this overview, complex inter-chunk references rely on accurate upfront summarization in the global context file.