How to Handle Long Texts Exceeding Single Sub-Agent Context Limits in Cangjie-Skill
Chunk oversized inputs at natural boundaries (chapters, volumes, parts), keep each slice under ≈50,000 characters, and inject the global BOOK_OVERVIEW.md into every chunk to maintain context consistency.
The kangarooking/cangjie-skill repository processes large documents through specialized sub-agents, each with finite context windows. When source material surpasses what a single sub-agent can ingest alongside the mandatory global overview, the system implements a deliberate chunking strategy rather than truncation or summarization. This preserves semantic integrity while keeping every agent operation within token limits.
The Core Problem: Context Constraints in Sub-Agent Architecture
Sub-agents in Cangjie-Skill operate as isolated workers. Each invocation must accommodate both the global BOOK_OVERVIEW.md (shared project context) and the actual content being processed. For multi-volume collections, long transcripts, or extensive source books, this creates a hard capacity boundary.
The repository addresses this through six coordinated steps documented in methodology/02-stage1-parallel-extract.md:
Step-by-Step Chunking Strategy
1. Detect Oversized Input
Before processing begins, the system evaluates whether the combined size of BOOK_OVERVIEW.md plus the full source text exceeds safe operational limits. If so, the pipeline flags the input for segmentation rather than attempting single-pass processing.
2. Chunk by Natural Boundaries
The text splits at semantic break points rather than arbitrary character positions:
- Chapter headings (第X章)
- Volume markers (卷X)
- Part divisions (Part, Section, or "P" segments in video transcripts)
This preserves narrative and logical coherence within each chunk, ensuring extractor sub-agents operate on self-contained units.
3. Enforce Size Constraints
Each chunk must satisfy:
chunk_size + BOOK_OVERVIEW_size ≤ sub_agent_context_limit
The repository establishes a practical upper bound of ≈50,000 characters (~5万字) per chunk as a conservative safeguard. This accounts for varying tokenization rates across different encodings and content types.
4. Inject Consistent Context
Every chunk receives identical header content: the complete BOOK_OVERVIEW.md file. This guarantees:
- No sub-agent loses awareness of the broader project scope
- Cross-chunk references remain interpretable
- Downstream merging produces coherent, non-contradictory results
5. Execute Parallel Extraction
Five extractor sub-agents process each chunk independently:
- Principle extractor: Core concepts and rules
- Glossary extractor: Terminology definitions
- Framework extractor: Structural models and relationships
- Counter-example extractor: Exception cases and edge conditions
- Case extractor: Concrete illustrations and applications
Parallel execution accelerates throughput for large documents.
6. Maintain Serial Fallback
If the runtime environment cannot spawn parallel processes, the identical chunk list processes sequentially without format changes. This ensures deterministic, reproducible behavior across different deployment contexts.
Implementation Example
Below is a reference Python implementation mirroring the methodology from methodology/02-stage1-parallel-extract.md. It demonstrates boundary-aware chunking with the repository's size constraints.
import os
from pathlib import Path
MAX_CHUNK_SIZE = 50_000 # characters, roughly 5万字
def load_overview(book_dir: Path) -> str:
"""Read the global BOOK_OVERVIEW.md that all sub-agents need."""
return (book_dir / "BOOK_OVERVIEW.md").read_text(encoding="utf-8")
def chunk_by_boundary(text: str) -> list[str]:
"""
Split a large text into natural chunks.
Heuristics target chapter headings (e.g. '第X章') or
volume markers (e.g. '卷X'), falling back to line length limits.
"""
chunks = []
current = []
for line in text.splitlines(keepends=True):
# Size guard: start new chunk if limit would be exceeded
if sum(len(l) for l in current) + len(line) > MAX_CHUNK_SIZE:
chunks.append("".join(current))
current = []
current.append(line)
# Natural boundary detection
if line.lstrip().startswith(("第", "卷", "Part", "Section")):
if len(current) > 1: # Avoid empty chunks
chunks.append("".join(current))
current = []
if current:
chunks.append("".join(current))
return chunks
def prepare_subagent_inputs(book_dir: Path):
overview = load_overview(book_dir)
full_text = (book_dir / "source.txt").read_text(encoding="utf-8")
chunks = chunk_by_boundary(full_text)
input_dir = book_dir / "subagent_inputs"
input_dir.mkdir(exist_ok=True)
for i, chunk in enumerate(chunks, start=1):
payload = f"{overview}\n\n---\n\n{chunk}"
(input_dir / f"chunk_{i:03}.md").write_text(payload, encoding="utf-8")
return input_dir
# Example invocation
book_path = Path("/path/to/books/example_book")
prepare_subagent_inputs(book_path)
Key implementation details aligned with source methodology:
| Aspect | Implementation | Source Reference |
|---|---|---|
| Size limit | MAX_CHUNK_SIZE = 50_000 |
methodology/02-stage1-parallel-extract.md "单块 ≤5 万字" |
| Boundary detection | Chapter/volume prefix matching | "按章节/卷/分P等自然边界切块" |
| Context injection | Overview prepended to every chunk | "每块配同一份 BOOK_OVERVIEW.md" |
| Output format | Numbered .md files in dedicated directory |
Stage 1 input specification |
Critical Design Considerations
Keeping BOOK_OVERVIEW.md Concise
As noted in methodology/07-stage5-deliver.md, the global overview must itself remain compact because it replicates into every sub-agent invocation. An bloated overview directly reduces available capacity for actual content chunks.
Testing Chunked Processing
The methodology/06-stage4-pressure-test.md specification defines validation criteria for sub-agent outputs. When testing chunked workflows, verify that:
- Cross-chunk entity references resolve consistently
- Merged partial results contain no duplication or contradiction
- Edge chunks (first and last) receive identical processing quality as middle chunks
Merging Partial Results
Downstream stages combine extractor outputs from all chunks into unified deliverables. The repository assumes chunk boundaries align with logical content divisions, making automated merging feasible without deep semantic reconciliation.
Summary
- Detect oversized inputs early by measuring against the 50,000-character practical limit
- Split at natural boundaries (chapters, volumes, parts) rather than arbitrary positions
- Inject identical
BOOK_OVERVIEW.mdcontext into every chunk to maintain coherence - Process chunks through all five extractors in parallel when possible, serially when necessary
- Validate that partial results merge cleanly into final deliverables
This methodology, formalized in methodology/02-stage1-parallel-extract.md, enables Cangjie-Skill to handle works of arbitrary length without modifying sub-agent architecture or exceeding context limits.
Frequently Asked Questions
What happens if a single chapter exceeds 50,000 characters?
The size constraint takes precedence: the chapter splits at the nearest natural sub-boundary (section, paragraph break) before the limit. If no sub-boundary exists, the split occurs at the character boundary with a line-break insertion. The BOOK_OVERVIEW.md injection ensures the sub-agent still understands chapter-wide context.
Does chunking affect the five extractor sub-agents differently?
No. All five extractors (principle, glossary, framework, counter-example, case) receive identically chunked inputs. Each runs independently per chunk, producing partial outputs that merge downstream. The methodology intentionally keeps extraction logic agnostic to chunking.
Can I adjust the 50,000-character limit for different models?
Yes, though the repository documents 50,000 as a conservative default. Adjust MAX_CHUNK_SIZE based on your sub-agent's actual context window, accounting for BOOK_OVERVIEW.md size and tokenization overhead. Verify through methodology/06-stage4-pressure-test.md testing that outputs remain complete.
How does the system handle dependencies between distant chapters?
The BOOK_OVERVIEW.md file must explicitly capture cross-cutting relationships, character arcs, or conceptual dependencies. Since sub-agents only see their chunk plus this overview, complex inter-chunk references rely on accurate upfront summarization in the global context file.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →