Token-in-Token-Out (TITO) in Miles: How Incremental Tokenization Works Across Model Architectures
Token-in-Token-Out (TITO) is Miles' unified mechanism for incremental tokenization that avoids re-tokenizing entire chat histories by merging pre-computed prefix tokens with newly appended messages.
TITO powers efficient inference in the radixark/miles codebase by eliminating redundant computation across multi-turn conversations. Instead of re-rendering and re-tokenizing the full chat template every turn, Miles maintains a pre-tokenized prefix and only processes incremental changes.
Core TITO Architecture
The TITOTokenizer base class in [miles/utils/chat_template_utils/tito_tokenizer.py](https://github.com/radixark/miles/blob/main/miles/utils/chat_template_utils/tito_tokenizer.py#L86‑L106) defines three essential operations:
| Method | Purpose |
|---|---|
tokenize_additional_messages |
Computes token IDs for messages after the stored prefix |
merge_tokens |
Concatenates prefix tokens with incremental tokens, handling model-specific boundary quirks |
postprocess_completion / preserve_server_message_state |
Hooks for adjusting assistant messages or restoring server-side state |
The default tokenize_additional_messages implementation renders old-plus-new chat using the model's chat template, then slices the rendered text to obtain only the incremental tokens (source).
Token Boundary Handling Across Model Families
Each model architecture in Miles registers a family that inherits from TITOTokenizer. Families override merge_tokens (and occasionally tokenize_additional_messages) to fix token-boundary mismatches specific to their chat templates.
Qwen 3: Newline Injection at im_end Boundaries
The Qwen 3 chat template emits <|im_end|\n>, but the model stops generation at im_end without the trailing newline. The Qwen3TITOTokenizer detects when the prefix ends with im_end and injects the missing newline during merge (source).
GLM 4.7: Ambiguous Token Overlap Resolution
GLM 4.7 uses <|user|> and <|observation|> as both stop and start tokens, creating boundary ambiguity. The GLM47TITOTokenizer strips the ambiguous trailing token before appending new tokens to prevent duplication (source).
Nemotron 3 and Kimi K2.5/K2.6: Shared Patterns with Customization
- Nemotron 3 inherits Qwen 3's newline handling unchanged via
Nemotron3TITOTokenizer - Kimi K2.5/K2.6 extends this pattern with a custom
special_token_idsset for segment separators while retaining the core newline logic
MiniMax M2.5/M2.7: Custom EOS Token Boundaries
These models end with a custom EOS token ([e~[) followed by a newline. The MinimaxM25TITOTokenizer and MinimaxM27TITOTokenizer override merge_tokens to append the newline after this EOS token.
DeepSeek V3.2 and V4: Bridge-Based Tokenization
| Version | Implementation |
|---|---|
| V3.2 | Uses vendored encoding_dsv32 bridge; no special merge_tokens needed since the bridge renders append-only strings |
| V4 | Extends V3.2 with additional merge_tokens override to trim trailing reasoning brackets |
How TITO Integrates with Inference
Miles' TITO system couples tightly with sglang parser bindings including reasoning_parser and tool_call_parser. Each family supplies:
- A fixed chat template (or native Hugging Face template fallback)
- Boundary handling for that model's special tokens
- Parser bindings for tool calling and reasoning extraction
This design ensures that incremental tokenization remains correct even when models add new special tokens or change template formatting between versions.
Summary
- TITO eliminates redundant tokenization by caching prefix tokens and only processing new messages
- Model families customize boundary handling through targeted
merge_tokensoverrides - Three core methods (
tokenize_additional_messages,merge_tokens,postprocess_completion) provide extension points for any architecture - DeepSeek architectures use bridge-based tokenization, bypassing typical template rendering entirely
Frequently Asked Questions
What problem does Token-in-Token-Out solve in LLM inference?
TITO solves the quadratic cost of chat template rendering in multi-turn conversations. Without incremental tokenization, every new message requires re-rendering and re-tokenizing the entire conversation history. TITO reduces this to processing only the delta, cutting both compute time and memory pressure for long contexts.
How does Miles handle models with identical token boundary quirks?
Miles uses inheritance chains across model families. For example, Nemotron3TITOTokenizer inherits directly from Qwen3TITOTokenizer without overrides, while Kimi25TITOTokenizer extends the same base with additional special_token_ids configuration. This avoids code duplication while allowing fine-grained customization.
Can TITO work with models that don't use Jinja chat templates?
Yes. DeepSeekV32TITOTokenizer demonstrates this approach: it uses a vendored encoding_dsv32 bridge for tokenization rather than template rendering. The base merge_tokens implementation works because the bridge already produces append-only strings, eliminating the boundary-matching problem entirely.
Why does merge_tokens require per-model overrides instead of a universal solution?
Special tokens vary dramatically across architectures—some require newlines, others strip ambiguous boundaries, and some use custom EOS markers. A universal heuristic would fail on edge cases like GLM 4.7's dual-purpose <|user|> token or MiniMax's [e~[ EOS sequence. Explicit per-family overrides ensure correctness without over-generalization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →