# Token-in-Token-Out (TITO) in Miles: How Incremental Tokenization Works Across Model Architectures

> Discover Token-in-Token-Out (TITO) in Miles. Learn how this incremental tokenization unifies chat histories across model architectures by merging prefix tokens with new messages.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: deep-dive
- Published: 2026-09-06

---

**Token-in-Token-Out (TITO) is Miles' unified mechanism for incremental tokenization that avoids re-tokenizing entire chat histories by merging pre-computed prefix tokens with newly appended messages.**

TITO powers efficient inference in the [radixark/miles](https://github.com/radixark/miles) codebase by eliminating redundant computation across multi-turn conversations. Instead of re-rendering and re-tokenizing the full chat template every turn, Miles maintains a **pre-tokenized prefix** and only processes incremental changes.

## Core TITO Architecture

The `TITOTokenizer` base class in [[`miles/utils/chat_template_utils/tito_tokenizer.py`](https://github.com/radixark/miles/blob/main/miles/utils/chat_template_utils/tito_tokenizer.py)](https://github.com/radixark/miles/blob/main/miles/utils/chat_template_utils/tito_tokenizer.py#L86‑L106) defines three essential operations:

| Method | Purpose |
|--------|---------|
| **`tokenize_additional_messages`** | Computes token IDs for messages after the stored prefix |
| **`merge_tokens`** | Concatenates prefix tokens with incremental tokens, handling model-specific boundary quirks |
| **`postprocess_completion`** / **`preserve_server_message_state`** | Hooks for adjusting assistant messages or restoring server-side state |

The default `tokenize_additional_messages` implementation renders old-plus-new chat using the model's chat template, then slices the rendered text to obtain only the incremental tokens ([source](https://github.com/radixark/miles/blob/main/miles/utils/chat_template_utils/tito_tokenizer.py#L22‑L34)).

## Token Boundary Handling Across Model Families

Each model architecture in Miles registers a **family** that inherits from `TITOTokenizer`. Families override `merge_tokens` (and occasionally `tokenize_additional_messages`) to fix token-boundary mismatches specific to their chat templates.

### Qwen 3: Newline Injection at `im_end` Boundaries

The Qwen 3 chat template emits `<|im_end|\n>`, but the model stops generation at `im_end` without the trailing newline. The `Qwen3TITOTokenizer` detects when the prefix ends with `im_end` and injects the missing newline during merge ([source](https://github.com/radixark/miles/blob/main/miles/utils/chat_template_utils/tito_tokenizer.py#L13‑L16)).

### GLM 4.7: Ambiguous Token Overlap Resolution

GLM 4.7 uses `<|user|>` and `<|observation|>` as both **stop** and **start** tokens, creating boundary ambiguity. The `GLM47TITOTokenizer` strips the ambiguous trailing token before appending new tokens to prevent duplication ([source](https://github.com/radixark/miles/blob/main/miles/utils/chat_template_utils/tito_tokenizer.py#L24‑L28)).

### Nemotron 3 and Kimi K2.5/K2.6: Shared Patterns with Customization

- **Nemotron 3** inherits Qwen 3's newline handling unchanged via `Nemotron3TITOTokenizer`
- **Kimi K2.5/K2.6** extends this pattern with a custom `special_token_ids` set for segment separators while retaining the core newline logic

### MiniMax M2.5/M2.7: Custom EOS Token Boundaries

These models end with a custom EOS token (`[e~[`) followed by a newline. The `MinimaxM25TITOTokenizer` and `MinimaxM27TITOTokenizer` override `merge_tokens` to append the newline after this EOS token.

### DeepSeek V3.2 and V4: Bridge-Based Tokenization

| Version | Implementation |
|---------|----------------|
| **V3.2** | Uses vendored `encoding_dsv32` bridge; no special `merge_tokens` needed since the bridge renders append-only strings |
| **V4** | Extends V3.2 with additional `merge_tokens` override to trim trailing reasoning brackets |

## How TITO Integrates with Inference

Miles' TITO system couples tightly with **sglang parser bindings** including `reasoning_parser` and `tool_call_parser`. Each family supplies:

- A **fixed chat template** (or native Hugging Face template fallback)
- **Boundary handling** for that model's special tokens
- Parser bindings for tool calling and reasoning extraction

This design ensures that incremental tokenization remains correct even when models add new special tokens or change template formatting between versions.

## Summary

- **TITO eliminates redundant tokenization** by caching prefix tokens and only processing new messages
- **Model families customize boundary handling** through targeted `merge_tokens` overrides
- **Three core methods** (`tokenize_additional_messages`, `merge_tokens`, `postprocess_completion`) provide extension points for any architecture
- **DeepSeek architectures** use bridge-based tokenization, bypassing typical template rendering entirely

## Frequently Asked Questions

### What problem does Token-in-Token-Out solve in LLM inference?

TITO solves the **quadratic cost of chat template rendering** in multi-turn conversations. Without incremental tokenization, every new message requires re-rendering and re-tokenizing the entire conversation history. TITO reduces this to processing only the delta, cutting both compute time and memory pressure for long contexts.

### How does Miles handle models with identical token boundary quirks?

Miles uses **inheritance chains** across model families. For example, `Nemotron3TITOTokenizer` inherits directly from `Qwen3TITOTokenizer` without overrides, while `Kimi25TITOTokenizer` extends the same base with additional `special_token_ids` configuration. This avoids code duplication while allowing fine-grained customization.

### Can TITO work with models that don't use Jinja chat templates?

Yes. `DeepSeekV32TITOTokenizer` demonstrates this approach: it uses a vendored `encoding_dsv32` bridge for tokenization rather than template rendering. The base `merge_tokens` implementation works because the bridge already produces append-only strings, eliminating the boundary-matching problem entirely.

### Why does `merge_tokens` require per-model overrides instead of a universal solution?

Special tokens vary dramatically across architectures—some require newlines, others strip ambiguous boundaries, and some use custom EOS markers. A universal heuristic would fail on edge cases like GLM 4.7's dual-purpose `<|user|>` token or MiniMax's `[e~[` EOS sequence. Explicit per-family overrides ensure correctness without over-generalization.