Token-in-Token-Out (TITO) in Miles: How Incremental Tokenization Works Across Model Architectures

Token-in-Token-Out (TITO) is Miles' unified mechanism for incremental tokenization that avoids re-tokenizing entire chat histories by merging pre-computed prefix tokens with newly appended messages.

TITO powers efficient inference in the radixark/miles codebase by eliminating redundant computation across multi-turn conversations. Instead of re-rendering and re-tokenizing the full chat template every turn, Miles maintains a pre-tokenized prefix and only processes incremental changes.

Core TITO Architecture

The TITOTokenizer base class in [miles/utils/chat_template_utils/tito_tokenizer.py](https://github.com/radixark/miles/blob/main/miles/utils/chat_template_utils/tito_tokenizer.py#L86‑L106) defines three essential operations:

Method Purpose
tokenize_additional_messages Computes token IDs for messages after the stored prefix
merge_tokens Concatenates prefix tokens with incremental tokens, handling model-specific boundary quirks
postprocess_completion / preserve_server_message_state Hooks for adjusting assistant messages or restoring server-side state

The default tokenize_additional_messages implementation renders old-plus-new chat using the model's chat template, then slices the rendered text to obtain only the incremental tokens (source).

Token Boundary Handling Across Model Families

Each model architecture in Miles registers a family that inherits from TITOTokenizer. Families override merge_tokens (and occasionally tokenize_additional_messages) to fix token-boundary mismatches specific to their chat templates.

Qwen 3: Newline Injection at im_end Boundaries

The Qwen 3 chat template emits <|im_end|\n>, but the model stops generation at im_end without the trailing newline. The Qwen3TITOTokenizer detects when the prefix ends with im_end and injects the missing newline during merge (source).

GLM 4.7: Ambiguous Token Overlap Resolution

GLM 4.7 uses <|user|> and <|observation|> as both stop and start tokens, creating boundary ambiguity. The GLM47TITOTokenizer strips the ambiguous trailing token before appending new tokens to prevent duplication (source).

Nemotron 3 and Kimi K2.5/K2.6: Shared Patterns with Customization

  • Nemotron 3 inherits Qwen 3's newline handling unchanged via Nemotron3TITOTokenizer
  • Kimi K2.5/K2.6 extends this pattern with a custom special_token_ids set for segment separators while retaining the core newline logic

MiniMax M2.5/M2.7: Custom EOS Token Boundaries

These models end with a custom EOS token ([e~[) followed by a newline. The MinimaxM25TITOTokenizer and MinimaxM27TITOTokenizer override merge_tokens to append the newline after this EOS token.

DeepSeek V3.2 and V4: Bridge-Based Tokenization

Version Implementation
V3.2 Uses vendored encoding_dsv32 bridge; no special merge_tokens needed since the bridge renders append-only strings
V4 Extends V3.2 with additional merge_tokens override to trim trailing reasoning brackets

How TITO Integrates with Inference

Miles' TITO system couples tightly with sglang parser bindings including reasoning_parser and tool_call_parser. Each family supplies:

  • A fixed chat template (or native Hugging Face template fallback)
  • Boundary handling for that model's special tokens
  • Parser bindings for tool calling and reasoning extraction

This design ensures that incremental tokenization remains correct even when models add new special tokens or change template formatting between versions.

Summary

  • TITO eliminates redundant tokenization by caching prefix tokens and only processing new messages
  • Model families customize boundary handling through targeted merge_tokens overrides
  • Three core methods (tokenize_additional_messages, merge_tokens, postprocess_completion) provide extension points for any architecture
  • DeepSeek architectures use bridge-based tokenization, bypassing typical template rendering entirely

Frequently Asked Questions

What problem does Token-in-Token-Out solve in LLM inference?

TITO solves the quadratic cost of chat template rendering in multi-turn conversations. Without incremental tokenization, every new message requires re-rendering and re-tokenizing the entire conversation history. TITO reduces this to processing only the delta, cutting both compute time and memory pressure for long contexts.

How does Miles handle models with identical token boundary quirks?

Miles uses inheritance chains across model families. For example, Nemotron3TITOTokenizer inherits directly from Qwen3TITOTokenizer without overrides, while Kimi25TITOTokenizer extends the same base with additional special_token_ids configuration. This avoids code duplication while allowing fine-grained customization.

Can TITO work with models that don't use Jinja chat templates?

Yes. DeepSeekV32TITOTokenizer demonstrates this approach: it uses a vendored encoding_dsv32 bridge for tokenization rather than template rendering. The base merge_tokens implementation works because the bridge already produces append-only strings, eliminating the boundary-matching problem entirely.

Why does merge_tokens require per-model overrides instead of a universal solution?

Special tokens vary dramatically across architectures—some require newlines, others strip ambiguous boundaries, and some use custom EOS markers. A universal heuristic would fail on edge cases like GLM 4.7's dual-purpose <|user|> token or MiniMax's [e~[ EOS sequence. Explicit per-family overrides ensure correctness without over-generalization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →