# How the ICN-Engine Scheduler Handles Speculative Decoding in Magnitude

> Learn how the ICN-Engine scheduler in magnitudedev/magnitude handles speculative decoding. Discover efficient KV caching and sequence reuse to eliminate redundant prefills and boost performance.

- Repository: [Magnitude/magnitude](https://github.com/magnitudedev/magnitude)
- Tags: internals
- Published: 2026-09-08

---

**The ICN-Engine scheduler orchestrates speculative decoding by translating logical prompt boundaries into draft-model positions via `PromptBoundary`, caching draft KV snapshots in `PromptCheckpointState::Speculative`, and reusing them through `SequencePool::acquire_matching` to eliminate redundant prefills.**

The `icn-engine` crate in the Magnitude inference stack manages the lifecycle of llama-cpp sequences and provides the scheduling primitives required for speculative decoding. In [`inference/crates/icn-engine/src/scheduler.rs`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs), the scheduler cleanly separates target and draft KV spaces while enabling the batch planner to interleave decode and prefill operations. This design lets a fast draft model run in parallel with the target model without recomputing shared prompt state.

## Core Abstractions for Speculative Decoding

The scheduler revolves around three core structures that bridge the logical prompt space with the llama-cpp speculative API.

### PromptBoundary and SpeculativePosition

A **`PromptBoundary`** tracks both the *logical* token count of a prompt and the *native* KV position of the target model. In [`inference/crates/icn-engine/src/scheduler.rs`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs) (lines 7–20), the method `PromptBoundary::speculative_position` constructs a **`SpeculativePosition`** that the llama-cpp speculative API consumes to align draft and target tensors. This mapping is the first step in enabling the draft model to resume generation from a cached state.

### PromptCheckpointState: Target vs. Draft KV

The **`PromptCheckpointState`** enum discriminates between normal target inference and speculative draft inference. It defines two variants: `Target(LlamaSequenceState)` for the standard KV snapshot and `Speculative(SpeculativePromptState)` for the draft model’s KV state (lines 75–80). A **`PromptCheckpoint`** pairs one of these states with the **`PromptBoundary`** where the checkpoint occurs, letting the scheduler splice the correct KV slice into a new sequence.

### ReusablePrefix for Cross-Request Cache Reuse

A **`ReusablePrefix`** stores a cached **`PromptLayout`** alongside its associated checkpoints, including speculative ones (lines 88–91). When a subsequent request shares a prefix with a cached layout, the scheduler attaches the stored checkpoints directly instead of re-running a full KV prefill. This abstraction is what makes multi-request speculative decoding efficient in production.

## Prefix Matching and Media-Aware Boundaries

Because prompts can contain indivisible media spans, the scheduler enforces strict boundary rules before reusing any cached state.

### Finding Shared Prefixes with common_prefix

The method **`PromptLayout::common_prefix`** computes the longest shared prefix between two prompt layouts while respecting media boundaries (lines 36–78). Media segments are treated as atomic units; the algorithm stops at the first mismatched media token, ensuring that a speculative draft cannot split a media span. The resulting **`PromptBoundary`** can then be converted into a **`SpeculativePosition`** for the draft model.

### Enforcing Media Boundaries with boundary_at_or_after

The helper **`PromptLayout::boundary_at_or_after`** locates the first legal boundary for a given logical token index (lines 111–148). If the index falls inside a media span, the boundary is advanced to the end of that span. This guarantee is essential for speculative decoding because the draft model may use a different tokenization granularity than the target model, and partial media reuse would corrupt the KV cache.

## Sequence Pool and Checkpoint Reuse

Before scheduling decode work, the engine attempts to find a reusable sequence in the pool.

### Selecting Cached Sequences with acquire_matching

**`SequencePool::acquire_matching`** scans the pool for an existing sequence whose cached **`ReusablePrefix`** exceeds a similarity threshold (≥ 0.1) with the incoming prompt (lines 70–91). When a match is found, the scheduler attaches the cached checkpoints—including those wrapped in `PromptCheckpointState::Speculative`—to the new sequence. This bypasses prefill for the shared portion and immediately sets up the draft model for speculative token generation.

## End-to-End Scheduling Flow

The full lifecycle from request arrival to speculative decode follows five stages:

1. **Prompt construction** — The request builds a **`PromptLayout`** containing text and optional media segments.
2. **Cache lookup** — `SequencePool::acquire_matching` searches for an existing sequence with a sufficiently similar cached prefix.
3. **Boundary calculation** — The shared prefix becomes a **`PromptBoundary`**. If `speculative_position` returns `Some(SpeculativePosition)`, the scheduler knows exactly where to splice the draft KV.
4. **Checkpoint creation** — A **`PromptCheckpoint`** is generated with the `Speculative` variant and stored inside the **`ReusablePrefix`**.
5. **Batch planning** — The **`BatchPlanner`** schedules decode work first, ensuring the speculative model’s next token is generated, followed by any remaining prefill tokens. This ordering preserves deterministic generation while keeping target and draft models in lock-step.

By keeping speculative state in `PromptCheckpointState` and exposing it through `PromptBoundary::speculative_position`, the scheduler maintains a clean separation between target and draft KV spaces.

## Rust Implementation Example

The following snippet demonstrates how a prompt layout is constructed, how a shared boundary is computed, and how it is promoted to a speculative position for reuse:

```rust
use llama_cpp_2::{
    token::LlamaToken,
    speculative::{SpeculativePosition, SpeculativePreflightParams},
};
use magnitude::inference::icn_engine::scheduler::{
    PromptLayout, PromptBoundary, PromptCheckpointState, ReusablePrefix,
};

/// Build a prompt that contains text followed by a media span (e.g. an image for M‑RoPE).
let prompt = PromptLayout::new(vec![
    PromptSegment::Text(vec![LlamaToken::new(1), LlamaToken::new(2)]),
    PromptSegment::Media {
        identity: "image‑a".into(),
        logical_tokens: 576,          // number of logical tokens the media expands to
        native_positions: 1,          // position offset in the native KV
    },
    PromptSegment::Text(vec![LlamaToken::new(3)]),
]);

/// Compute the shared prefix with a previously‑cached prompt.
let cached = /* previously cached PromptLayout */;
let shared_boundary: PromptBoundary = cached.common_prefix(&prompt);

/// Turn the boundary into a speculative position for the draft model.
if let Some(spec_pos) = shared_boundary.speculative_position() {
    // `spec_pos` can now be passed to the llama-cpp speculative API.
    println!("Speculative target={}; draft={}", spec_pos.target, spec_pos.draft);
}

/// When re‑using a sequence, attach its cached checkpoints (including speculative ones).
let reusable = ReusablePrefix {
    layout: cached,
    checkpoints: vec![
        PromptCheckpoint {
            state: PromptCheckpointState::Speculative(/* draft KV state */),
            boundary: shared_boundary,
        },
    ],
};

```

This example mirrors the logic found in [`inference/crates/icn-engine/src/scheduler.rs`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs) and shows the exact integration point with the llama-cpp speculative API.

## Supporting Crates and Files

Several neighboring crates complete the speculative decoding pipeline:

- **[`inference/crates/icn-speculative/src/lib.rs`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-speculative/src/lib.rs)** — Provides low-level speculative pre-flight and validation helpers consumed by the scheduler when constructing `SpeculativePosition`.
- **[`inference/crates/icn-models/src/package_service.rs`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-models/src/package_service.rs)** — Generates deterministic bundle keys for speculative model bundles, ensuring the draft and target models are correctly paired at load time.
- **[`inference/crates/icn-models/src/catalog.rs`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-models/src/catalog.rs)** — Declares `CatalogSpeculativeDecoding` and ties speculative configuration parameters into the model catalog.

## Summary

- The **`PromptBoundary`** abstraction bridges logical prompt tokens and native KV positions, exposing `speculative_position` for the llama-cpp draft API.
- **`PromptCheckpointState`** cleanly separates target KV snapshots from draft KV snapshots via the `Target` and `Speculative` variants.
- **`PromptLayout::common_prefix`** and **`boundary_at_or_after`** enforce media-aware boundary rules so that indivisible media spans are never split during reuse.
- **`SequencePool::acquire_matching`** enables sub-linear prefix reuse across requests by attaching cached speculative checkpoints when similarity exceeds the ≥ 0.1 threshold.
- The batch planner interleaves decode and prefill work so that the draft model remains one step ahead of the target model without redundant computation.

## Frequently Asked Questions

### How does `PromptBoundary` enable speculative decoding?

`PromptBoundary` tracks the logical token count and native KV offset of a shared prompt prefix. Its `speculative_position` method translates that boundary into a `SpeculativePosition` struct consumed by the llama-cpp speculative API, allowing the draft model to resume generation from a known KV index without re-running prefill.

### Why must media spans remain indivisible during prefix matching?

Media tokens in layouts like M-RoPE represent contiguous embeddings that cannot be split across requests. `PromptLayout::common_prefix` stops prefix computation at the first mismatched media token, and `boundary_at_or_after` rounds any interior boundary up to the end of the media span. This prevents the draft model—which may tokenize media differently—from corrupting the target KV cache.

### What happens when `SequencePool::acquire_matching` finds a reusable prefix?

When the pool locates a cached sequence whose `ReusablePrefix` shares at least 0.1 similarity with the incoming prompt, the scheduler splices the stored checkpoints—including those in `PromptCheckpointState::Speculative`—into the new sequence. This bypasses KV prefill for the matched portion and immediately prepares the draft model for speculative token generation.

### Where is the speculative draft model paired with the target model?

The pairing is handled in [`inference/crates/icn-models/src/package_service.rs`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-models/src/package_service.rs), which generates deterministic bundle keys for draft-target bundles, and in [`inference/crates/icn-models/src/catalog.rs`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-models/src/catalog.rs), where `CatalogSpeculativeDecoding` declares the speculative configuration. These catalog entries ensure the scheduler loads the correct draft weights alongside the target model.