How the ICN-Engine Scheduler Handles Speculative Decoding in Magnitude

The ICN-Engine scheduler orchestrates speculative decoding by translating logical prompt boundaries into draft-model positions via PromptBoundary, caching draft KV snapshots in PromptCheckpointState::Speculative, and reusing them through SequencePool::acquire_matching to eliminate redundant prefills.

The icn-engine crate in the Magnitude inference stack manages the lifecycle of llama-cpp sequences and provides the scheduling primitives required for speculative decoding. In inference/crates/icn-engine/src/scheduler.rs, the scheduler cleanly separates target and draft KV spaces while enabling the batch planner to interleave decode and prefill operations. This design lets a fast draft model run in parallel with the target model without recomputing shared prompt state.

Core Abstractions for Speculative Decoding

The scheduler revolves around three core structures that bridge the logical prompt space with the llama-cpp speculative API.

PromptBoundary and SpeculativePosition

A PromptBoundary tracks both the logical token count of a prompt and the native KV position of the target model. In inference/crates/icn-engine/src/scheduler.rs (lines 7–20), the method PromptBoundary::speculative_position constructs a SpeculativePosition that the llama-cpp speculative API consumes to align draft and target tensors. This mapping is the first step in enabling the draft model to resume generation from a cached state.

PromptCheckpointState: Target vs. Draft KV

The PromptCheckpointState enum discriminates between normal target inference and speculative draft inference. It defines two variants: Target(LlamaSequenceState) for the standard KV snapshot and Speculative(SpeculativePromptState) for the draft model’s KV state (lines 75–80). A PromptCheckpoint pairs one of these states with the PromptBoundary where the checkpoint occurs, letting the scheduler splice the correct KV slice into a new sequence.

ReusablePrefix for Cross-Request Cache Reuse

A ReusablePrefix stores a cached PromptLayout alongside its associated checkpoints, including speculative ones (lines 88–91). When a subsequent request shares a prefix with a cached layout, the scheduler attaches the stored checkpoints directly instead of re-running a full KV prefill. This abstraction is what makes multi-request speculative decoding efficient in production.

Prefix Matching and Media-Aware Boundaries

Because prompts can contain indivisible media spans, the scheduler enforces strict boundary rules before reusing any cached state.

Finding Shared Prefixes with common_prefix

The method PromptLayout::common_prefix computes the longest shared prefix between two prompt layouts while respecting media boundaries (lines 36–78). Media segments are treated as atomic units; the algorithm stops at the first mismatched media token, ensuring that a speculative draft cannot split a media span. The resulting PromptBoundary can then be converted into a SpeculativePosition for the draft model.

Enforcing Media Boundaries with boundary_at_or_after

The helper PromptLayout::boundary_at_or_after locates the first legal boundary for a given logical token index (lines 111–148). If the index falls inside a media span, the boundary is advanced to the end of that span. This guarantee is essential for speculative decoding because the draft model may use a different tokenization granularity than the target model, and partial media reuse would corrupt the KV cache.

Sequence Pool and Checkpoint Reuse

Before scheduling decode work, the engine attempts to find a reusable sequence in the pool.

Selecting Cached Sequences with acquire_matching

SequencePool::acquire_matching scans the pool for an existing sequence whose cached ReusablePrefix exceeds a similarity threshold (≥ 0.1) with the incoming prompt (lines 70–91). When a match is found, the scheduler attaches the cached checkpoints—including those wrapped in PromptCheckpointState::Speculative—to the new sequence. This bypasses prefill for the shared portion and immediately sets up the draft model for speculative token generation.

End-to-End Scheduling Flow

The full lifecycle from request arrival to speculative decode follows five stages:

  1. Prompt construction — The request builds a PromptLayout containing text and optional media segments.
  2. Cache lookup — SequencePool::acquire_matching searches for an existing sequence with a sufficiently similar cached prefix.
  3. Boundary calculation — The shared prefix becomes a PromptBoundary. If speculative_position returns Some(SpeculativePosition), the scheduler knows exactly where to splice the draft KV.
  4. Checkpoint creation — A PromptCheckpoint is generated with the Speculative variant and stored inside the ReusablePrefix.
  5. Batch planning — The BatchPlanner schedules decode work first, ensuring the speculative model’s next token is generated, followed by any remaining prefill tokens. This ordering preserves deterministic generation while keeping target and draft models in lock-step.

By keeping speculative state in PromptCheckpointState and exposing it through PromptBoundary::speculative_position, the scheduler maintains a clean separation between target and draft KV spaces.

Rust Implementation Example

The following snippet demonstrates how a prompt layout is constructed, how a shared boundary is computed, and how it is promoted to a speculative position for reuse:

use llama_cpp_2::{
    token::LlamaToken,
    speculative::{SpeculativePosition, SpeculativePreflightParams},
};
use magnitude::inference::icn_engine::scheduler::{
    PromptLayout, PromptBoundary, PromptCheckpointState, ReusablePrefix,
};

/// Build a prompt that contains text followed by a media span (e.g. an image for M‑RoPE).
let prompt = PromptLayout::new(vec![
    PromptSegment::Text(vec![LlamaToken::new(1), LlamaToken::new(2)]),
    PromptSegment::Media {
        identity: "image‑a".into(),
        logical_tokens: 576,          // number of logical tokens the media expands to
        native_positions: 1,          // position offset in the native KV
    },
    PromptSegment::Text(vec![LlamaToken::new(3)]),
]);

/// Compute the shared prefix with a previously‑cached prompt.
let cached = /* previously cached PromptLayout */;
let shared_boundary: PromptBoundary = cached.common_prefix(&prompt);

/// Turn the boundary into a speculative position for the draft model.
if let Some(spec_pos) = shared_boundary.speculative_position() {
    // `spec_pos` can now be passed to the llama-cpp speculative API.
    println!("Speculative target={}; draft={}", spec_pos.target, spec_pos.draft);
}

/// When re‑using a sequence, attach its cached checkpoints (including speculative ones).
let reusable = ReusablePrefix {
    layout: cached,
    checkpoints: vec![
        PromptCheckpoint {
            state: PromptCheckpointState::Speculative(/* draft KV state */),
            boundary: shared_boundary,
        },
    ],
};

This example mirrors the logic found in inference/crates/icn-engine/src/scheduler.rs and shows the exact integration point with the llama-cpp speculative API.

Supporting Crates and Files

Several neighboring crates complete the speculative decoding pipeline:

Summary

  • The PromptBoundary abstraction bridges logical prompt tokens and native KV positions, exposing speculative_position for the llama-cpp draft API.
  • PromptCheckpointState cleanly separates target KV snapshots from draft KV snapshots via the Target and Speculative variants.
  • PromptLayout::common_prefix and boundary_at_or_after enforce media-aware boundary rules so that indivisible media spans are never split during reuse.
  • SequencePool::acquire_matching enables sub-linear prefix reuse across requests by attaching cached speculative checkpoints when similarity exceeds the ≥ 0.1 threshold.
  • The batch planner interleaves decode and prefill work so that the draft model remains one step ahead of the target model without redundant computation.

Frequently Asked Questions

How does PromptBoundary enable speculative decoding?

PromptBoundary tracks the logical token count and native KV offset of a shared prompt prefix. Its speculative_position method translates that boundary into a SpeculativePosition struct consumed by the llama-cpp speculative API, allowing the draft model to resume generation from a known KV index without re-running prefill.

Why must media spans remain indivisible during prefix matching?

Media tokens in layouts like M-RoPE represent contiguous embeddings that cannot be split across requests. PromptLayout::common_prefix stops prefix computation at the first mismatched media token, and boundary_at_or_after rounds any interior boundary up to the end of the media span. This prevents the draft model—which may tokenize media differently—from corrupting the target KV cache.

What happens when SequencePool::acquire_matching finds a reusable prefix?

When the pool locates a cached sequence whose ReusablePrefix shares at least 0.1 similarity with the incoming prompt, the scheduler splices the stored checkpoints—including those in PromptCheckpointState::Speculative—into the new sequence. This bypasses KV prefill for the matched portion and immediately prepares the draft model for speculative token generation.

Where is the speculative draft model paired with the target model?

The pairing is handled in inference/crates/icn-models/src/package_service.rs, which generates deterministic bundle keys for draft-target bundles, and in inference/crates/icn-models/src/catalog.rs, where CatalogSpeculativeDecoding declares the speculative configuration. These catalog entries ensure the scheduler loads the correct draft weights alongside the target model.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →