How the ICN-Engine Scheduler Handles Speculative Decoding in Magnitude
The ICN-Engine scheduler orchestrates speculative decoding by translating logical prompt boundaries into draft-model positions via PromptBoundary, caching draft KV snapshots in PromptCheckpointState::Speculative, and reusing them through SequencePool::acquire_matching to eliminate redundant prefills.
The icn-engine crate in the Magnitude inference stack manages the lifecycle of llama-cpp sequences and provides the scheduling primitives required for speculative decoding. In inference/crates/icn-engine/src/scheduler.rs, the scheduler cleanly separates target and draft KV spaces while enabling the batch planner to interleave decode and prefill operations. This design lets a fast draft model run in parallel with the target model without recomputing shared prompt state.
Core Abstractions for Speculative Decoding
The scheduler revolves around three core structures that bridge the logical prompt space with the llama-cpp speculative API.
PromptBoundary and SpeculativePosition
A PromptBoundary tracks both the logical token count of a prompt and the native KV position of the target model. In inference/crates/icn-engine/src/scheduler.rs (lines 7–20), the method PromptBoundary::speculative_position constructs a SpeculativePosition that the llama-cpp speculative API consumes to align draft and target tensors. This mapping is the first step in enabling the draft model to resume generation from a cached state.
PromptCheckpointState: Target vs. Draft KV
The PromptCheckpointState enum discriminates between normal target inference and speculative draft inference. It defines two variants: Target(LlamaSequenceState) for the standard KV snapshot and Speculative(SpeculativePromptState) for the draft model’s KV state (lines 75–80). A PromptCheckpoint pairs one of these states with the PromptBoundary where the checkpoint occurs, letting the scheduler splice the correct KV slice into a new sequence.
ReusablePrefix for Cross-Request Cache Reuse
A ReusablePrefix stores a cached PromptLayout alongside its associated checkpoints, including speculative ones (lines 88–91). When a subsequent request shares a prefix with a cached layout, the scheduler attaches the stored checkpoints directly instead of re-running a full KV prefill. This abstraction is what makes multi-request speculative decoding efficient in production.
Prefix Matching and Media-Aware Boundaries
Because prompts can contain indivisible media spans, the scheduler enforces strict boundary rules before reusing any cached state.
Finding Shared Prefixes with common_prefix
The method PromptLayout::common_prefix computes the longest shared prefix between two prompt layouts while respecting media boundaries (lines 36–78). Media segments are treated as atomic units; the algorithm stops at the first mismatched media token, ensuring that a speculative draft cannot split a media span. The resulting PromptBoundary can then be converted into a SpeculativePosition for the draft model.
Enforcing Media Boundaries with boundary_at_or_after
The helper PromptLayout::boundary_at_or_after locates the first legal boundary for a given logical token index (lines 111–148). If the index falls inside a media span, the boundary is advanced to the end of that span. This guarantee is essential for speculative decoding because the draft model may use a different tokenization granularity than the target model, and partial media reuse would corrupt the KV cache.
Sequence Pool and Checkpoint Reuse
Before scheduling decode work, the engine attempts to find a reusable sequence in the pool.
Selecting Cached Sequences with acquire_matching
SequencePool::acquire_matching scans the pool for an existing sequence whose cached ReusablePrefix exceeds a similarity threshold (≥ 0.1) with the incoming prompt (lines 70–91). When a match is found, the scheduler attaches the cached checkpoints—including those wrapped in PromptCheckpointState::Speculative—to the new sequence. This bypasses prefill for the shared portion and immediately sets up the draft model for speculative token generation.
End-to-End Scheduling Flow
The full lifecycle from request arrival to speculative decode follows five stages:
- Prompt construction — The request builds a
PromptLayoutcontaining text and optional media segments. - Cache lookup —
SequencePool::acquire_matchingsearches for an existing sequence with a sufficiently similar cached prefix. - Boundary calculation — The shared prefix becomes a
PromptBoundary. Ifspeculative_positionreturnsSome(SpeculativePosition), the scheduler knows exactly where to splice the draft KV. - Checkpoint creation — A
PromptCheckpointis generated with theSpeculativevariant and stored inside theReusablePrefix. - Batch planning — The
BatchPlannerschedules decode work first, ensuring the speculative model’s next token is generated, followed by any remaining prefill tokens. This ordering preserves deterministic generation while keeping target and draft models in lock-step.
By keeping speculative state in PromptCheckpointState and exposing it through PromptBoundary::speculative_position, the scheduler maintains a clean separation between target and draft KV spaces.
Rust Implementation Example
The following snippet demonstrates how a prompt layout is constructed, how a shared boundary is computed, and how it is promoted to a speculative position for reuse:
use llama_cpp_2::{
token::LlamaToken,
speculative::{SpeculativePosition, SpeculativePreflightParams},
};
use magnitude::inference::icn_engine::scheduler::{
PromptLayout, PromptBoundary, PromptCheckpointState, ReusablePrefix,
};
/// Build a prompt that contains text followed by a media span (e.g. an image for M‑RoPE).
let prompt = PromptLayout::new(vec![
PromptSegment::Text(vec![LlamaToken::new(1), LlamaToken::new(2)]),
PromptSegment::Media {
identity: "image‑a".into(),
logical_tokens: 576, // number of logical tokens the media expands to
native_positions: 1, // position offset in the native KV
},
PromptSegment::Text(vec![LlamaToken::new(3)]),
]);
/// Compute the shared prefix with a previously‑cached prompt.
let cached = /* previously cached PromptLayout */;
let shared_boundary: PromptBoundary = cached.common_prefix(&prompt);
/// Turn the boundary into a speculative position for the draft model.
if let Some(spec_pos) = shared_boundary.speculative_position() {
// `spec_pos` can now be passed to the llama-cpp speculative API.
println!("Speculative target={}; draft={}", spec_pos.target, spec_pos.draft);
}
/// When re‑using a sequence, attach its cached checkpoints (including speculative ones).
let reusable = ReusablePrefix {
layout: cached,
checkpoints: vec![
PromptCheckpoint {
state: PromptCheckpointState::Speculative(/* draft KV state */),
boundary: shared_boundary,
},
],
};
This example mirrors the logic found in inference/crates/icn-engine/src/scheduler.rs and shows the exact integration point with the llama-cpp speculative API.
Supporting Crates and Files
Several neighboring crates complete the speculative decoding pipeline:
inference/crates/icn-speculative/src/lib.rs— Provides low-level speculative pre-flight and validation helpers consumed by the scheduler when constructingSpeculativePosition.inference/crates/icn-models/src/package_service.rs— Generates deterministic bundle keys for speculative model bundles, ensuring the draft and target models are correctly paired at load time.inference/crates/icn-models/src/catalog.rs— DeclaresCatalogSpeculativeDecodingand ties speculative configuration parameters into the model catalog.
Summary
- The
PromptBoundaryabstraction bridges logical prompt tokens and native KV positions, exposingspeculative_positionfor the llama-cpp draft API. PromptCheckpointStatecleanly separates target KV snapshots from draft KV snapshots via theTargetandSpeculativevariants.PromptLayout::common_prefixandboundary_at_or_afterenforce media-aware boundary rules so that indivisible media spans are never split during reuse.SequencePool::acquire_matchingenables sub-linear prefix reuse across requests by attaching cached speculative checkpoints when similarity exceeds the ≥ 0.1 threshold.- The batch planner interleaves decode and prefill work so that the draft model remains one step ahead of the target model without redundant computation.
Frequently Asked Questions
How does PromptBoundary enable speculative decoding?
PromptBoundary tracks the logical token count and native KV offset of a shared prompt prefix. Its speculative_position method translates that boundary into a SpeculativePosition struct consumed by the llama-cpp speculative API, allowing the draft model to resume generation from a known KV index without re-running prefill.
Why must media spans remain indivisible during prefix matching?
Media tokens in layouts like M-RoPE represent contiguous embeddings that cannot be split across requests. PromptLayout::common_prefix stops prefix computation at the first mismatched media token, and boundary_at_or_after rounds any interior boundary up to the end of the media span. This prevents the draft model—which may tokenize media differently—from corrupting the target KV cache.
What happens when SequencePool::acquire_matching finds a reusable prefix?
When the pool locates a cached sequence whose ReusablePrefix shares at least 0.1 similarity with the incoming prompt, the scheduler splices the stored checkpoints—including those in PromptCheckpointState::Speculative—into the new sequence. This bypasses KV prefill for the matched portion and immediately prepares the draft model for speculative token generation.
Where is the speculative draft model paired with the target model?
The pairing is handled in inference/crates/icn-models/src/package_service.rs, which generates deterministic bundle keys for draft-target bundles, and in inference/crates/icn-models/src/catalog.rs, where CatalogSpeculativeDecoding declares the speculative configuration. These catalog entries ensure the scheduler loads the correct draft weights alongside the target model.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →