How the icn‑engine Crate Handles Inference Execution and Scheduling in Magnitude

The icn‑engine crate separates inference execution from scheduling: a dedicated executor thread runs native llama.cpp workloads while a lightweight, deterministic Rust scheduler manages prompt caching, sequence pooling, and batch planning with decode‑first ordering.

The icn‑engine crate is the core inference runtime in magnitudedev/magnitude, a Rust‑based AI inference engine. It orchestrates the full model lifecycle—from loading the native llama.cpp backend through scheduling concurrent requests to streaming token events back to SDK consumers. This article examines how inference execution and scheduling work, with direct references to the source implementation.


Architecture Overview: Execution vs. Scheduling

The icn‑engine crate deliberately decouples execution (native GPU/CPU inference) from scheduling (deciding what to run and when). This separation keeps the scheduler policy‑free and fully testable in pure Rust, while the executor handles unsafe native calls.

Component Responsibility Source Location
NativeBackend Process‑wide LlamaBackend instance, hardware discovery, load planning lib.rs#L34-L44
PreparedModelLoad Fully resolved load plan ready for instantiation lib.rs#L84-L115
LlamaCompletionBackend Public CompletionBackend implementation for SDK lib.rs#L862-L870
Executor thread Runs native context, processes commands, emits events lib.rs#L885-L894
Scheduler (scheduler.rs) Prompt layout, sequence pooling, batch planning [scheduler.rs](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs)
BatchPlanner Orders work: decode tokens first, then pre‑fill quanta scheduler.rs#L63-L82
SequencePool Reusable sequence IDs with optional KV cache prefix reuse scheduler.rs#L70-L91

Backend Initialization and Model Loading

Inference execution begins with a one‑time backend initialization. The NativeBackend::initialize method creates a process‑wide LlamaBackend instance wrapped in an Arc, ensuring all executor threads share a single native library instance.

let backend = NativeBackend::initialize()?;
// Source: https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/lib.rs#L34-L44

Load planning happens through prepare_load, which validates configuration, runs hardware‑aware planning via icn_hardware::plan_load_with_backend, and produces a PreparedModelLoad. This prepared plan includes:

  • ModelLoadPhase values describing each loading step
  • A command channel (commands) for the executor thread
  • A start channel for the observer
let intent = execution_intent(model_path, projector_path, &defaults);
let prepared = backend.prepare_load(
    model_id,
    intent,
    speculative_cfg,
    hardware,
    caps,
    profile,
    fingerprint,
    expects_vision
)?;

The Executor Thread: Native Inference Loop

The executor thread is where actual inference execution occurs. Created via thread::Builder, it runs executor_main for the lifetime of the model:

let (commands, command_receiver) = sync_channel(COMMAND_QUEUE_CAPACITY);
let executor = thread::Builder::new()
    .name(format!("icn-llama-{model_id}"))
    .spawn(move || executor_main(
        backend,
        planned,
        acceleration,
        command_receiver,
        // ...
    ));
// Source: https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/lib.rs#L885-L894

The executor_main function:

  1. Creates a LlamaThreadPool for parallel tokenization
  2. Initializes optional speculative decoding session
  3. Loads multimodal runtime if mtmd feature is enabled
  4. Enters run_initialized_executor to process ExecutorCommands and emit ExecutorItems

The thread persists until LlamaCompletionBackend is dropped, ensuring model context stability across many requests.


Scheduling: Prompt Layout and Cache Reuse

The scheduler operates on prompt layouts—semantic representations that track both logical token counts and native KV positions.

PromptBoundary and PromptLayout

pub(crate) struct PromptBoundary {
    pub(crate) logical_tokens: usize,
    pub(crate) native_position: i32,
}
// Source: https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs#L7-L20

A PromptLayout consists of PromptSegments, each knowing its logical_tokens and native_positions. The common_prefix method computes shared prefixes between prompts, enabling cache reuse.

SequencePool: Reusable Cache Prefixes

When a request completes pre‑fill, its ReusablePrefix (prompt layout plus checkpoint states) is retained. For new requests, SequencePool::acquire_matching attempts to find a sequence with sufficient prefix similarity:

let best = self.available.iter()
    .enumerate()
    .filter_map(|(index, sequence)| {
        let prefix = sequence.reusable_prefix.as_ref()?;
        let common_prefix = prefix.layout.common_prefix(prompt).logical_tokens;
        let similarity = common_prefix as f32 / prompt_tokens as f32;
        (similarity > SLOT_PROMPT_SIMILARITY_THRESHOLD).then_some((index, common_prefix))
    })
    .max_by_key(|(_, common_prefix)| *common_prefix)
    .map(|(index, _)| index);
// Source: https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs#L70-L91

The threshold SLOT_PROMPT_SIMILARITY_THRESHOLD is set to 0.1 (10%). Matching sequences skip redundant pre‑fill work, dramatically reducing latency for similar prompts.


Batch Planning: Decode‑First Scheduling

Each scheduling tick produces a vector of BatchWork items through BatchPlanner::plan. The policy is intentionally simple and deterministic:

  1. Decode work first — at most one token per sequence, sorted by sequence ID for stability
  2. Pre‑fill quanta — round‑robin across sequences, limited to prefill_quantum tokens each
// Decode-first: sorted by sequence_id
let mut decode_candidates = candidates.iter()
    .filter(|c| c.kind == WorkKind::Decode)
    .collect::<Vec<_>>();
decode_candidates.sort_unstable_by_key(|c| c.sequence_id);

// Prefill round-robin with cursor-based fairness
let start = self.cursor % prefill_candidates.len();
// ...
while available > 0 { /* distribute quanta */ }
// Source: https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs#L63-L82

This decode‑first approach minimizes latency for active generations while the round‑robin pre‑fill prevents starvation of new requests. The cursor (self.cursor) advances each tick, ensuring fair distribution without external policy dependencies.


Request Lifecycle and Event Streaming

A complete inference request flows through these stages:

  1. Acquire sequence — pool.acquire_matching(&prompt) attempts cache reuse
  2. Tokenize prompt — convert chat to TokenizedPrompt (tokens + layout)
  3. Send command — ExecutorCommand::Complete dispatched to executor thread
  4. Batch loop — repeatedly call BatchPlanner::plan, feed to LlamaBatch, commit via BatchCommit::apply
  5. Emit events — ExecutorItem::Event for each token, ExecutorItem::Completed at termination
  6. Return to pool — sequence released with updated ReusablePrefix

The ActiveRequest struct holds per‑request mutable state including phase, token history, speculative data, and timing measurements. BatchCommit::apply mutates this state and records progress for the next scheduling tick.


Advanced Features: Speculative Decoding and Multimodal

Speculative Decoding

When ExecutionIntent::speculative is enabled, prepare_native_plan creates a SpeculativeSession. The executor splits target and draft contexts, attaches separate thread pools if needed, and passes the session to the scheduler. The scheduler records speculative indices in BatchCommit via record_speculative_indices for later merge operations.

Multimodal Support

With the mtmd feature flag, MultimodalRuntime loads during executor initialization:

// In run_initialized_executor, multimodal block
let multimodal = if config.multimodal_enabled {
    Some(MultimodalRuntime::load(...)?)
} else { None };
// Lines 161-184 in lib.rs

The multimodal pipeline integrates into prompt tokenization, handling projector models and image token injection without modifying core scheduling logic.


Complete Example: Running Inference Through the Backend

use icn_engine::lib::{NativeBackend, LlamaCompletionBackend};
use icn_contracts::inference::ResolvedInferenceRequest;
use std::path::PathBuf;
use std::sync::Arc;

// 1️⃣ Initialize once per process
let backend = NativeBackend::initialize()?;

// 2️⃣ Prepare load plan
let defaults = model_plan_defaults();
let intent = execution_intent(
    PathBuf::from("model.gguf"),
    None,  // no projector
    &defaults
);
let prepared = backend.prepare_load(
    "my-model".to_string(),
    intent,
    icn_contracts::SpeculativeDecodingConfig::Disabled { /* ... */ },
    hardware_snapshot,
    TemplateCapabilities::default(),
    ReasoningProfile::default(),
    "fingerprint".to_string(),
    false,  // expects_vision
)?;

// 3️⃣ Create completion backend with observer
struct MyObserver;
impl InferenceObserver for MyObserver {
    fn on_event(&self, event: InferenceEvent) {
        println!("Event: {:?}", event);
    }
}

let observer = Arc::new(MyObserver {});
let completion_backend = prepared.execute(observer)?;

// 4️⃣ Issue inference request
let request = ResolvedInferenceRequest::new(/* ... */);
completion_backend.complete(
    request,
    |admitted| {
        println!("Prompt accepted: {} tokens", admitted);
        Ok(())
    },
    |event| {
        println!("Token: {}", event);
        Ok(())
    }
)?;
// Sources: lib.rs lines 34, 26-30, 862-870

This pattern—initialize once, prepare load, execute with observer—provides a clean async‑native interface while the complex scheduling and execution machinery runs internally.


Summary

  • NativeBackend provides process‑wide llama.cpp access with hardware‑aware load planning
  • Executor thread isolates native inference work, processing commands and streaming events
  • SequencePool enables KV cache prefix reuse when prompt similarity exceeds 10%
  • BatchPlanner implements deterministic decode‑first, round‑robin pre‑fill scheduling
  • PromptLayout/PromptBoundary track logical and native positions for accurate cache management
  • Speculative decoding and multimodal features extend capabilities without scheduler changes

Frequently Asked Questions

How does icn‑engine ensure fair scheduling across concurrent requests?

The BatchPlanner uses a decode‑first, round‑robin pre‑fill policy. Decode tokens for active sequences are always processed first and sorted by sequence ID for determinism. Remaining batch capacity is distributed as quanta across pre‑fill sequences using a rotating cursor, preventing any single request from monopolizing resources.

Can icn‑engine reuse KV cache across different requests?

Yes, through SequencePool::acquire_matching. When a sequence completes, its ReusablePrefix (prompt layout + checkpoint) is retained. New requests compare their prompt against available prefixes; if similarity exceeds SLOT_PROMPT_SIMILARITY_THRESHOLD (0.1), the cached KV is reused, skipping redundant pre‑fill computation.

What happens when speculative decoding is enabled?

The load planner creates a SpeculativeSession with separate target and draft contexts. The scheduler records speculative indices in each BatchCommit, which the decoder uses to merge draft tokens with target KV cache. This adds minimal overhead to the core scheduling loop while potentially doubling throughput for compatible models.

How is multimodal input handled in the scheduler?

Multimodal support is orthogonal to scheduling. When the mtmd feature is compiled, MultimodalRuntime loads during executor initialization and integrates into prompt tokenization. The scheduler continues to operate on PromptLayout and PromptSegment abstractions, with image tokens handled transparently as additional segments with proper native_positions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →