# How the icn‑engine Crate Handles Inference Execution and Scheduling in Magnitude

> Discover how the icn-engine crate in Magnitude separates inference execution and scheduling. Learn about its executor thread and Rust scheduler for efficient prompt caching and batch planning.

- Repository: [Magnitude/magnitude](https://github.com/magnitudedev/magnitude)
- Tags: internals
- Published: 2026-09-06

---

**The icn‑engine crate separates inference execution from scheduling: a dedicated executor thread runs native llama.cpp workloads while a lightweight, deterministic Rust scheduler manages prompt caching, sequence pooling, and batch planning with decode‑first ordering.**

The **icn‑engine** crate is the core inference runtime in [magnitudedev/magnitude](https://github.com/magnitudedev/magnitude), a Rust‑based AI inference engine. It orchestrates the full model lifecycle—from loading the native [`llama.cpp`](https://github.com/magnitudedev/magnitude/blob/main/llama.cpp) backend through scheduling concurrent requests to streaming token events back to SDK consumers. This article examines how inference execution and scheduling work, with direct references to the source implementation.

---

## Architecture Overview: Execution vs. Scheduling

The icn‑engine crate deliberately decouples **execution** (native GPU/CPU inference) from **scheduling** (deciding what to run and when). This separation keeps the scheduler policy‑free and fully testable in pure Rust, while the executor handles unsafe native calls.

| Component | Responsibility | Source Location |
|-----------|--------------|---------------|
| **NativeBackend** | Process‑wide `LlamaBackend` instance, hardware discovery, load planning | [`lib.rs#L34-L44`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/lib.rs#L34-L44) |
| **PreparedModelLoad** | Fully resolved load plan ready for instantiation | [`lib.rs#L84-L115`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/lib.rs#L84-L115) |
| **LlamaCompletionBackend** | Public `CompletionBackend` implementation for SDK | [`lib.rs#L862-L870`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/lib.rs#L862-L870) |
| **Executor thread** | Runs native context, processes commands, emits events | [`lib.rs#L885-L894`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/lib.rs#L885-L894) |
| **Scheduler** ([`scheduler.rs`](https://github.com/magnitudedev/magnitude/blob/main/scheduler.rs)) | Prompt layout, sequence pooling, batch planning | [[`scheduler.rs`](https://github.com/magnitudedev/magnitude/blob/main/scheduler.rs)](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs) |
| **BatchPlanner** | Orders work: decode tokens first, then pre‑fill quanta | [`scheduler.rs#L63-L82`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs#L63-L82) |
| **SequencePool** | Reusable sequence IDs with optional KV cache prefix reuse | [`scheduler.rs#L70-L91`](https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs#L70-L91) |

---

## Backend Initialization and Model Loading

Inference execution begins with a one‑time backend initialization. The `NativeBackend::initialize` method creates a process‑wide `LlamaBackend` instance wrapped in an `Arc`, ensuring all executor threads share a single native library instance.

```rust
let backend = NativeBackend::initialize()?;
// Source: https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/lib.rs#L34-L44

```

Load planning happens through `prepare_load`, which validates configuration, runs hardware‑aware planning via `icn_hardware::plan_load_with_backend`, and produces a `PreparedModelLoad`. This prepared plan includes:

- **ModelLoadPhase** values describing each loading step
- A command channel (`commands`) for the executor thread
- A `start` channel for the observer

```rust
let intent = execution_intent(model_path, projector_path, &defaults);
let prepared = backend.prepare_load(
    model_id,
    intent,
    speculative_cfg,
    hardware,
    caps,
    profile,
    fingerprint,
    expects_vision
)?;

```

---

## The Executor Thread: Native Inference Loop

The executor thread is where actual inference execution occurs. Created via `thread::Builder`, it runs `executor_main` for the lifetime of the model:

```rust
let (commands, command_receiver) = sync_channel(COMMAND_QUEUE_CAPACITY);
let executor = thread::Builder::new()
    .name(format!("icn-llama-{model_id}"))
    .spawn(move || executor_main(
        backend,
        planned,
        acceleration,
        command_receiver,
        // ...
    ));
// Source: https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/lib.rs#L885-L894

```

The `executor_main` function:

1. Creates a `LlamaThreadPool` for parallel tokenization
2. Initializes optional speculative decoding session
3. Loads multimodal runtime if `mtmd` feature is enabled
4. Enters `run_initialized_executor` to process `ExecutorCommand`s and emit `ExecutorItem`s

The thread persists until `LlamaCompletionBackend` is dropped, ensuring model context stability across many requests.

---

## Scheduling: Prompt Layout and Cache Reuse

The scheduler operates on **prompt layouts**—semantic representations that track both logical token counts and native KV positions.

### PromptBoundary and PromptLayout

```rust
pub(crate) struct PromptBoundary {
    pub(crate) logical_tokens: usize,
    pub(crate) native_position: i32,
}
// Source: https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs#L7-L20

```

A `PromptLayout` consists of `PromptSegment`s, each knowing its `logical_tokens` and `native_positions`. The `common_prefix` method computes shared prefixes between prompts, enabling cache reuse.

### SequencePool: Reusable Cache Prefixes

When a request completes pre‑fill, its `ReusablePrefix` (prompt layout plus checkpoint states) is retained. For new requests, `SequencePool::acquire_matching` attempts to find a sequence with sufficient prefix similarity:

```rust
let best = self.available.iter()
    .enumerate()
    .filter_map(|(index, sequence)| {
        let prefix = sequence.reusable_prefix.as_ref()?;
        let common_prefix = prefix.layout.common_prefix(prompt).logical_tokens;
        let similarity = common_prefix as f32 / prompt_tokens as f32;
        (similarity > SLOT_PROMPT_SIMILARITY_THRESHOLD).then_some((index, common_prefix))
    })
    .max_by_key(|(_, common_prefix)| *common_prefix)
    .map(|(index, _)| index);
// Source: https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs#L70-L91

```

The threshold `SLOT_PROMPT_SIMILARITY_THRESHOLD` is set to **0.1** (10%). Matching sequences skip redundant pre‑fill work, dramatically reducing latency for similar prompts.

---

## Batch Planning: Decode‑First Scheduling

Each scheduling tick produces a vector of `BatchWork` items through `BatchPlanner::plan`. The policy is intentionally simple and deterministic:

1. **Decode work first** — at most one token per sequence, sorted by sequence ID for stability
2. **Pre‑fill quanta** — round‑robin across sequences, limited to `prefill_quantum` tokens each

```rust
// Decode-first: sorted by sequence_id
let mut decode_candidates = candidates.iter()
    .filter(|c| c.kind == WorkKind::Decode)
    .collect::<Vec<_>>();
decode_candidates.sort_unstable_by_key(|c| c.sequence_id);

// Prefill round-robin with cursor-based fairness
let start = self.cursor % prefill_candidates.len();
// ...
while available > 0 { /* distribute quanta */ }
// Source: https://github.com/magnitudedev/magnitude/blob/main/inference/crates/icn-engine/src/scheduler.rs#L63-L82

```

This **decode‑first** approach minimizes latency for active generations while the **round‑robin pre‑fill** prevents starvation of new requests. The cursor (`self.cursor`) advances each tick, ensuring fair distribution without external policy dependencies.

---

## Request Lifecycle and Event Streaming

A complete inference request flows through these stages:

1. **Acquire sequence** — `pool.acquire_matching(&prompt)` attempts cache reuse
2. **Tokenize prompt** — convert chat to `TokenizedPrompt` (tokens + layout)
3. **Send command** — `ExecutorCommand::Complete` dispatched to executor thread
4. **Batch loop** — repeatedly call `BatchPlanner::plan`, feed to `LlamaBatch`, commit via `BatchCommit::apply`
5. **Emit events** — `ExecutorItem::Event` for each token, `ExecutorItem::Completed` at termination
6. **Return to pool** — sequence released with updated `ReusablePrefix`

The `ActiveRequest` struct holds per‑request mutable state including phase, token history, speculative data, and timing measurements. `BatchCommit::apply` mutates this state and records progress for the next scheduling tick.

---

## Advanced Features: Speculative Decoding and Multimodal

### Speculative Decoding

When `ExecutionIntent::speculative` is enabled, `prepare_native_plan` creates a `SpeculativeSession`. The executor splits target and draft contexts, attaches separate thread pools if needed, and passes the session to the scheduler. The scheduler records speculative indices in `BatchCommit` via `record_speculative_indices` for later merge operations.

### Multimodal Support

With the `mtmd` feature flag, `MultimodalRuntime` loads during executor initialization:

```rust
// In run_initialized_executor, multimodal block
let multimodal = if config.multimodal_enabled {
    Some(MultimodalRuntime::load(...)?)
} else { None };
// Lines 161-184 in lib.rs

```

The multimodal pipeline integrates into prompt tokenization, handling projector models and image token injection without modifying core scheduling logic.

---

## Complete Example: Running Inference Through the Backend

```rust
use icn_engine::lib::{NativeBackend, LlamaCompletionBackend};
use icn_contracts::inference::ResolvedInferenceRequest;
use std::path::PathBuf;
use std::sync::Arc;

// 1️⃣ Initialize once per process
let backend = NativeBackend::initialize()?;

// 2️⃣ Prepare load plan
let defaults = model_plan_defaults();
let intent = execution_intent(
    PathBuf::from("model.gguf"),
    None,  // no projector
    &defaults
);
let prepared = backend.prepare_load(
    "my-model".to_string(),
    intent,
    icn_contracts::SpeculativeDecodingConfig::Disabled { /* ... */ },
    hardware_snapshot,
    TemplateCapabilities::default(),
    ReasoningProfile::default(),
    "fingerprint".to_string(),
    false,  // expects_vision
)?;

// 3️⃣ Create completion backend with observer
struct MyObserver;
impl InferenceObserver for MyObserver {
    fn on_event(&self, event: InferenceEvent) {
        println!("Event: {:?}", event);
    }
}

let observer = Arc::new(MyObserver {});
let completion_backend = prepared.execute(observer)?;

// 4️⃣ Issue inference request
let request = ResolvedInferenceRequest::new(/* ... */);
completion_backend.complete(
    request,
    |admitted| {
        println!("Prompt accepted: {} tokens", admitted);
        Ok(())
    },
    |event| {
        println!("Token: {}", event);
        Ok(())
    }
)?;
// Sources: lib.rs lines 34, 26-30, 862-870

```

This pattern—initialize once, prepare load, execute with observer—provides a clean async‑native interface while the complex scheduling and execution machinery runs internally.

---

## Summary

- **NativeBackend** provides process‑wide [`llama.cpp`](https://github.com/magnitudedev/magnitude/blob/main/llama.cpp) access with hardware‑aware load planning
- **Executor thread** isolates native inference work, processing commands and streaming events
- **SequencePool** enables KV cache prefix reuse when prompt similarity exceeds 10%
- **BatchPlanner** implements deterministic decode‑first, round‑robin pre‑fill scheduling
- **PromptLayout/PromptBoundary** track logical and native positions for accurate cache management
- **Speculative decoding and multimodal** features extend capabilities without scheduler changes

---

## Frequently Asked Questions

### How does icn‑engine ensure fair scheduling across concurrent requests?

The **BatchPlanner** uses a **decode‑first, round‑robin pre‑fill** policy. Decode tokens for active sequences are always processed first and sorted by sequence ID for determinism. Remaining batch capacity is distributed as quanta across pre‑fill sequences using a rotating cursor, preventing any single request from monopolizing resources.

### Can icn‑engine reuse KV cache across different requests?

Yes, through **SequencePool::acquire_matching**. When a sequence completes, its `ReusablePrefix` (prompt layout + checkpoint) is retained. New requests compare their prompt against available prefixes; if similarity exceeds `SLOT_PROMPT_SIMILARITY_THRESHOLD` (0.1), the cached KV is reused, skipping redundant pre‑fill computation.

### What happens when speculative decoding is enabled?

The load planner creates a **SpeculativeSession** with separate target and draft contexts. The scheduler records speculative indices in each **BatchCommit**, which the decoder uses to merge draft tokens with target KV cache. This adds minimal overhead to the core scheduling loop while potentially doubling throughput for compatible models.

### How is multimodal input handled in the scheduler?

Multimodal support is **orthogonal to scheduling**. When the `mtmd` feature is compiled, **MultimodalRuntime** loads during executor initialization and integrates into prompt tokenization. The scheduler continues to operate on `PromptLayout` and `PromptSegment` abstractions, with image tokens handled transparently as additional segments with proper `native_positions`.