What Is the icn-api Crate in Magnitude's ICN? HTTP Interface Explained

The icn-api crate is the HTTP façade of Magnitude's Inference Compute Node (ICN), transforming core inference runtime traits into a standards-compliant web service using Axum.

The icn-api crate sits at the network boundary of Magnitude's distributed inference stack. It exposes the internal ICN capabilities—model loading, chat completions, hardware introspection, and streaming events—as REST endpoints that the SDK, client applications, and external tools can consume. This article breaks down its architecture, key responsibilities, and how it bridges the low-level inference engine with higher-level client abstractions.

Core Responsibilities of icn-api

The crate handles eight primary concerns, each implemented in inference/crates/icn-api/src/lib.rs.

Routing and Request Handling with Axum

The app(state: AppState) -> Router function constructs the Axum router, mapping HTTP routes to async handlers. Key endpoints include /health, /api/v1/models/{model_id}/load-plan, and /v1/chat/completions【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L70-L77】.

The router pattern enables clean separation of concerns: handlers focus on request/response translation while the AppState provides access to backend implementations.

State Aggregation via AppState

The AppState struct aggregates optional trait objects for all ICN subsystems: catalog models, discovery, hardware, downloads, the model controller, and identity metadata【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L72-L84】.

This design allows icn-api to operate in degraded modes. A node can start without a catalog provider or with a stubbed hardware backend, then gain capabilities as components initialize.

Model Lifecycle Management

Endpoints for model management delegate to the ModelInstanceController trait. The ensure_model_instance handler triggers model loading through configured backends, with StaticModelInstanceController providing a concrete implementation for static deployments【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L62-L73】.

Additional endpoints cover load plan previews, catalog installation/removal, and graceful instance shutdown.

Hardware and Catalog Introspection

The hardware handler exposes system capabilities through the HardwareProvider trait, while catalog and discovery endpoints surface available models via CatalogModels and DiscoveredModels implementations【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L216-L224】.

This introspection layer enables the Magnitude SDK to match client requests against node capabilities before dispatching inference workloads.

OpenAPI Specification Generation

The serve_openapi handler generates machine-readable API documentation at /openapi.json using the utoipa crate【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L74-L78】.

This supports client code generation, automated testing, and third-party integration without manual API documentation maintenance.

Server-Sent Event Streaming

The watch_inference_events handler publishes SSE streams for reactive client updates, multiplexing broadcasts from the model controller, download manager, catalog, discovery, and assessment subsystems【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L96-L120】.

Clients receive real-time notifications for model load completion, download progress, and resource invalidations without polling.

Authorization Middleware

When an authorization token is configured, the authorize middleware validates Bearer <token> headers before allowing access to protected endpoints【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L52-L70】.

This enables deployment scenarios ranging from open internal networks to authenticated multi-tenant clusters.

Unified Error Handling

The ApiError type converts internal InferenceError and InventoryError variants into consistent JSON responses, ensuring API consumers receive predictable error structures regardless of which backend subsystem fails【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L83-L125】.

Architecture Position in Magnitude's Stack

The icn-api crate occupies a specific layer in Magnitude's dependency hierarchy:


clients → client-common → SDK → icn-api → acn-protocol / acn → inference engine

The acn crate implements low-level RPC protocol details, while icn-api provides the higher-level HTTP/JSON abstraction that the SDK consumes. This separation yields three architectural benefits:

  • Pluggable backends — Any CompletionBackend implementation can be wrapped by StaticModelInstanceController and exposed via HTTP without modifying icn-api's routing logic
  • Observability integration — SSE streams, OpenAPI exports, and structured health responses make ICN nodes compatible with monitoring platforms and orchestration tools
  • Version negotiation — ServerIdentity (instance ID, API version, native build) returned by /health allows clients to verify compatibility before issuing requests

Practical Example: Starting an ICN Server

The following demonstrates bootstrapping a minimal ICN HTTP server with a static test backend:

use icn_api::{app, AppState};
use icn_contracts::dummy::StaticBackend;

#[tokio::main]
async fn main() {
    let backend = StaticBackend::new("demo-model".to_string());
    
    let state = AppState::model_free().with_model_controller(
        std::sync::Arc::new(icn_api::StaticModelInstanceController::new(
            std::sync::Arc::new(backend),
        )),
    );
    
    let router = app(state);
    axum::Server::bind(&"127.0.0.1:8080".parse().unwrap())
        .serve(router.into_make_service())
        .await
        .unwrap();
}

After startup, verify the node with the health endpoint:

curl -s http://localhost:8080/health | jq .

Expected response:

{
  "status": "ok",
  "ready": true,
  "version": "0.1.0",
  "api_version": 1,
  "instance_id": "embedded",
  "native_build": "unknown"
}

Request an OpenAI-compatible chat completion:

curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "demo-model", "messages": [{"role": "user", "content": "Hello, ICN!"}]}'

Key Source Files

File Purpose
inference/crates/icn-api/Cargo.toml Crate metadata and dependencies (axum, utoipa, serde)
inference/crates/icn-api/src/lib.rs Core implementation: AppState, router, handlers, SSE, errors
design/icn/provider.md Design document for ICN provider layer exposure
design/inference/engine.md High-level inference engine architecture
design/inference/http-protocol-compatibility.md OpenAI-style endpoint compatibility guarantees

Summary

  • icn-api transforms Magnitude's ICN runtime into an HTTP service using Axum
  • The AppState struct aggregates trait objects for all ICN subsystems, enabling flexible deployment configurations
  • Endpoints cover model lifecycle, hardware introspection, chat completions, and real-time event streaming
  • OpenAPI generation and structured error handling support production integration
  • Authorization middleware and version negotiation enable multi-tenant and clustered deployments

Frequently Asked Questions

What protocol does icn-api use for client communication?

icn-api uses standard HTTP/1.1 and HTTP/2 with JSON request/response bodies. For real-time updates, it implements Server-Sent Events (SSE) via the watch_inference_events handler, streaming resource invalidations and state changes to connected clients.

How does icn-api relate to the acn crate?

The acn crate implements Magnitude's low-level RPC protocol between nodes, while icn-api provides the higher-level HTTP/JSON interface consumed by the SDK. The SDK communicates with icn-api endpoints, which may internally coordinate with other nodes through acn's RPC mechanisms.

Can icn-api operate without all backend providers initialized?

Yes. The AppState struct uses Option<Arc<dyn Trait>> for all provider fields, allowing nodes to start with partial capabilities. A node might begin with only hardware introspection, then gain model serving capabilities once a CompletionBackend becomes available.

Where is the OpenAPI specification served?

The machine-readable OpenAPI 3.0 specification is available at /openapi.json, generated at runtime by the utoipa crate. This endpoint requires no authentication and enables automated client generation and API discovery.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →