What Is the icn-api Crate in Magnitude's ICN? HTTP Interface Explained
The icn-api crate is the HTTP façade of Magnitude's Inference Compute Node (ICN), transforming core inference runtime traits into a standards-compliant web service using Axum.
The icn-api crate sits at the network boundary of Magnitude's distributed inference stack. It exposes the internal ICN capabilities—model loading, chat completions, hardware introspection, and streaming events—as REST endpoints that the SDK, client applications, and external tools can consume. This article breaks down its architecture, key responsibilities, and how it bridges the low-level inference engine with higher-level client abstractions.
Core Responsibilities of icn-api
The crate handles eight primary concerns, each implemented in inference/crates/icn-api/src/lib.rs.
Routing and Request Handling with Axum
The app(state: AppState) -> Router function constructs the Axum router, mapping HTTP routes to async handlers. Key endpoints include /health, /api/v1/models/{model_id}/load-plan, and /v1/chat/completions【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L70-L77】.
The router pattern enables clean separation of concerns: handlers focus on request/response translation while the AppState provides access to backend implementations.
State Aggregation via AppState
The AppState struct aggregates optional trait objects for all ICN subsystems: catalog models, discovery, hardware, downloads, the model controller, and identity metadata【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L72-L84】.
This design allows icn-api to operate in degraded modes. A node can start without a catalog provider or with a stubbed hardware backend, then gain capabilities as components initialize.
Model Lifecycle Management
Endpoints for model management delegate to the ModelInstanceController trait. The ensure_model_instance handler triggers model loading through configured backends, with StaticModelInstanceController providing a concrete implementation for static deployments【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L62-L73】.
Additional endpoints cover load plan previews, catalog installation/removal, and graceful instance shutdown.
Hardware and Catalog Introspection
The hardware handler exposes system capabilities through the HardwareProvider trait, while catalog and discovery endpoints surface available models via CatalogModels and DiscoveredModels implementations【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L216-L224】.
This introspection layer enables the Magnitude SDK to match client requests against node capabilities before dispatching inference workloads.
OpenAPI Specification Generation
The serve_openapi handler generates machine-readable API documentation at /openapi.json using the utoipa crate【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L74-L78】.
This supports client code generation, automated testing, and third-party integration without manual API documentation maintenance.
Server-Sent Event Streaming
The watch_inference_events handler publishes SSE streams for reactive client updates, multiplexing broadcasts from the model controller, download manager, catalog, discovery, and assessment subsystems【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L96-L120】.
Clients receive real-time notifications for model load completion, download progress, and resource invalidations without polling.
Authorization Middleware
When an authorization token is configured, the authorize middleware validates Bearer <token> headers before allowing access to protected endpoints【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L52-L70】.
This enables deployment scenarios ranging from open internal networks to authenticated multi-tenant clusters.
Unified Error Handling
The ApiError type converts internal InferenceError and InventoryError variants into consistent JSON responses, ensuring API consumers receive predictable error structures regardless of which backend subsystem fails【/cache/repos/github.com/magnitudedev/magnitude/main/inference/crates/icn-api/src/lib.rs#L83-L125】.
Architecture Position in Magnitude's Stack
The icn-api crate occupies a specific layer in Magnitude's dependency hierarchy:
clients → client-common → SDK → icn-api → acn-protocol / acn → inference engine
The acn crate implements low-level RPC protocol details, while icn-api provides the higher-level HTTP/JSON abstraction that the SDK consumes. This separation yields three architectural benefits:
- Pluggable backends — Any
CompletionBackendimplementation can be wrapped byStaticModelInstanceControllerand exposed via HTTP without modifying icn-api's routing logic - Observability integration — SSE streams, OpenAPI exports, and structured health responses make ICN nodes compatible with monitoring platforms and orchestration tools
- Version negotiation —
ServerIdentity(instance ID, API version, native build) returned by/healthallows clients to verify compatibility before issuing requests
Practical Example: Starting an ICN Server
The following demonstrates bootstrapping a minimal ICN HTTP server with a static test backend:
use icn_api::{app, AppState};
use icn_contracts::dummy::StaticBackend;
#[tokio::main]
async fn main() {
let backend = StaticBackend::new("demo-model".to_string());
let state = AppState::model_free().with_model_controller(
std::sync::Arc::new(icn_api::StaticModelInstanceController::new(
std::sync::Arc::new(backend),
)),
);
let router = app(state);
axum::Server::bind(&"127.0.0.1:8080".parse().unwrap())
.serve(router.into_make_service())
.await
.unwrap();
}
After startup, verify the node with the health endpoint:
curl -s http://localhost:8080/health | jq .
Expected response:
{
"status": "ok",
"ready": true,
"version": "0.1.0",
"api_version": 1,
"instance_id": "embedded",
"native_build": "unknown"
}
Request an OpenAI-compatible chat completion:
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "demo-model", "messages": [{"role": "user", "content": "Hello, ICN!"}]}'
Key Source Files
| File | Purpose |
|---|---|
inference/crates/icn-api/Cargo.toml |
Crate metadata and dependencies (axum, utoipa, serde) |
inference/crates/icn-api/src/lib.rs |
Core implementation: AppState, router, handlers, SSE, errors |
design/icn/provider.md |
Design document for ICN provider layer exposure |
design/inference/engine.md |
High-level inference engine architecture |
design/inference/http-protocol-compatibility.md |
OpenAI-style endpoint compatibility guarantees |
Summary
- icn-api transforms Magnitude's ICN runtime into an HTTP service using Axum
- The
AppStatestruct aggregates trait objects for all ICN subsystems, enabling flexible deployment configurations - Endpoints cover model lifecycle, hardware introspection, chat completions, and real-time event streaming
- OpenAPI generation and structured error handling support production integration
- Authorization middleware and version negotiation enable multi-tenant and clustered deployments
Frequently Asked Questions
What protocol does icn-api use for client communication?
icn-api uses standard HTTP/1.1 and HTTP/2 with JSON request/response bodies. For real-time updates, it implements Server-Sent Events (SSE) via the watch_inference_events handler, streaming resource invalidations and state changes to connected clients.
How does icn-api relate to the acn crate?
The acn crate implements Magnitude's low-level RPC protocol between nodes, while icn-api provides the higher-level HTTP/JSON interface consumed by the SDK. The SDK communicates with icn-api endpoints, which may internally coordinate with other nodes through acn's RPC mechanisms.
Can icn-api operate without all backend providers initialized?
Yes. The AppState struct uses Option<Arc<dyn Trait>> for all provider fields, allowing nodes to start with partial capabilities. A node might begin with only hardware introspection, then gain model serving capabilities once a CompletionBackend becomes available.
Where is the OpenAPI specification served?
The machine-readable OpenAPI 3.0 specification is available at /openapi.json, generated at runtime by the utoipa crate. This endpoint requires no authentication and enables automated client generation and API discovery.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →