How Magnitude Profiles Hardware for Optimal AI Model Selection
Magnitude implements a hardware-aware model selection pipeline that captures real-time CPU, GPU, and memory snapshots through a local Inference Control Node (ICN), derives a normalized hardware-memory view, and ranks AI models using a weighted blend of intelligence, speed, and fidelity scores to ensure optimal placement without resource exhaustion.
Magnitude (magnitudedev/magnitude) executes a complete hardware profiling workflow locally to match AI inference workloads with available system capabilities. By transforming raw hardware snapshots into structured memory views and applying predictive ranking algorithms, Magnitude filters incompatible models and selects the best-fit candidate for your specific configuration. Understanding how Magnitude profiles hardware for optimal AI model selection reveals a sophisticated approach to preventing out-of-memory crashes while maximizing model performance.
Capturing the Hardware Snapshot with ICN
The profiling process begins with the ICN (Inference Control Node) daemon, which polls the operating system to detect all available compute resources. This daemon captures detailed information about CPUs, GPUs, memory domains, and accelerator hardware, then exposes this data through the IcnHardware effect.
In packages/acn/src/local-inference-hardware.ts (lines 30‑62), the system projects the raw ICN data into a typed LocalInferenceHardware value. The implementation exposes two critical operations: hardware.get to retrieve the current snapshot and hardware.refresh to update the view when hardware changes occur. This projection runs as a background stream, ensuring the hardware view remains current when devices are added or removed, such as plugging in an external GPU.
Deriving the Hardware-Memory View
Raw hardware snapshots undergo normalization in packages/client-common/src/utils/hardware-memory.ts through the deriveHardwareMemoryView function (lines 55‑66). This transformation produces a hardware-memory view that standardizes device names across platforms and calculates precise free and used byte counts for every memory domain.
The function aggregates totalBytes and availableBytes from system RAM and accelerator memory (identified by memoryDomainId), flags inconsistencies such as missing free-memory information, and caches the result. When invoked, it accepts configuration options like fallbackToAccelerators to determine how to handle heterogeneous memory pools.
// Pull the current hardware view (client-common)
const hardwareView = deriveHardwareMemoryView(
hardwareSnapshot, // LocalInferenceHardware from IcnHardware
Option.none(), // No active model allocation yet
{ fallbackToAccelerators: true }
);
Filtering and Ranking Models
Once Magnitude establishes the hardware-memory view, it applies a two-stage validation and scoring system to select viable models.
Memory Constraint Validation
Before ranking candidates, Magnitude enforces hard memory limits using targetPhysicalMemoryBytes. In packages/client-common/src/local-models/setup.ts (line 362), the system compares the model's requiredBytes against available hardware memory. If the model exceeds capacity, Magnitude throws an insufficient_resources error, preventing out-of-memory crashes before inference begins.
// Check if a model fits the available memory
if (targetPhysicalMemoryBytes(hardwareView) < model.requiredBytes) {
throw new Error(
"This model does not fit available hardware."
);
}
Normalized Speed and Intelligence Scoring
For models that pass memory validation, Magnitude computes ranking scores in packages/acn/src/local-model-ranking-policy.ts (lines 30‑45) via the modelRankingScores function. This algorithm blends three weighted factors:
- Intelligence (50% weight): Baseline capability score of the model
- Speed (30% weight): Normalized token-per-second performance via
normalizedModelSpeedScore - Fidelity (20% weight): Quality rank of the model's output
The speed component applies a ceiling constant (SPEED_SCORE_CEILING) to raw token-per-second benchmarks, then maps the result to a 0‑1 scale. This normalization ensures slower GPUs remain viable options when they satisfy latency budgets, rather than being filtered out by absolute performance thresholds.
// Compute a ranking score for a candidate model
const scores = modelRankingScores({
profile: model.servingProfile,
performance: benchmarkSamples,
intelligenceScore: model.intelligence,
fidelityRank: model.fidelity,
});
if (Option.isSome(scores)) {
const { intelligence, speed, fidelity } = Option.getOrThrow(scores);
const overall = 0.5 * intelligence + 0.3 * speed + 0.2 * fidelity;
console.log(`Overall ranking for ${model.id}: ${overall}`);
}
Final Model Selection and Caching
The selection layer in packages/acn/src/model-selection.ts (lines 62‑94) combines algorithmic scores with user preferences through makeModelSelection. This function weights the hardware-derived rankings against recency and favorite model flags to present the optimal choice in the Hardware tab interface.
The final selection persists in the ModelSelection state, ensuring subsequent inference requests automatically utilize the best-fit model without re-evaluating the hardware profile. This stateful approach eliminates redundant profiling overhead while maintaining responsiveness to hardware changes through the ICN daemon's background refresh cycle.
Summary
- ICN daemon polls local hardware and projects snapshots into
LocalInferenceHardwarevia theIcnHardwareeffect inpackages/acn/src/local-inference-hardware.ts - Memory normalization occurs in
deriveHardwareMemoryView(packages/client-common/src/utils/hardware-memory.ts), which calculates available bytes across all memory domains and accelerator pools - Hard constraints prevent OOM errors by rejecting models exceeding
targetPhysicalMemoryByteswithinsufficient_resourceserrors - Ranking algorithm blends intelligence (50%), normalized speed (30%), and fidelity (20%) scores, with speed normalized against
SPEED_SCORE_CEILING - Stateful selection stores the optimal model in
ModelSelectionstate viamakeModelSelection, balancing hardware capabilities with user preferences
Frequently Asked Questions
How does Magnitude detect hardware changes in real-time?
The ICN daemon runs a background stream that calls hardware.refresh whenever the OS reports hardware events, such as GPU hot-plugging or memory allocation changes. This updates the LocalInferenceHardware projection in packages/acn/src/local-inference-hardware.ts without requiring application restarts, ensuring model recommendations reflect current system capabilities.
What prevents Magnitude from selecting a model that exhausts available memory?
Magnitude implements a hard constraint check in packages/client-common/src/local-models/setup.ts (line 362) using targetPhysicalMemoryBytes. Before ranking occurs, the system verifies that the model's requiredBytes fits within the aggregated availableBytes of all memory domains. If insufficient, Magnitude raises an insufficient_resources error, blocking deployment of oversized models.
How is the model ranking score calculated?
The modelRankingScores function in packages/acn/src/local-model-ranking-policy.ts calculates a composite score using the formula: 0.5 * intelligence + 0.3 * speed + 0.2 * fidelity. The speed component uses normalizedModelSpeedScore to cap raw token-per-second values at SPEED_SCORE_CEILING and map them to a 0‑1 scale, enabling fair comparison across disparate hardware.
Does hardware profiling require external services or cloud calls?
No. According to the magnitudedev/magnitude source code, the entire hardware profiling pipeline runs locally through the ICN daemon. All hardware snapshots, memory calculations, and ranking computations execute on the host machine, ensuring privacy and eliminating network latency from the model selection process.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →