How Magnitude Ensures Privacy and Cost-Efficiency with Local AI Models
Magnitude guarantees data privacy by processing all inference locally through the local provider unless a cloud API key is explicitly configured, while delivering cost efficiency through zero per-request fees and intelligent hardware-aware model ranking.
The magnitudedev/magnitude repository implements a privacy-first AI architecture that keeps user data on-device by default. By routing all model inference through the local Agent Communication Network (ACN) daemon and requiring explicit opt-in for external APIs, Magnitude ensures privacy and cost-efficiency with local AI models without sacrificing usability.
Local-First Architecture for Guaranteed Privacy
On-Device Inference Isolation
Magnitude’s privacy guarantee rests on the local provider identifier defined in info/local-inference-capacity-and-ranking.md. When the SDK receives a request with provider: 'local', it instructs the ACN daemon to load models exclusively from the ICN-managed model store on the user’s disk rather than contacting remote endpoints.
Because the client communicates only with the local daemon, no prompts or generated text traverse external networks. The daemon projects only model metadata—including catalog entries, memory-fit calculations, and speed/quality scores—to the client via the LocalModel union described in design/model-management/local-model-product-projection.md. This projection ensures the client receives authoritative ranking data without exposing inference payloads to the network.
Explicit Cloud Opt-In Requirements
Cloud model usage is strictly optional per info/cloud-model-usage.md. The SDK enforces this by requiring a valid Magnitude API key to enable cloud providers; in the absence of this key, the system automatically falls back to the local provider. This design prevents accidental data leakage—unless the user explicitly supplies an API key, all setModel calls resolve to on-device inference.
Cost-Efficiency Through Local Compute
Zero Per-Request Pricing
When using the local provider, Magnitude eliminates per-token and per-call fees entirely. Unlike cloud APIs that bill per request according to the Pro subscription pricing outlined in info/cloud-model-usage.md, local inference costs only the compute resources of the host machine. This transforms variable API costs into fixed infrastructure utilization.
Memory-Aware Model Selection
The ACN runtime implements smart resource management to maximize hardware utilization. Before loading a model, the system re-checks available machine memory (see lines 23-28 of info/local-inference-capacity-and-ranking.md) and may prompt users to close memory-heavy applications to prevent out-of-memory failures.
Additionally, the local-model projection ranks candidates by Speed and Intelligence while respecting the device’s memory ceiling. The interface displays only the top-10 downloadable models that fit current hardware constraints, preventing wasteful downloads of oversized or underperforming models. This selective projection, detailed in design/model-management/local-model-product-projection.md, ensures users never allocate resources to incompatible model artifacts.
Implementing Privacy-First Model Selection
Selecting Local Models from the Catalog
The following pattern demonstrates how to query the local catalog and select the highest-ranked model that fits your hardware:
import { ProviderClient } from '@magnitudedev/sdk'
// Create a client (the SDK automatically connects to the local ACN daemon)
const client = ProviderClient.make()
// List local catalog entries (only models with `provider === 'local'` are shown)
const catalog = await client.listModels({ provider: 'local' })
// Pick the top‑ranked model that fits the device memory
const chosen = catalog.models.find(m => m.memoryFit && m.rank <= 10)
// Instruct the daemon to use the selected model for inference
await client.setModel({ modelId: chosen.modelId, provider: 'local' })
This implementation ensures all inference remains on-device by explicitly filtering for the local provider and verifying memoryFit before selection.
Enforcing Cloud Fallback Protection
The SDK automatically prevents cloud leakage when no API key is configured:
import { ProviderClient } from '@magnitudedev/sdk'
// The client reads the API key from the `.env` file (if any)
const client = ProviderClient.make()
// If no key is configured, `setModel` will reject cloud providers
try {
await client.setModel({ modelId: 'gpt-4o', provider: 'openai' })
} catch (e) {
console.log('Cloud model unavailable – using local model instead')
}
This try-catch pattern demonstrates the enforced fallback behavior that protects user data when cloud credentials are absent.
Monitoring Optional Cloud Usage
For users who do configure cloud access, the SDK provides visibility into consumption:
import { ProviderClient } from '@magnitudedev/sdk'
const client = ProviderClient.make()
const usage = await client.getUsage() // Returns current window usage & limits
console.log(`Used ${usage.tokens} tokens this week; limit ${usage.weeklyLimit}`)
This optional monitoring complements the privacy-first default by making cloud costs transparent and auditable.
Summary
- Local provider isolation: The
localprovider ininfo/local-inference-capacity-and-ranking.mdensures all inference occurs on-device with no external HTTP calls. - Explicit cloud opt-in: Cloud models require a Magnitude API key per
info/cloud-model-usage.md, with automatic fallback to local inference when keys are absent. - Zero per-request costs: Local inference incurs no token fees, converting variable API costs to fixed compute utilization.
- Hardware-aware optimization: Memory checks (lines 23-28) and ranked model projection prevent resource waste and out-of-memory failures.
Frequently Asked Questions
How does Magnitude prevent accidental data transmission to cloud APIs?
Magnitude requires an explicit API key configuration to enable cloud providers. As implemented in the SDK client logic, the setModel method rejects cloud provider requests when no key is present, automatically falling back to the local provider per info/cloud-model-usage.md. This ensures that unless the user deliberately supplies credentials, all data remains on the host machine.
What files control how Magnitude ranks local models for hardware compatibility?
The ranking and filtering logic resides in info/local-inference-capacity-and-ranking.md (which defines memory-fit calculations and hardware constraints) and design/model-management/local-model-product-projection.md (which implements the ACN projection layer that unifies catalog data into the LocalModel union). These files ensure only compatible, high-performing models are presented to the user.
Does using local models incur any usage fees?
No. When using the local provider, Magnitude charges no per-request, per-token, or per-call fees. The only costs are the computational resources consumed on the user’s own machine, contrasting with cloud providers that require Pro subscription pricing as detailed in info/cloud-model-usage.md.
How does Magnitude handle insufficient memory for local model inference?
Before loading a model, the ACN runtime checks available system memory (see implementation in lines 23-28 of info/local-inference-capacity-and-ranking.md). If resources are constrained, the system may prompt users to close memory-intensive applications or select a smaller model from the ranked catalog, preventing costly out-of-memory crashes and maintaining system responsiveness.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →