How to Monitor Resource Usage of a microsandbox Instance: Complete Guide with SDK and CLI
Microsandbox records CPU, memory, disk I/O, network I/O, and uptime metrics in a shared-memory registry that you can query via the Rust SDK or CLI commands.
This guide covers how to monitor resource usage of microsandbox instances using both programmatic interfaces and command-line tools. The microsandbox project implements a lightweight, real-time metrics system that balances granularity with minimal overhead.
Understanding the Metrics Architecture
Microsandbox stores runtime metrics in a shared-memory registry (microsandbox_metrics crate). The collector writes samples from the runtime, enabling low-latency reads without interrupting the sandbox itself.
The registry lives in crates/metrics/src/lib.rs and is populated by the collector in crates/metrics-collector/src/lib.rs. This design lets you query metrics from any language SDK or the CLI without requiring in-process instrumentation.
Configuring Metrics Sampling
Before monitoring, ensure metrics collection is enabled. You control this through the sandbox configuration.
In sdk/rust/lib/sandbox/config.rs (lines 192-199), the SandboxConfig::effective_metrics_interval resolves the actual sampling rate:
// From config.rs - effective interval resolution
pub fn effective_metrics_interval(&self) -> Option<Duration> {
self.metrics_sample_interval_ms
.map(|ms| Duration::from_millis(ms as u64))
.or_else(|| Some(DEFAULT_METRICS_INTERVAL))
}
Set the interval during sandbox creation:
use microsandbox::SandboxBuilder;
use std::time::Duration;
let sandbox = SandboxBuilder::new()
.name("production-api")
.image("docker.io/library/alpine:latest")
.metrics_sample_interval(Duration::from_millis(750)) // 750ms sampling
.run()
.await?;
- Shorter intervals (100-500ms): Higher granularity for debugging or bursty workloads
- Longer intervals (5-10s): Lower overhead for steady-state monitoring
- Disabled: Set
metrics_sample_interval_ms: Noneto turn off collection entirely
The effective configuration is resolved live, so metrics reflect any runtime resizes without restarting the sandbox.
Fetching a Single Metrics Snapshot
For point-in-time monitoring, use Sandbox::metrics() defined in sdk/rust/lib/sandbox/metrics.rs (lines 101-108):
// metrics.rs L101-108
pub async fn metrics(&self) -> Result<SandboxMetrics> {
if !self.config.metrics_enabled() {
return Err(Error::MetricsDisabled);
}
self.backend.metrics(&self.id).await
}
Example usage:
let snapshot = sandbox.metrics().await?;
println!("CPU: {:.1}%", snapshot.cpu_percent);
println!("Memory: {} MiB", snapshot.memory_bytes / 1024 / 1024);
This validates that metrics are enabled, then forwards to the backend implementation.
Streaming Real-Time Metrics
For continuous monitoring, Sandbox::metrics_stream() creates an async stream. From sdk/rust/lib/sandbox/metrics.rs (lines 15-26):
// metrics.rs L15-26
pub fn metrics_stream(&self, interval: Duration) -> impl Stream<Item = Result<SandboxMetrics>> {
let backend = self.backend.clone();
let id = self.id.clone();
stream::unfold((backend, id), move |(backend, id)| {
let interval = interval;
async move {
tokio::time::sleep(interval).await;
let result = backend.metrics(&id).await;
Some((result, (backend, id)))
}
})
}
Key detail: The stream interval is independent of the sandbox's configured sampling rate. You can stream every second even if the sandbox samples every 5 seconds—the backend returns the latest available sample.
Example streaming implementation:
use futures::stream::StreamExt;
let mut stream = sandbox.metrics_stream(Duration::from_secs(1));
while let Some(result) = stream.next().await {
match result {
Ok(m) => println!(
"{} | CPU: {:5.1}% | Net TX: {} B",
m.timestamp, m.cpu_percent, m.net_tx_bytes
),
Err(e) => eprintln!("Metrics error: {}", e),
}
}
The backend functions local_metrics (lines 34-55) and local_metrics_stream (lines 60-73) in the same file handle the actual registry reads for local backends.
Understanding SandboxMetrics Fields
The SandboxMetrics struct in sdk/rust/lib/sandbox/metrics.rs (lines 27-60) exposes these fields:
| Field | Type | Description |
|---|---|---|
cpu_percent |
f32 |
CPU usage percentage across host CPUs |
vcpu_time_ns |
u64 |
Guest vCPU execution time in nanoseconds |
memory_bytes |
u64 |
Resident memory usage in bytes |
memory_available_bytes |
Option<u64> |
Memory available to the guest |
memory_host_resident_bytes |
Option<u64> |
Host-side resident set size |
memory_limit_bytes |
u64 |
Configured guest memory limit |
disk_read_bytes |
u64 |
Cumulative disk bytes read |
disk_write_bytes |
u64 |
Cumulative disk bytes written |
net_rx_bytes |
u64 |
Cumulative network bytes received |
net_tx_bytes |
u64 |
Cumulative network bytes transmitted |
uptime |
Duration |
Time since sandbox start |
timestamp |
DateTime<Utc> |
When this sample was recorded |
Use disk_read_bytes and disk_write_bytes for I/O rate calculations by comparing consecutive samples. Network counters are cumulative—divide by elapsed time for throughput rates.
Monitoring Multiple Sandboxes
For dashboard or inventory use cases, all_sandbox_metrics() returns snapshots for every known sandbox. From sdk/rust/lib/sandbox/metrics.rs (lines 212-222):
// metrics.rs L212-222
pub async fn all_sandbox_metrics() -> Result<HashMap<String, SandboxMetrics>> {
let backend = default_backend().await?;
let sandboxes = backend.list().await?;
let mut results = HashMap::new();
for sandbox in sandboxes {
if let Ok(metrics) = sandbox.metrics().await {
results.insert(sandbox.name().to_string(), metrics);
}
}
Ok(results)
}
This iterates all sandboxes, skipping those without metrics enabled.
CLI Commands for Resource Monitoring
The msb CLI exposes the same backend calls. Commands are generated from the Rust SDK in crates/cli/src/commands/sandbox.rs.
Single snapshot:
msb sandbox metrics <sandbox-name>
Streaming with custom interval:
msb sandbox metrics-stream <sandbox-name> --interval 500ms
All sandboxes summary:
msb sandbox list-metrics
Equivalent to the SDK's all_sandbox_metrics() function.
Complete Monitoring Example
use microsandbox::{SandboxBuilder, Result};
use std::time::Duration;
use futures::stream::StreamExt;
#[tokio::main]
async fn main() -> Result<()> {
// Create sandbox with 750ms sampling
let sandbox = SandboxBuilder::new()
.name("api-server")
.image("docker.io/library/alpine:latest")
.metrics_sample_interval(Duration::from_millis(750))
.run()
.await?;
// Initial health check
let initial = sandbox.metrics().await?;
println!("Started with {} MiB memory", initial.memory_bytes / 1024 / 1024);
// Monitor for 60 seconds
let mut stream = sandbox.metrics_stream(Duration::from_secs(2));
let start = std::time::Instant::now();
while start.elapsed() < Duration::from_secs(60) {
if let Some(Ok(m)) = stream.next().await {
let mem_mib = m.memory_bytes / 1024 / 1024;
let mem_pct = 100.0 * m.memory_bytes as f32 / m.memory_limit_bytes as f32;
println!(
"[{}] CPU: {:5.1}% | Mem: {:4} MiB ({:4.1}%) | Net: {} TX / {} RX",
m.timestamp.format("%H:%M:%S"),
m.cpu_percent, mem_mib, mem_pct,
m.net_tx_bytes, m.net_rx_bytes
);
}
}
Ok(())
}
Summary
- Configuration: Set
metrics_sample_intervalinSandboxBuilderorSandboxConfig; effective interval resolved viaeffective_metrics_interval()inconfig.rs - Single query: Use
sandbox.metrics()→SandboxMetrics(lines 101-108) - Streaming: Use
sandbox.metrics_stream(interval)for async iteration (lines 15-26) - Backend reads:
local_metricsandlocal_metrics_streamquery the shared-memory registry (lines 34-73) - Batch operations:
all_sandbox_metrics()returnsHashMap<String, SandboxMetrics>for all sandboxes (lines 212-222) - CLI:
msb sandbox metrics,msb sandbox metrics-stream,msb sandbox list-metrics
Frequently Asked Questions
How do I disable metrics collection to reduce overhead?
Set metrics_sample_interval_ms: None in your SandboxConfig, or omit the metrics_sample_interval call in SandboxBuilder. The metrics() and metrics_stream() methods will return Error::MetricsDisabled if you attempt to query disabled metrics.
Can I stream metrics faster than the sandbox samples?
Yes. The stream interval is independent of the sampling rate. If you stream every 100ms but the sandbox samples every 5 seconds, you'll receive the same cached value repeatedly until a new sample is written. The backend in local_metrics always returns the latest available snapshot.
What happens to metrics during a live resize?
Metrics reflect effective configuration immediately. The local_metrics function resolves SandboxConfig::effective_metrics_interval() on each call, so CPU/memory limits in the returned SandboxMetrics update without restarting the sandbox or monitoring client.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →