How Rigging Provides Health Checks and Server Runtime Primitives in Marin
Rigging supplies Marin services with a telemetry-based health-checking subsystem and lightweight server-runtime primitives including exponential backoff, rate limiting, and secure tunneling.
The rigging library serves as the foundational infrastructure layer for the Marin ecosystem, powering observability and runtime reliability for services such as Iris, Zephyr, and VLLM. Located in the marin-community/marin repository, this low-level Python package implements health checks through a process-local telemetry exporter while providing essential server runtime primitives that higher-level services consume without heavy transitive dependencies.
Telemetry-Based Health Checks in Rigging
Rigging implements health checks via a lightweight telemetry subsystem that exports metrics to Finelog or any HTTP endpoint. The architecture centers on a process-wide singleton _Runtime initialized through rigging.telemetry.configure().
Configuring the Telemetry Exporter
The telemetry system creates a global _Runtime instance on the first call to configure() in lib/rigging/src/rigging/telemetry/__init__.py (lines 505-543). Subsequent invocations are ignored to maintain consistent configuration throughout the process lifetime.
The exporter maintains a bounded queue to prevent OOM conditions, enforcing limits via max_queue_records and max_queue_bytes parameters in the _Runtime.emit() method (lines 74-92). A background delivery thread batches records and retries failed exports using exponential backoff.
Recording Runtime Health Metrics
The record_runtime_health() function (lines 86-113) captures internal exporter state and emits it as telemetry gauges. This function reads runtime_status() to gather:
queue_depth: Current number of pending recordstelemetry_lost_records: Count of dropped records due to queue limits- Export attempt counters
- Age of the oldest queued record
These metrics provide operators with immediate visibility into whether the telemetry pipeline itself is healthy.
Prometheus Integration
Rigging exposes these health gauges through PrometheusCollector and PrometheusScraper classes defined in lib/rigging/src/rigging/telemetry/prometheus.py (lines 13-45). These thin wrappers allow Marin services to expose telemetry data on standard /metrics endpoints for Prometheus scraping, creating a complete observability loop from process health to external monitoring systems.
Server Runtime Primitives
Beyond observability, rigging provides dependency-free utilities for building resilient server applications.
Retry Logic with Exponential Backoff
The ExponentialBackoff class in lib/rigging/src/rigging/timing.py (lines 58-78) generates retry intervals with configurable jitter. This primitive supports the initial, maximum, factor, and jitter parameters, enabling services to implement resilient retry loops for transient failures.
Rate Limiting
The RateLimiter class (lines 84-108) implements a token-bucket algorithm for throttling operations. Services invoke acquire() before executing rate-sensitive calls, protecting downstream systems from overload.
Secure Tunneling and Authentication
Rigging provides open_tunnel() in lib/rigging/src/rigging/tunnel.py (lines 31-71) to establish SSH-style TCP tunnels to remote GCP or CoreWeave nodes, optionally routed through IAP. For service-to-service authentication, rigging.server_auth.get_token() (lines 12-38) obtains short-lived JWTs, eliminating the need for persistent long-lived secrets when communicating with Iris controllers or VLLM servers.
Integration Pattern in Marin Services
Marin services typically initialize rigging at startup and leverage these primitives throughout their lifecycle:
from rigging.telemetry import configure, record_runtime_health
from rigging.timing import ExponentialBackoff, RateLimiter
from rigging.tunnel import open_tunnel
import time
# Initialize telemetry once per process
configure(
endpoint="https://finelog.example.com/api/v1/telemetry",
service="iris-controller",
attributes={"region": "us-central1"}
)
# Emit health snapshots periodically
record_runtime_health()
# Implement resilient retry logic
backoff = ExponentialBackoff(initial=0.2, maximum=5.0, factor=2.0, jitter=0.1)
while True:
try:
# Perform remote operation
break
except Exception:
time.sleep(backoff.next_interval())
# Rate-limit external API calls
limiter = RateLimiter(rate=100, per_second=True)
limiter.acquire()
This pattern allows services like Iris and Zephyr to maintain operational visibility while executing reliable, authenticated, and rate-limited operations against external resources.
Summary
- Rigging provides Marin’s foundational telemetry and runtime infrastructure through a lightweight, dependency-free library.
- Health checks rely on
record_runtime_health()exposing queue depth, lost records, and exporter status via the telemetry subsystem inrigging/telemetry/__init__.py. - The
configure()function establishes a process-wide singleton with bounded queues and background delivery threads to prevent memory exhaustion. - Server primitives include
ExponentialBackofffor retries,RateLimiterfor throttling,open_tunnel()for secure connectivity, andserver_auth.get_token()for JWT-based authentication. - Prometheus integration via
PrometheusCollectorbridges internal telemetry to external monitoring stacks.
Frequently Asked Questions
What is the purpose of the bounded queue in rigging's telemetry system?
The bounded queue prevents out-of-memory errors during telemetry spikes by enforcing max_queue_records and max_queue_bytes limits in _Runtime.emit(). When capacity is exceeded, new records are dropped and counted in telemetry_lost_records, ensuring the process remains stable even under high load.
How does record_runtime_health() help monitor Marin services?
This function samples the telemetry exporter's internal state—including queue depth, lost record counts, and export attempt metrics—and emits them as gauges. Operators can inspect these values through Finelog or Prometheus to determine if the telemetry pipeline itself is functioning correctly, providing a meta-level health check for the observability system.
When should I use ExponentialBackoff versus RateLimiter?
Use ExponentialBackoff when handling transient failures in retry loops, such as failed HTTP requests or storage operations. Use RateLimiter when you need to proactively throttle outgoing requests to respect API quotas or prevent overwhelming downstream services, regardless of whether previous requests failed.
Where are the rigging source files located in the Marin repository?
The core implementation resides in lib/rigging/src/rigging/, with telemetry logic in telemetry/__init__.py, timing utilities in timing.py, tunnel functionality in tunnel.py, and authentication helpers in server_auth.py. These paths provide the low-level primitives consumed by higher-level Marin services like Iris and Zephyr.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →