# How Rigging Provides Health Checks and Server Runtime Primitives in Marin

> Discover how Rigging enhances Marin services with health checks and server runtime primitives like rate limiting and secure tunneling. Improve your application's reliability and performance today.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-28

---

**Rigging supplies Marin services with a telemetry-based health-checking subsystem and lightweight server-runtime primitives including exponential backoff, rate limiting, and secure tunneling.**

The `rigging` library serves as the foundational infrastructure layer for the Marin ecosystem, powering observability and runtime reliability for services such as Iris, Zephyr, and VLLM. Located in the `marin-community/marin` repository, this low-level Python package implements health checks through a process-local telemetry exporter while providing essential server runtime primitives that higher-level services consume without heavy transitive dependencies.

## Telemetry-Based Health Checks in Rigging

Rigging implements health checks via a lightweight telemetry subsystem that exports metrics to Finelog or any HTTP endpoint. The architecture centers on a process-wide singleton `_Runtime` initialized through `rigging.telemetry.configure()`.

### Configuring the Telemetry Exporter

The telemetry system creates a global `_Runtime` instance on the first call to `configure()` in [`lib/rigging/src/rigging/telemetry/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/rigging/src/rigging/telemetry/__init__.py) (lines 505-543). Subsequent invocations are ignored to maintain consistent configuration throughout the process lifetime.

The exporter maintains a **bounded queue** to prevent OOM conditions, enforcing limits via `max_queue_records` and `max_queue_bytes` parameters in the `_Runtime.emit()` method (lines 74-92). A background delivery thread batches records and retries failed exports using exponential backoff.

### Recording Runtime Health Metrics

The `record_runtime_health()` function (lines 86-113) captures internal exporter state and emits it as telemetry gauges. This function reads `runtime_status()` to gather:

- `queue_depth`: Current number of pending records
- `telemetry_lost_records`: Count of dropped records due to queue limits
- Export attempt counters
- Age of the oldest queued record

These metrics provide operators with immediate visibility into whether the telemetry pipeline itself is healthy.

### Prometheus Integration

Rigging exposes these health gauges through `PrometheusCollector` and `PrometheusScraper` classes defined in [`lib/rigging/src/rigging/telemetry/prometheus.py`](https://github.com/marin-community/marin/blob/main/lib/rigging/src/rigging/telemetry/prometheus.py) (lines 13-45). These thin wrappers allow Marin services to expose telemetry data on standard `/metrics` endpoints for Prometheus scraping, creating a complete observability loop from process health to external monitoring systems.

## Server Runtime Primitives

Beyond observability, rigging provides dependency-free utilities for building resilient server applications.

### Retry Logic with Exponential Backoff

The `ExponentialBackoff` class in [`lib/rigging/src/rigging/timing.py`](https://github.com/marin-community/marin/blob/main/lib/rigging/src/rigging/timing.py) (lines 58-78) generates retry intervals with configurable jitter. This primitive supports the `initial`, `maximum`, `factor`, and `jitter` parameters, enabling services to implement resilient retry loops for transient failures.

### Rate Limiting

The `RateLimiter` class (lines 84-108) implements a token-bucket algorithm for throttling operations. Services invoke `acquire()` before executing rate-sensitive calls, protecting downstream systems from overload.

### Secure Tunneling and Authentication

Rigging provides `open_tunnel()` in [`lib/rigging/src/rigging/tunnel.py`](https://github.com/marin-community/marin/blob/main/lib/rigging/src/rigging/tunnel.py) (lines 31-71) to establish SSH-style TCP tunnels to remote GCP or CoreWeave nodes, optionally routed through IAP. For service-to-service authentication, `rigging.server_auth.get_token()` (lines 12-38) obtains short-lived JWTs, eliminating the need for persistent long-lived secrets when communicating with Iris controllers or VLLM servers.

## Integration Pattern in Marin Services

Marin services typically initialize rigging at startup and leverage these primitives throughout their lifecycle:

```python
from rigging.telemetry import configure, record_runtime_health
from rigging.timing import ExponentialBackoff, RateLimiter
from rigging.tunnel import open_tunnel
import time

# Initialize telemetry once per process

configure(
    endpoint="https://finelog.example.com/api/v1/telemetry",
    service="iris-controller",
    attributes={"region": "us-central1"}
)

# Emit health snapshots periodically

record_runtime_health()

# Implement resilient retry logic

backoff = ExponentialBackoff(initial=0.2, maximum=5.0, factor=2.0, jitter=0.1)
while True:
    try:
        # Perform remote operation

        break
    except Exception:
        time.sleep(backoff.next_interval())

# Rate-limit external API calls

limiter = RateLimiter(rate=100, per_second=True)
limiter.acquire()

```

This pattern allows services like Iris and Zephyr to maintain operational visibility while executing reliable, authenticated, and rate-limited operations against external resources.

## Summary

- **Rigging** provides Marin’s foundational telemetry and runtime infrastructure through a lightweight, dependency-free library.
- Health checks rely on `record_runtime_health()` exposing queue depth, lost records, and exporter status via the telemetry subsystem in [`rigging/telemetry/__init__.py`](https://github.com/marin-community/marin/blob/main/rigging/telemetry/__init__.py).
- The `configure()` function establishes a process-wide singleton with bounded queues and background delivery threads to prevent memory exhaustion.
- Server primitives include `ExponentialBackoff` for retries, `RateLimiter` for throttling, `open_tunnel()` for secure connectivity, and `server_auth.get_token()` for JWT-based authentication.
- Prometheus integration via `PrometheusCollector` bridges internal telemetry to external monitoring stacks.

## Frequently Asked Questions

### What is the purpose of the bounded queue in rigging's telemetry system?

The bounded queue prevents out-of-memory errors during telemetry spikes by enforcing `max_queue_records` and `max_queue_bytes` limits in `_Runtime.emit()`. When capacity is exceeded, new records are dropped and counted in `telemetry_lost_records`, ensuring the process remains stable even under high load.

### How does `record_runtime_health()` help monitor Marin services?

This function samples the telemetry exporter's internal state—including queue depth, lost record counts, and export attempt metrics—and emits them as gauges. Operators can inspect these values through Finelog or Prometheus to determine if the telemetry pipeline itself is functioning correctly, providing a meta-level health check for the observability system.

### When should I use `ExponentialBackoff` versus `RateLimiter`?

Use **ExponentialBackoff** when handling transient failures in retry loops, such as failed HTTP requests or storage operations. Use **RateLimiter** when you need to proactively throttle outgoing requests to respect API quotas or prevent overwhelming downstream services, regardless of whether previous requests failed.

### Where are the rigging source files located in the Marin repository?

The core implementation resides in `lib/rigging/src/rigging/`, with telemetry logic in [`telemetry/__init__.py`](https://github.com/marin-community/marin/blob/main/telemetry/__init__.py), timing utilities in [`timing.py`](https://github.com/marin-community/marin/blob/main/timing.py), tunnel functionality in [`tunnel.py`](https://github.com/marin-community/marin/blob/main/tunnel.py), and authentication helpers in [`server_auth.py`](https://github.com/marin-community/marin/blob/main/server_auth.py). These paths provide the low-level primitives consumed by higher-level Marin services like Iris and Zephyr.