How to Monitor Hindsight Using Prometheus Metrics: A Complete Setup Guide
Hindsight exposes Prometheus-compatible metrics through built-in OpenTelemetry instrumentation, providing a /metrics endpoint that can be scraped by any Prometheus server or the included Grafana LGTM stack.
The vectorize-io/hindsight repository ships with a complete observability layer that leverages OpenTelemetry to generate Prometheus metrics. This instrumentation tracks operation latency, LLM call durations, HTTP request timing, and system metrics without requiring external dependencies beyond the standard FastAPI application stack.
Initializing the Metrics Collector
The monitoring stack begins with initialize_metrics() in hindsight_api_slim/hindsight_api/metrics.py (lines 94-52). This function creates a PrometheusMetricReader, configures custom histogram views for different latency types, and establishes a global meter provider.
# hindsight_api_slim/hindsight_api/metrics.py
def initialize_metrics(service_name: str = "hindsight-api", service_version: str = "1.0.0"):
"""
Initialise OpenTelemetry metrics with a Prometheus exporter.
Returns a PrometheusMetricReader for exposing the /metrics endpoint.
"""
global _meter
resource = Resource.create(
{"service.name": service_name, "service.version": service_version}
)
prometheus_reader = PrometheusMetricReader()
# Custom histogram buckets for operation, LLM and HTTP latency
duration_view = View(
instrument_name="hindsight.operation.duration",
aggregation=ExplicitBucketHistogramAggregation(boundaries=DURATION_BUCKETS),
)
llm_duration_view = View(
instrument_name="hindsight.llm.duration",
aggregation=ExplicitBucketHistogramAggregation(boundaries=LLM_DURATION_BUCKETS),
)
http_duration_view = View(
instrument_name="hindsight.http.duration",
aggregation=ExplicitBucketHistogramAggregation(boundaries=HTTP_DURATION_BUCKETS),
)
provider = MeterProvider(
resource=resource,
metric_readers=[prometheus_reader],
views=[duration_view, llm_duration_view, http_duration_view],
)
metrics.set_meter_provider(provider)
_meter = metrics.get_meter(__name__)
return prometheus_reader
Both the worker and API server invoke this function during startup. The worker calls it in worker/main.py (lines 48-52), while the API initializes it during app creation in api/http.py.
Exposing the /metrics HTTP Endpoint
Once initialized, the application exposes metrics through a FastAPI endpoint at GET /metrics. This endpoint uses generate_latest() from the prometheus_client library to return the current metric snapshot in Prometheus exposition format.
The worker implementation in hindsight_api_slim/hindsight_api/worker/main.py (lines 78-86) defines the endpoint as follows:
# hindsight_api_slim/hindsight_api/worker/main.py
@app.get(
"/metrics",
summary="Prometheus metrics endpoint",
description="Exports metrics in Prometheus format for scraping",
tags=["Monitoring"],
)
async def metrics_endpoint():
"""Return Prometheus metrics."""
metrics_data = generate_latest()
return Response(content=metrics_data, media_type=CONTENT_TYPE_LATEST)
The API server provides an identical implementation in hindsight_api_slim/hindsight_api/api/http.py (lines 21-14):
# hindsight_api_slim/hindsight_api/api/http.py
@app.get(
"/metrics",
summary="Prometheus metrics endpoint",
description="Exports metrics in Prometheus format for scraping",
tags=["Monitoring"],
)
async def metrics_endpoint():
"""Return Prometheus metrics."""
from prometheus_client import CONTENT_TYPE_LATEST, generate_latest
metrics_data = generate_latest()
return Response(content=metrics_data, media_type=CONTENT_TYPE_LATEST)
When the returned PrometheusMetricReader is stored in app.state.prometheus_reader, the endpoint can generate fresh snapshots of all collected telemetry data.
Configuring Prometheus Scraping
The repository includes a pre-configured Prometheus job definition in scripts/dev/monitoring/prometheus.yml (lines 24-30) that targets the Hindsight API endpoint:
# scripts/dev/monitoring/prometheus.yml
scrape_configs:
- job_name: 'hindsight-api'
static_configs:
- targets: ['host.docker.internal:8888']
metrics_path: '/metrics'
scrape_interval: 5s
This configuration scrapes the /metrics endpoint every 5 seconds. The host.docker.internal reference allows the containerized Prometheus instance to reach the Hindsight service running on the host machine.
Running the Complete Monitoring Stack
Hindsight provides a Docker Compose configuration that deploys the Grafana OTel LGTM (Loki, Grafana, Tempo, Mimir/Prometheus) stack with pre-built dashboards. The compose file at scripts/dev/monitoring/docker-compose.yaml (lines 17-25) mounts the Prometheus configuration and three specialized dashboards:
# scripts/dev/monitoring/docker-compose.yaml
services:
grafana-lgtm:
image: grafana/otel-lgtm:latest
ports:
- "3000:3000"
environment:
- GF_AUTH_ANONYMOUS_ENABLED=true
- GF_AUTH_ANONYMOUS_ORG_ROLE=Admin
- GF_AUTH_DISABLE_LOGIN_FORM=true
volumes:
- ./prometheus.yml:/otel-lgtm/prometheus.yaml:ro
- ../../../monitoring/grafana/dashboards/hindsight-operations.json:/otel-lgtm/hindsight-operations.json:ro
- ../../../monitoring/grafana/dashboards/hindsight-llm.json:/otel-lgtm/hindsight-llm.json:ro
- ../../../monitoring/grafana/dashboards/hindsight-api-service.json:/otel-lgtm/hindsight-api-service.json:ro
- ./grafana-dashboards.yaml:/otel-lgtm/grafana/conf/provisioning/dashboards/grafana-dashboards.yaml:ro
The mounted dashboards located in monitoring/grafana/dashboards/ provide visualizations for:
- Operation latency percentiles (
hindsight.operation.duration) - LLM call latency and token usage (
hindsight.llm.duration) - HTTP request latency (
hindsight.http.duration) - Process and database pool metrics (auto-collected by OpenTelemetry)
Integrating Custom Metrics in Your Application
To add Prometheus monitoring to a custom Hindsight-based service, import the metrics module and follow this pattern:
from hindsight_api.metrics import initialize_metrics, get_meter
from fastapi import FastAPI, Response
from prometheus_client import CONTENT_TYPE_LATEST, generate_latest
app = FastAPI()
# Initialise metrics once during startup
prom_reader = initialize_metrics(service_name="my-custom-service", service_version="0.1.0")
app.state.prometheus_reader = prom_reader
# Expose the Prometheus endpoint
@app.get("/metrics")
async def prometheus_metrics():
return Response(content=generate_latest(),
media_type=CONTENT_TYPE_LATEST)
After initialization, use get_meter() to create custom instruments:
from hindsight_api.metrics import get_meter
from contextlib import contextmanager
import time
meter = get_meter()
request_latency = meter.create_histogram(
name="my_custom.http.duration",
description="Latency of my custom HTTP handler",
unit="s",
)
@contextmanager
def record_latency():
start = time.time()
try:
yield
finally:
request_latency.record(time.time() - start, {"handler": "my_endpoint"})
Summary
- Initialize once: Call
initialize_metrics()at application startup to create thePrometheusMetricReaderand configure histogram buckets for operation, LLM, and HTTP durations. - Expose endpoint: Both the worker (
worker/main.py) and API (api/http.py) serve Prometheus-formatted metrics atGET /metricsusinggenerate_latest(). - Scrape configuration: Use the provided
prometheus.ymlto configure scraping intervals and targets, or integrate with existing Prometheus infrastructure. - Visualize: Deploy the LGTM stack via
docker-compose.yamlto access pre-built Grafana dashboards for operations, LLM metrics, and API service monitoring without additional configuration.
Frequently Asked Questions
What metrics does Hindsight expose to Prometheus?
Hindsight exposes three primary histogram metrics: hindsight.operation.duration for general operations, hindsight.llm.duration for language model call latency, and hindsight.http.duration for HTTP request timing. Additionally, OpenTelemetry auto-instrumentation collects process metrics, database connection pool statistics, and custom counters defined throughout the application.
Do I need to modify the Hindsight source code to enable Prometheus monitoring?
No modifications are required. The worker and API implementations in hindsight_api_slim/hindsight_api/worker/main.py and hindsight_api_slim/hindsight_api/api/http.py already include the initialize_metrics() call and /metrics endpoint. Simply start the application and configure your Prometheus instance to scrape the endpoint.
How do I customize the histogram buckets for my specific deployment?
Edit the bucket definitions in hindsight_api_slim/hindsight_api/metrics.py before calling initialize_metrics(). The file defines DURATION_BUCKETS, LLM_DURATION_BUCKETS, and HTTP_DURATION_BUCKETS as constants used in the ExplicitBucketHistogramAggregation views. Adjust these arrays to match your specific latency requirements and SLOs.
Can I use my existing Prometheus instance instead of the provided LGTM stack?
Yes. Configure your existing Prometheus server to scrape the Hindsight /metrics endpoint using the job definition from scripts/dev/monitoring/prometheus.yml. The metrics use standard Prometheus exposition format, so any compatible scraper can consume them. Import the dashboard JSON files from monitoring/grafana/dashboards/ into your existing Grafana instance to visualize the data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →