# How to Monitor Hindsight Using Prometheus Metrics: A Complete Setup Guide

> Monitor Hindsight using Prometheus metrics with this complete setup guide. Learn how to leverage built-in OpenTelemetry instrumentation for powerful insights.

- Repository: [vectorize-io/hindsight](https://github.com/vectorize-io/hindsight)
- Tags: how-to-guide
- Published: 2026-03-13

---

**Hindsight exposes Prometheus-compatible metrics through built-in OpenTelemetry instrumentation, providing a `/metrics` endpoint that can be scraped by any Prometheus server or the included Grafana LGTM stack.**

The `vectorize-io/hindsight` repository ships with a complete observability layer that leverages OpenTelemetry to generate Prometheus metrics. This instrumentation tracks operation latency, LLM call durations, HTTP request timing, and system metrics without requiring external dependencies beyond the standard FastAPI application stack.

## Initializing the Metrics Collector

The monitoring stack begins with `initialize_metrics()` in [`hindsight_api_slim/hindsight_api/metrics.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api_slim/hindsight_api/metrics.py) (lines 94-52). This function creates a **PrometheusMetricReader**, configures custom histogram views for different latency types, and establishes a global meter provider.

```python

# hindsight_api_slim/hindsight_api/metrics.py

def initialize_metrics(service_name: str = "hindsight-api", service_version: str = "1.0.0"):
    """
    Initialise OpenTelemetry metrics with a Prometheus exporter.
    Returns a PrometheusMetricReader for exposing the /metrics endpoint.
    """
    global _meter
    resource = Resource.create(
        {"service.name": service_name, "service.version": service_version}
    )
    prometheus_reader = PrometheusMetricReader()
    # Custom histogram buckets for operation, LLM and HTTP latency

    duration_view = View(
        instrument_name="hindsight.operation.duration",
        aggregation=ExplicitBucketHistogramAggregation(boundaries=DURATION_BUCKETS),
    )
    llm_duration_view = View(
        instrument_name="hindsight.llm.duration",
        aggregation=ExplicitBucketHistogramAggregation(boundaries=LLM_DURATION_BUCKETS),
    )
    http_duration_view = View(
        instrument_name="hindsight.http.duration",
        aggregation=ExplicitBucketHistogramAggregation(boundaries=HTTP_DURATION_BUCKETS),
    )
    provider = MeterProvider(
        resource=resource,
        metric_readers=[prometheus_reader],
        views=[duration_view, llm_duration_view, http_duration_view],
    )
    metrics.set_meter_provider(provider)
    _meter = metrics.get_meter(__name__)
    return prometheus_reader

```

Both the **worker** and **API server** invoke this function during startup. The worker calls it in [`worker/main.py`](https://github.com/vectorize-io/hindsight/blob/main/worker/main.py) (lines 48-52), while the API initializes it during app creation in [`api/http.py`](https://github.com/vectorize-io/hindsight/blob/main/api/http.py).

## Exposing the /metrics HTTP Endpoint

Once initialized, the application exposes metrics through a FastAPI endpoint at `GET /metrics`. This endpoint uses `generate_latest()` from the `prometheus_client` library to return the current metric snapshot in Prometheus exposition format.

The worker implementation in [`hindsight_api_slim/hindsight_api/worker/main.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api_slim/hindsight_api/worker/main.py) (lines 78-86) defines the endpoint as follows:

```python

# hindsight_api_slim/hindsight_api/worker/main.py

@app.get(
    "/metrics",
    summary="Prometheus metrics endpoint",
    description="Exports metrics in Prometheus format for scraping",
    tags=["Monitoring"],
)
async def metrics_endpoint():
    """Return Prometheus metrics."""
    metrics_data = generate_latest()
    return Response(content=metrics_data, media_type=CONTENT_TYPE_LATEST)

```

The API server provides an identical implementation in [`hindsight_api_slim/hindsight_api/api/http.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api_slim/hindsight_api/api/http.py) (lines 21-14):

```python

# hindsight_api_slim/hindsight_api/api/http.py

@app.get(
    "/metrics",
    summary="Prometheus metrics endpoint",
    description="Exports metrics in Prometheus format for scraping",
    tags=["Monitoring"],
)
async def metrics_endpoint():
    """Return Prometheus metrics."""
    from prometheus_client import CONTENT_TYPE_LATEST, generate_latest
    metrics_data = generate_latest()
    return Response(content=metrics_data, media_type=CONTENT_TYPE_LATEST)

```

When the returned `PrometheusMetricReader` is stored in `app.state.prometheus_reader`, the endpoint can generate fresh snapshots of all collected telemetry data.

## Configuring Prometheus Scraping

The repository includes a pre-configured Prometheus job definition in [`scripts/dev/monitoring/prometheus.yml`](https://github.com/vectorize-io/hindsight/blob/main/scripts/dev/monitoring/prometheus.yml) (lines 24-30) that targets the Hindsight API endpoint:

```yaml

# scripts/dev/monitoring/prometheus.yml

scrape_configs:
  - job_name: 'hindsight-api'
    static_configs:
      - targets: ['host.docker.internal:8888']
    metrics_path: '/metrics'
    scrape_interval: 5s

```

This configuration scrapes the `/metrics` endpoint every 5 seconds. The `host.docker.internal` reference allows the containerized Prometheus instance to reach the Hindsight service running on the host machine.

## Running the Complete Monitoring Stack

Hindsight provides a Docker Compose configuration that deploys the **Grafana OTel LGTM** (Loki, Grafana, Tempo, Mimir/Prometheus) stack with pre-built dashboards. The compose file at [`scripts/dev/monitoring/docker-compose.yaml`](https://github.com/vectorize-io/hindsight/blob/main/scripts/dev/monitoring/docker-compose.yaml) (lines 17-25) mounts the Prometheus configuration and three specialized dashboards:

```yaml

# scripts/dev/monitoring/docker-compose.yaml

services:
  grafana-lgtm:
    image: grafana/otel-lgtm:latest
    ports:
      - "3000:3000"
    environment:
      - GF_AUTH_ANONYMOUS_ENABLED=true
      - GF_AUTH_ANONYMOUS_ORG_ROLE=Admin
      - GF_AUTH_DISABLE_LOGIN_FORM=true
    volumes:
      - ./prometheus.yml:/otel-lgtm/prometheus.yaml:ro
      - ../../../monitoring/grafana/dashboards/hindsight-operations.json:/otel-lgtm/hindsight-operations.json:ro
      - ../../../monitoring/grafana/dashboards/hindsight-llm.json:/otel-lgtm/hindsight-llm.json:ro
      - ../../../monitoring/grafana/dashboards/hindsight-api-service.json:/otel-lgtm/hindsight-api-service.json:ro
      - ./grafana-dashboards.yaml:/otel-lgtm/grafana/conf/provisioning/dashboards/grafana-dashboards.yaml:ro

```

The mounted dashboards located in `monitoring/grafana/dashboards/` provide visualizations for:
- **Operation latency percentiles** (`hindsight.operation.duration`)
- **LLM call latency and token usage** (`hindsight.llm.duration`)
- **HTTP request latency** (`hindsight.http.duration`)
- **Process and database pool metrics** (auto-collected by OpenTelemetry)

## Integrating Custom Metrics in Your Application

To add Prometheus monitoring to a custom Hindsight-based service, import the metrics module and follow this pattern:

```python
from hindsight_api.metrics import initialize_metrics, get_meter
from fastapi import FastAPI, Response
from prometheus_client import CONTENT_TYPE_LATEST, generate_latest

app = FastAPI()

# Initialise metrics once during startup

prom_reader = initialize_metrics(service_name="my-custom-service", service_version="0.1.0")
app.state.prometheus_reader = prom_reader

# Expose the Prometheus endpoint

@app.get("/metrics")
async def prometheus_metrics():
    return Response(content=generate_latest(),
                    media_type=CONTENT_TYPE_LATEST)

```

After initialization, use `get_meter()` to create custom instruments:

```python
from hindsight_api.metrics import get_meter
from contextlib import contextmanager
import time

meter = get_meter()
request_latency = meter.create_histogram(
    name="my_custom.http.duration",
    description="Latency of my custom HTTP handler",
    unit="s",
)

@contextmanager
def record_latency():
    start = time.time()
    try:
        yield
    finally:
        request_latency.record(time.time() - start, {"handler": "my_endpoint"})

```

## Summary

- **Initialize once**: Call `initialize_metrics()` at application startup to create the `PrometheusMetricReader` and configure histogram buckets for operation, LLM, and HTTP durations.
- **Expose endpoint**: Both the worker ([`worker/main.py`](https://github.com/vectorize-io/hindsight/blob/main/worker/main.py)) and API ([`api/http.py`](https://github.com/vectorize-io/hindsight/blob/main/api/http.py)) serve Prometheus-formatted metrics at `GET /metrics` using `generate_latest()`.
- **Scrape configuration**: Use the provided [`prometheus.yml`](https://github.com/vectorize-io/hindsight/blob/main/prometheus.yml) to configure scraping intervals and targets, or integrate with existing Prometheus infrastructure.
- **Visualize**: Deploy the LGTM stack via [`docker-compose.yaml`](https://github.com/vectorize-io/hindsight/blob/main/docker-compose.yaml) to access pre-built Grafana dashboards for operations, LLM metrics, and API service monitoring without additional configuration.

## Frequently Asked Questions

### What metrics does Hindsight expose to Prometheus?

Hindsight exposes three primary histogram metrics: `hindsight.operation.duration` for general operations, `hindsight.llm.duration` for language model call latency, and `hindsight.http.duration` for HTTP request timing. Additionally, OpenTelemetry auto-instrumentation collects process metrics, database connection pool statistics, and custom counters defined throughout the application.

### Do I need to modify the Hindsight source code to enable Prometheus monitoring?

No modifications are required. The worker and API implementations in [`hindsight_api_slim/hindsight_api/worker/main.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api_slim/hindsight_api/worker/main.py) and [`hindsight_api_slim/hindsight_api/api/http.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api_slim/hindsight_api/api/http.py) already include the `initialize_metrics()` call and `/metrics` endpoint. Simply start the application and configure your Prometheus instance to scrape the endpoint.

### How do I customize the histogram buckets for my specific deployment?

Edit the bucket definitions in [`hindsight_api_slim/hindsight_api/metrics.py`](https://github.com/vectorize-io/hindsight/blob/main/hindsight_api_slim/hindsight_api/metrics.py) before calling `initialize_metrics()`. The file defines `DURATION_BUCKETS`, `LLM_DURATION_BUCKETS`, and `HTTP_DURATION_BUCKETS` as constants used in the `ExplicitBucketHistogramAggregation` views. Adjust these arrays to match your specific latency requirements and SLOs.

### Can I use my existing Prometheus instance instead of the provided LGTM stack?

Yes. Configure your existing Prometheus server to scrape the Hindsight `/metrics` endpoint using the job definition from [`scripts/dev/monitoring/prometheus.yml`](https://github.com/vectorize-io/hindsight/blob/main/scripts/dev/monitoring/prometheus.yml). The metrics use standard Prometheus exposition format, so any compatible scraper can consume them. Import the dashboard JSON files from `monitoring/grafana/dashboards/` into your existing Grafana instance to visualize the data.