# How FreeLLM API Admin Dashboard Analytics Track p95 Latency and TTFT

> Discover how FreeLLM API analytics track p95 latency and TTFT using nearest-rank percentile queries and AVG functions on the SQLite requests table, serving real-time data to the React frontend.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: internals
- Published: 2026-08-31

---

**The FreeLLM API admin dashboard calculates p95 latency using a nearest-rank percentile query on the SQLite `requests` table and computes average time-to-first-token (TTFT) via `AVG(ttfb_ms)`, serving both metrics through the `/api/analytics/by-platform` endpoint to the React frontend.**

The `tashfeenahmed/freellmapi` repository implements a high-precision analytics pipeline for monitoring LLM provider performance. The server-side analytics router queries raw request logs to compute percentile latencies and TTFT averages, while the React-based admin UI renders these metrics in real-time stat cards and per-provider data tables.

## Server-Side Metric Calculation

The backend logic resides in [`server/src/routes/analytics.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/analytics.ts), where two dedicated SQL aggregation strategies handle latency and TTFT separately.

### Calculating p95 Latency via Nearest-Rank Method

Instead of using a built-in percentile function, the router implements a **nearest-rank algorithm** to find the 95th percentile of `latency_ms` values. For each provider, it first counts non-null latency rows, then executes a targeted query with a calculated offset:

```typescript
// Inside /api/analytics/by-platform (server/src/routes/analytics.ts)
const p95Stmt = db.prepare(`
  SELECT r.latency_ms FROM requests r
  LEFT JOIN api_keys k ON k.id = r.key_id
  WHERE r.created_at >= ? AND r.platform = ? AND ${ENDPOINT_ID_SQL} = ?
    AND r.latency_ms IS NOT NULL
  ORDER BY r.latency_ms ASC
  LIMIT 1 OFFSET ?
`);

```

The offset is calculated as `Math.floor((latencyCount - 1) * 0.95)`, selecting the row at the 95th percentile position. The router rounds the resulting `latency_ms` to an integer and returns it as `p95LatencyMs`【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/routes/analytics.ts#L1520-L1535】.

### Averaging Time-to-First-Token (TTFT)

TTFT—measuring the delay until the first token streams back—is computed as a simple average across the `ttfb_ms` column. The router uses `AVG(ttfb_ms)` in the SQL query for both summary and per-platform endpoints:

```typescript
const ttfbRow = db.prepare(`
  SELECT AVG(ttfb_ms) as avg_ttfb_ms FROM requests
  WHERE created_at >= ? AND ttfb_ms IS NOT NULL
`).get(since) as { avg_ttfb_ms: number | null } | undefined;

const avgTtfbMs = ttfbRow?.avg_ttfb_ms != null
  ? Math.round(ttfbRow.avg_ttfb_ms)
  : null;

```

This value is returned as `avgTtfbMs` and rounded to the nearest integer【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/routes/analytics.ts#L63‑L68】【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/routes/analytics.ts#L150‑L154】.

### Handling Data Pruning and Null States

When the raw `requests` table is pruned (data older than the retention window), the router returns `null` for both `p95LatencyMs` and `avgTtfbMs`. This prevents the UI from displaying misleading zero values and allows the frontend to render a placeholder instead.

## Frontend Data Flow and Visualization

The admin dashboard consumes these metrics through the React component defined in [`client/src/pages/AnalyticsPage.tsx`](https://github.com/tashfeenahmed/freellmapi/blob/main/client/src/pages/AnalyticsPage.tsx).

### Fetching Analytics in the React Dashboard

The frontend uses a TanStack Query hook to fetch per-platform analytics:

```tsx
const { data: byPlatform = [] } = useQuery({
  queryKey: ['analytics', 'by-platform', range],
  queryFn: () => apiFetch<ByPlatformRow[]>(
    `/api/analytics/by-platform?range=${range}`
  ),
});

```

The `ByPlatformRow` TypeScript interface explicitly types the latency fields as potentially null:

```typescript
{
  platform: string;
  providerId: string;
  endpoint?: string;
  p95LatencyMs: number | null;   // 95th percentile latency
  avgTtfbMs: number | null;      // Average TTFT
  // ... other fields
}

```

### Rendering Latency Cards and Provider Tables

The UI renders these values in dedicated `Stat` components and tabular views, applying a fallback display when metrics are unavailable:

```tsx
{/* Summary stat cards */}
<Stat
  icon={Clock}
  label={t('analytics.p95Latency')}
  value={p95Value}  // "—" if null, otherwise "123 ms"
/>
<Stat
  icon={Zap}
  label={t('analytics.avgTtft')}
  value={ttftValue}  // "—" if null, otherwise "85 ms"
/>

{/* Per-provider table cells */}
<TableCell className="text-right tabular-nums">
  {p.p95LatencyMs != null ? `${p.p95LatencyMs} ms` : '—'}
</TableCell>
<TableCell className="text-right tabular-nums">
  {p.avgTtfbMs != null ? `${p.avgTtfbMs} ms` : '—'}
</TableCell>

```

This ensures that pruned or empty datasets display neutral dashes rather than invalid zeros【/cache/repos/github.com/tashfeenahmed/freellmapi/main/client/src/pages/AnalyticsPage.tsx#L31‑L38】【/cache/repos/github.com/tashfeenahmed/freellmapi/main/client/src/pages/AnalyticsPage.tsx#L84‑L92】.

## Database Architecture and Query Optimization

The analytics router maintains **per-provider granularity** by grouping results using both the platform slug and a custom endpoint identifier (`ENDPOINT_ID_SQL`). This distinction ensures that multiple custom relays for the same provider do not conflate their metrics【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/routes/analytics.ts#L10‑L16】.

While summary statistics like total request counts and token volumes are optimized through the `request_hourly` aggregate table, **latency-specific metrics necessarily query the raw `requests` table** because the hourly bucket does not retain per-request latency distributions.

## Summary

- **p95 Latency**: Calculated via nearest-rank algorithm on `latency_ms` with `OFFSET ⌊(count‑1) × 0.95⌋` in [`server/src/routes/analytics.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/analytics.ts).
- **TTFT**: Computed as `ROUND(AVG(ttfb_ms))` from non-null rows in the `requests` table.
- **Data Integrity**: Returns `null` for pruned windows, allowing the UI to show "—" placeholders.
- **Per-Provider Tracking**: Uses composite grouping by platform and endpoint ID to isolate custom relays.
- **Frontend**: React dashboard in [`client/src/pages/AnalyticsPage.tsx`](https://github.com/tashfeenahmed/freellmapi/blob/main/client/src/pages/AnalyticsPage.tsx) consumes `/api/analytics/by-platform` and `/api/analytics/summary` to render stat cards and tables.

## Frequently Asked Questions

### How does FreeLLM API calculate p95 latency without using a percentile function?

The analytics router implements a nearest-rank method manually. It counts valid latency rows, calculates the target index using `Math.floor((count - 1) * 0.95)`, and uses a prepared SQLite statement with `LIMIT 1 OFFSET ?` to select the exact row at the 95th percentile position. This approach avoids SQLite window function dependencies while maintaining statistical accuracy.

### What happens to TTFT and latency metrics when old requests are pruned?

When the retention window prunes old records from the `requests` table, the aggregation queries return `null` for both `avgTtfbMs` and `p95LatencyMs`. The React frontend detects these null values and renders an em-dash ("—") in the stat cards and table cells to indicate unavailable data rather than displaying a misleading zero.

### Can the dashboard track p50 (median) latency in addition to p95?

The current implementation specifically tracks the 95th percentile for latency bounds. However, the nearest-rank algorithm used in [`server/src/routes/analytics.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/analytics.ts) accepts any decimal multiplier (0.50 for median). Extending the dashboard to support p50 would require duplicating the p95 query logic with a 0.50 offset factor and adding the corresponding frontend display fields.

### How does the analytics router distinguish between different providers for metrics?

The router groups metrics by both the platform identifier (e.g., "openai", "anthropic") and a derived endpoint ID calculated via `ENDPOINT_ID_SQL`. This ensures that custom proxy endpoints or multiple configurations of the same provider are tracked as separate rows in the analytics output, preventing metric conflation across distinct API gateways.