How FreeLLM API Admin Dashboard Analytics Track p95 Latency and TTFT
The FreeLLM API admin dashboard calculates p95 latency using a nearest-rank percentile query on the SQLite requests table and computes average time-to-first-token (TTFT) via AVG(ttfb_ms), serving both metrics through the /api/analytics/by-platform endpoint to the React frontend.
The tashfeenahmed/freellmapi repository implements a high-precision analytics pipeline for monitoring LLM provider performance. The server-side analytics router queries raw request logs to compute percentile latencies and TTFT averages, while the React-based admin UI renders these metrics in real-time stat cards and per-provider data tables.
Server-Side Metric Calculation
The backend logic resides in server/src/routes/analytics.ts, where two dedicated SQL aggregation strategies handle latency and TTFT separately.
Calculating p95 Latency via Nearest-Rank Method
Instead of using a built-in percentile function, the router implements a nearest-rank algorithm to find the 95th percentile of latency_ms values. For each provider, it first counts non-null latency rows, then executes a targeted query with a calculated offset:
// Inside /api/analytics/by-platform (server/src/routes/analytics.ts)
const p95Stmt = db.prepare(`
SELECT r.latency_ms FROM requests r
LEFT JOIN api_keys k ON k.id = r.key_id
WHERE r.created_at >= ? AND r.platform = ? AND ${ENDPOINT_ID_SQL} = ?
AND r.latency_ms IS NOT NULL
ORDER BY r.latency_ms ASC
LIMIT 1 OFFSET ?
`);
The offset is calculated as Math.floor((latencyCount - 1) * 0.95), selecting the row at the 95th percentile position. The router rounds the resulting latency_ms to an integer and returns it as p95LatencyMs【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/routes/analytics.ts#L1520-L1535】.
Averaging Time-to-First-Token (TTFT)
TTFT—measuring the delay until the first token streams back—is computed as a simple average across the ttfb_ms column. The router uses AVG(ttfb_ms) in the SQL query for both summary and per-platform endpoints:
const ttfbRow = db.prepare(`
SELECT AVG(ttfb_ms) as avg_ttfb_ms FROM requests
WHERE created_at >= ? AND ttfb_ms IS NOT NULL
`).get(since) as { avg_ttfb_ms: number | null } | undefined;
const avgTtfbMs = ttfbRow?.avg_ttfb_ms != null
? Math.round(ttfbRow.avg_ttfb_ms)
: null;
This value is returned as avgTtfbMs and rounded to the nearest integer【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/routes/analytics.ts#L63‑L68】【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/routes/analytics.ts#L150‑L154】.
Handling Data Pruning and Null States
When the raw requests table is pruned (data older than the retention window), the router returns null for both p95LatencyMs and avgTtfbMs. This prevents the UI from displaying misleading zero values and allows the frontend to render a placeholder instead.
Frontend Data Flow and Visualization
The admin dashboard consumes these metrics through the React component defined in client/src/pages/AnalyticsPage.tsx.
Fetching Analytics in the React Dashboard
The frontend uses a TanStack Query hook to fetch per-platform analytics:
const { data: byPlatform = [] } = useQuery({
queryKey: ['analytics', 'by-platform', range],
queryFn: () => apiFetch<ByPlatformRow[]>(
`/api/analytics/by-platform?range=${range}`
),
});
The ByPlatformRow TypeScript interface explicitly types the latency fields as potentially null:
{
platform: string;
providerId: string;
endpoint?: string;
p95LatencyMs: number | null; // 95th percentile latency
avgTtfbMs: number | null; // Average TTFT
// ... other fields
}
Rendering Latency Cards and Provider Tables
The UI renders these values in dedicated Stat components and tabular views, applying a fallback display when metrics are unavailable:
{/* Summary stat cards */}
<Stat
icon={Clock}
label={t('analytics.p95Latency')}
value={p95Value} // "—" if null, otherwise "123 ms"
/>
<Stat
icon={Zap}
label={t('analytics.avgTtft')}
value={ttftValue} // "—" if null, otherwise "85 ms"
/>
{/* Per-provider table cells */}
<TableCell className="text-right tabular-nums">
{p.p95LatencyMs != null ? `${p.p95LatencyMs} ms` : '—'}
</TableCell>
<TableCell className="text-right tabular-nums">
{p.avgTtfbMs != null ? `${p.avgTtfbMs} ms` : '—'}
</TableCell>
This ensures that pruned or empty datasets display neutral dashes rather than invalid zeros【/cache/repos/github.com/tashfeenahmed/freellmapi/main/client/src/pages/AnalyticsPage.tsx#L31‑L38】【/cache/repos/github.com/tashfeenahmed/freellmapi/main/client/src/pages/AnalyticsPage.tsx#L84‑L92】.
Database Architecture and Query Optimization
The analytics router maintains per-provider granularity by grouping results using both the platform slug and a custom endpoint identifier (ENDPOINT_ID_SQL). This distinction ensures that multiple custom relays for the same provider do not conflate their metrics【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/routes/analytics.ts#L10‑L16】.
While summary statistics like total request counts and token volumes are optimized through the request_hourly aggregate table, latency-specific metrics necessarily query the raw requests table because the hourly bucket does not retain per-request latency distributions.
Summary
- p95 Latency: Calculated via nearest-rank algorithm on
latency_mswithOFFSET ⌊(count‑1) × 0.95⌋inserver/src/routes/analytics.ts. - TTFT: Computed as
ROUND(AVG(ttfb_ms))from non-null rows in therequeststable. - Data Integrity: Returns
nullfor pruned windows, allowing the UI to show "—" placeholders. - Per-Provider Tracking: Uses composite grouping by platform and endpoint ID to isolate custom relays.
- Frontend: React dashboard in
client/src/pages/AnalyticsPage.tsxconsumes/api/analytics/by-platformand/api/analytics/summaryto render stat cards and tables.
Frequently Asked Questions
How does FreeLLM API calculate p95 latency without using a percentile function?
The analytics router implements a nearest-rank method manually. It counts valid latency rows, calculates the target index using Math.floor((count - 1) * 0.95), and uses a prepared SQLite statement with LIMIT 1 OFFSET ? to select the exact row at the 95th percentile position. This approach avoids SQLite window function dependencies while maintaining statistical accuracy.
What happens to TTFT and latency metrics when old requests are pruned?
When the retention window prunes old records from the requests table, the aggregation queries return null for both avgTtfbMs and p95LatencyMs. The React frontend detects these null values and renders an em-dash ("—") in the stat cards and table cells to indicate unavailable data rather than displaying a misleading zero.
Can the dashboard track p50 (median) latency in addition to p95?
The current implementation specifically tracks the 95th percentile for latency bounds. However, the nearest-rank algorithm used in server/src/routes/analytics.ts accepts any decimal multiplier (0.50 for median). Extending the dashboard to support p50 would require duplicating the p95 query logic with a 0.50 offset factor and adding the corresponding frontend display fields.
How does the analytics router distinguish between different providers for metrics?
The router groups metrics by both the platform identifier (e.g., "openai", "anthropic") and a derived endpoint ID calculated via ENDPOINT_ID_SQL. This ensures that custom proxy endpoints or multiple configurations of the same provider are tracked as separate rows in the analytics output, preventing metric conflation across distinct API gateways.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →