# How llmfit-core bench.rs Measures Tokens-per-Second Without Blocking the TUI

> Discover how llmfit-core bench.rs measures tokens-per-second without blocking the TUI. Learn how synchronous HTTP benchmarks achieve responsiveness through background threads and callbacks.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-09-11

---

**The [`bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/bench.rs) module performs synchronous HTTP benchmarks against Ollama and OpenAI-compatible endpoints while keeping the TUI responsive by running all blocking I/O in a background thread and communicating progress via callbacks.**

The `llmfit-core` crate provides high-precision throughput measurement for LLM inference servers, calculating tokens-per-second (TPS) for both Ollama and OpenAI-compatible APIs. Because the benchmark logic relies on synchronous HTTP calls using `ureq`, the architecture decouples the blocking measurement code from the interactive terminal user interface (TUI) through strategic thread delegation. This design allows real-time progress updates while maintaining accurate wall-clock and server-reported timing metrics.

## TPS Calculation Strategies by Endpoint Type

The benchmark module handles two distinct API families with different timing metadata availability. Each requires a specialized approach to calculate throughput accurately.

### Ollama Server Timing (Native Metrics)

For Ollama endpoints (`POST /api/generate`), the server provides precise per-request timing fields that enable exact TPS calculation without relying solely on wall-clock time. According to the source code in [`llmfit-core/src/bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/bench.rs) (lines 24-31), the response body contains `eval_count` (generated tokens) and `eval_duration` (nanoseconds).

The calculation at lines 97-102 converts nanoseconds to seconds and computes:

```rust
let tps = if let (Some(eval_count), Some(eval_dur)) =
    (resp_body.eval_count, resp_body.eval_duration)
{
    // tokens / seconds
    eval_count as f64 / (eval_dur as f64 / 1_000_000_000.0)
} else if output_tokens > 0 {
    // fallback to wall‑clock
    output_tokens as f64 / total_wall.as_secs_f64()
} else {
    0.0
};

```

The **time-to-first-token (TTFT)** metric extracts `prompt_eval_duration` when available, providing insight into prompt processing latency. If the server omits timing metadata, the code falls back to `total_wall` measurement.

### OpenAI-Compatible Approximation (Wall-Clock Fallback)

For OpenAI-compatible endpoints (`POST /v1/chat/completions`)—including vLLM, Ferrum, MLX, and llama-cpp—the API returns usage statistics but lacks per-token timing granularity. As implemented in lines 61-70 and 133-138 of [`bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/bench.rs), the benchmark approximates TPS using wall-clock duration divided by `completion_tokens` from the usage field.

```rust
let total_wall = start.elapsed();      // wall‑clock duration
let output_tokens = usage.completion_tokens;
let tps = if output_tokens > 0 && total_wall.as_secs_f64() > 0.0 {
    // tokens / seconds (no per‑token timing available)
    output_tokens as f64 / total_wall.as_secs_f64()
} else {
    0.0
};

```

Because these APIs do not expose streaming timing data, **TTFT cannot be measured** and is recorded as `None`.

## Architecture: Isolating Blocking I/O from the TUI

The benchmark functions `bench_ollama` and `bench_openai_compat` are inherently blocking operations. They execute synchronous HTTP requests via `ureq` and wait for complete JSON responses before calculating metrics. The TUI remains interactive through a thread-based concurrency model rather than async/await.

### Synchronous Benchmark Functions in bench.rs

The core measurement logic in [`llmfit-core/src/bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/bench.rs) performs blocking network calls to ensure accurate timing without the overhead of runtime scheduling. This approach eliminates jitter from async task switching but would freeze the interface if executed on the main thread.

### Background Thread Delegation in the TUI Layer

The TUI implementation in [`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs) and [`llmfit-tui/src/tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_app.rs) spawns the benchmark in a separate thread, passing a progress callback that allows the UI to update without blocking on network latency.

In [`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs), the application launches the benchmark using `std::thread::spawn`:

```rust
let progress = |run: usize, total: usize| {
    // UI updates (progress bar, status line) happen here
    app.update_progress(run, total);
};

std::thread::spawn(move || {
    let result = bench::benchmark_target(&target, runs, &progress);
    // Send the result back to the UI thread via a channel or event queue
    tx.send(result).unwrap();
});

```

The `benchmark_target` function invokes the provided `on_progress` closure before each warm-up and measurement iteration. This callback mechanism ensures the progress bar advances while the HTTP request blocks the background thread. The main event loop continues processing user input, maintaining a responsive interface even during high-latency API calls.

## Implementation Details and Code Walkthrough

The three-file architecture separates concerns between measurement accuracy and UI responsiveness:

- **[`llmfit-core/src/bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/bench.rs)**: Contains the concrete benchmark implementation, HTTP client logic, and TPS calculations. This file remains agnostic to the UI layer, exposing only the `benchmark_target` function with a generic progress callback.

- **[`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs)**: Entry point responsible for thread lifecycle management and inter-thread communication via channels.

- **[`llmfit-tui/src/tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_app.rs)**: Handles UI state updates and event processing, invoking `bench::benchmark_target` with a closure that mutates the application state safely.

This pattern allows [`bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/bench.rs) to focus on precise timing measurement while the TUI layer handles the concurrency model. The synchronous design of the benchmark functions simplifies error handling and ensures consistent wall-clock measurements without runtime interference.

## Summary

- **Ollama endpoints** utilize server-reported `eval_duration` and `eval_count` for exact TPS calculation, with fallback to wall-clock timing.
- **OpenAI-compatible endpoints** rely on wall-clock duration divided by `completion_tokens` due to API limitations, with TTFT unavailable.
- **Blocking I/O** is confined to background threads in [`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs), preventing TUI freezing.
- **Progress callbacks** provide real-time UI updates without async complexity, bridging the synchronous benchmark and interactive interface.

## Frequently Asked Questions

### How does llmfit-core handle missing timing data from Ollama?

When Ollama responses omit `eval_duration` or `eval_count` fields, the code at lines 97-102 of [`bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/bench.rs) falls back to wall-clock measurement. It calculates TPS using `output_tokens as f64 / total_wall.as_secs_f64()`, ensuring the benchmark completes even with incomplete server metadata.

### Why can't time-to-first-token (TTFT) be measured for OpenAI-compatible endpoints?

The OpenAI-compatible API specification returns aggregate usage statistics (`prompt_tokens`, `completion_tokens`) but does not expose per-token timing or streaming metadata in the standard REST response. Without access to when the first token was generated versus when the request was sent, the benchmark records TTFT as `None` for these providers.

### What HTTP client does bench.rs use for synchronous requests?

The benchmark module uses `ureq`, a synchronous HTTP client library for Rust. This choice eliminates async runtime overhead and provides deterministic timing measurements, though it requires the TUI to wrap calls in `std::thread::spawn` to maintain responsiveness.

### How does the progress callback communicate with the main TUI thread?

The `on_progress` closure passed to `benchmark_target` captures a sender channel or application state reference from the spawning thread. When invoked before each benchmark iteration, it dispatches progress updates through Rust's standard library channels (`std::sync::mpsc`) or shared-state primitives, allowing the main thread to update the interface without polling.