# How LLMFIT's Benchmarking Subsystem Measures Real tok/s and Submits Community Data

> Discover how LLMFIT measures real tok/s by testing live inference requests to providers like Ollama and vLLM. Learn how it submits community data to enrich the public leaderboard.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-08-21

---

**LLMFIT measures real tokens‑per‑second by issuing live inference requests to running providers like Ollama, vLLM, MLX, and llama‑cpp, calculating TPS from native timing fields or wall‑clock duration, then aggregates results and submits them via automated GitHub pull requests to enrich the public leaderboard.**

LLMFIT is an open‑source framework designed to match large language models with optimal hardware configurations. Understanding how the benchmarking subsystem measures real tok/s against running providers and submits community data requires examining the provider‑specific timing logic in [`llmfit-core/src/bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/bench.rs) and the submission pipeline in [`llmfit-core/src/share.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/share.rs).

## Measuring Real Tokens‑per‑Second Against Live Providers

The core benchmarking logic resides in **[`llmfit-core/src/bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/bench.rs)**, which implements distinct measurement strategies depending on the provider's API capabilities.

### Ollama Native Timing (TTFT and TPS)

For **Ollama** endpoints, the `bench_ollama` function leverages native timing fields exposed by the generate API to calculate precise throughput metrics.

The implementation sends a request to `/{base_url}/api/generate` and extracts three key measurements from the response:

- **Time‑to‑first‑token (TTFT)**: Directly obtained from the `prompt_eval_duration` field
- **Tokens‑per‑second (TPS)**: Computed as `eval_count / eval_duration` using Ollama's native timing, falling back to wall‑clock calculation (`output_tokens / total_wall`) when native data is unavailable
- **Total latency**: The complete round‑trip duration in milliseconds

The single‑run logic lives in the `ollama_generate` helper (line 144). The function performs a mandatory warmup request to stabilize the inference engine, then executes the configured number of benchmark runs, aggregating results through `BenchSummary::from_runs` (lines 47‑78).

### OpenAI‑Compatible Providers (vLLM, MLX, llama‑cpp)

For **vLLM**, **MLX**, and **llama‑cpp** (OpenAI‑compatible endpoints), `bench_openai_compat` sends requests to `/v1/chat/completions`.

Since these providers do not expose granular token timing fields in non‑streaming responses, LLMFIT estimates TPS strictly from wall‑clock measurement: `output_tokens / total_wall`. TTFT is recorded as `None` with an internal note that streaming would be required for true first‑token latency measurement. The request construction resides in `openai_chat` (line 82).

Both provider handlers execute a warmup request (excluded from statistics) before running the configured iteration count, ensuring measurements reflect steady‑state performance rather than cold‑start overhead.

## Aggregating Benchmark Statistics

After completing the run series, each provider returns a structured `BenchResult` containing the model identifier, provider tag, individual `BenchRun` instances, and a statistical summary. The `BenchSummary` computes aggregate metrics including mean, median, and standard deviation across all successful iterations, filtering out warmup data automatically.

## Storing and Submitting Community Data

Once benchmarking completes, the `--share` flag triggers the community submission workflow implemented in **[`llmfit-core/src/share.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/share.rs)**.

### Local Storage with Hardware Detection

The `store_local` function (line 397) persists benchmark results to a local JSON store before upstream submission. This process:

1. Detects system specifications via `SystemSpecs::detect()` from [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs), capturing RAM, CPU, and GPU configurations
2. Constructs a payload containing the benchmark summary, hardware identification, and tool version metadata
3. Writes the data to a pending directory for batch processing

### Automated PR Submission Workflow

The `share_all_pending` function (line 96) orchestrates community contributions:

1. Enumerates all pending benchmark files in the local store
2. Optionally displays a dry‑run preview for user verification
3. Triggers `submit_stored` (line 81) for each payload

The submission process creates a deterministic hardware slug via `build_submission` (line 41), generates a stable content hash to prevent duplicate uploads, and manages the GitHub workflow:

- Creates a fork of the upstream repository if the user lacks direct write access
- Generates a stable branch name using `short_hash` of the content
- Opens a new pull request or reuses an existing one for the hardware configuration
- Uploads the JSON payload via `put_file`

Upon merge, the benchmark data populates **`llmfit-core/data/community`** in the upstream repository, enriching the public leaderboard and enabling future LLMFIT runs to utilize locally‑measured TPS values rather than estimates.

## Code Example: Running a Full Benchmark Cycle

The following Rust snippet demonstrates the complete workflow from live measurement through community submission:

```rust
// Run Ollama benchmark with 3 iterations
let ollama_result = bench_ollama(
    "http://localhost:11434",
    "llama3.1:8b",
    3,
    &|run, total| println!("Ollama run {}/{}", run, total)
)?;

// Run vLLM benchmark with 4 iterations  
let vllm_result = bench_openai_compat(
    "http://localhost:8000",
    "llama3.1:8b",
    "vllm",
    4,
    &|run, total| println!("vLLM run {}/{}", run, total)
)?;

// Detect hardware and store locally
let specs = SystemSpecs::detect()?;
store_local(&[ollama_result, vllm_result], &specs)?;

// Submit to upstream repository (creates PR)
let share_opts = ShareOptions { dry_run: false, assume_yes: true };
share_all_pending(&share_opts, None)?;

```

This pattern executes real inference against running providers, aggregates statistically valid throughput data, and contributes findings to the community dataset according to the LLMFIT source code architecture.

## Summary

- **Real measurements**: `bench_ollama` and `bench_openai_compat` in [`llmfit-core/src/bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/bench.rs) issue live requests to calculate TPS using native timing (Ollama) or wall‑clock division (OpenAI‑compatible providers).
- **Statistical rigor**: Warmup requests precede timed runs, with results aggregated through `BenchSummary::from_runs` to eliminate cold‑start variance.
- **Hardware attribution**: `SystemSpecs::detect()` in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) captures RAM, CPU, and GPU details for accurate performance context.
- **Community contribution**: `store_local` and `share_all_pending` in [`llmfit-core/src/share.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/share.rs) automate GitHub PR creation to populate `llmfit-core/data/community` with validated benchmarks.

## Frequently Asked Questions

### How does LLMFIT calculate tokens‑per‑second for providers without native timing APIs?

For OpenAI‑compatible providers like vLLM, MLX, and llama‑cpp, LLMFIT calculates TPS by dividing the total output token count by the wall‑clock duration of the request. This method lacks granularity for time‑to‑first‑token measurement, which remains unreported unless the provider exposes streaming metadata.

### What is the purpose of the warmup request in LLMFIT benchmarks?

The warmup request stabilizes the inference engine's caches and GPU memory allocations before measurement begins. This ensures that subsequent timed runs reflect steady‑state performance rather than initialization overhead, producing more consistent and reproducible tok/s measurements.

### Where does LLMFIT store benchmark data before submitting to the community?

Pending benchmark data resides in a local JSON store managed by `store_local` in [`llmfit-core/src/share.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/share.rs) (line 397). The system groups results by hardware configuration, creating a deterministic `payload_slug` that prevents duplicate submissions while allowing batch uploads via `share_all_pending`.

### How does LLMFIT prevent duplicate benchmark submissions?

The `build_submission` function (line 41) generates a stable content hash for each benchmark payload. When `submit_stored` processes uploads, it checks for existing files with matching hashes in the target repository, skipping redundant entries and ensuring the community dataset contains only unique hardware‑model‑performance combinations.