# How the llmfit Benchmark Command Works: Automated LLM Inference Testing

> Learn how the llmfit bench command automates LLM inference testing. Measure throughput, detect servers, and submit results to the community repository.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-09-13

---

**The `llmfit bench` command automatically detects running LLM inference servers, measures tokens-per-second throughput across configurable runs, and optionally submits results to the community repository via GitHub pull requests.**

The `llmfit` CLI tool from the AlexsJones/llmfit repository provides a comprehensive benchmarking subsystem that eliminates manual configuration when testing large language model performance. The `llmfit bench` command orchestrates a complete workflow from server detection through statistical aggregation, supporting major inference runtimes including Ollama, vLLM, Ferrum, MLX, and llama-cpp without requiring explicit endpoint setup.

## Overview of the llmfit Benchmark Architecture

The benchmarking system implements a three-phase pipeline defined in [`llmfit-core/src/bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/bench.rs). First, it discovers which inference provider is active and which model it serves. Next, it executes a warm-up request followed by multiple measurement runs to ensure stable results. Finally, it aggregates performance statistics and persists results locally or submits them to the community database when requested.

## Phase 1: Automatic Provider Detection

The detection logic centers on the `auto_detect_target()` function (lines 730-779 in [`llmfit-core/src/bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/bench.rs)). This probe sequence checks known ports and endpoints for active inference servers in a specific priority order, returning a concrete `BenchTarget` struct.

Detection proceeds through provider-specific helper functions:

- `detect_ollama_model()` queries the `/api/tags` endpoint to enumerate available models
- `detect_vllm_model()` and `detect_ferrum_model()` target OpenAI-compatible `/v1/models` endpoints
- `detect_llamacpp_model()` checks for llama.cpp server instances on standard ports
- `detect_openai_model()` serves as a generic fallback for compatible servers

Each function returns the first model matching an optional user-supplied hint from the `--model` flag. If no specific provider is indicated, the system probes vLLM/Ferrum first, followed by Ollama and llama-cpp, ultimately falling back to MLX.

## Phase 2: Executing Performance Benchmarks

Once detection returns a concrete target, the `benchmark_target()` dispatcher (lines 661-682) routes execution to provider-specific implementations. The system performs a warm-up request to stabilize the runtime before collecting timing data across the configured number of runs (defaulting to 3, configurable via `--runs`).

### Ollama-Specific Benchmarking

For Ollama deployments, the `bench_ollama()` function (lines 124-158) sends prompts to the `/api/generate` endpoint. This implementation extracts **native timing metrics** including `eval_duration` and `prompt_eval_duration` directly from the response JSON. If these fields are absent, it falls back to wall-clock timing calculations to ensure compatibility across Ollama versions.

### OpenAI-Compatible Providers

The `bench_openai_compat()` function (lines 261-295) handles vLLM, Ferrum, MLX, and llama-cpp through their shared OpenAI-compatible API surface. It posts requests to `/v1/chat/completions` and calculates tokens-per-second from wall-clock elapsed time. Note that **Time-To-First-Token (TTFT)** metrics are unavailable for these providers because the current implementation does not use streaming responses.

## Phase 3: Result Aggregation and Storage

Each benchmark run produces a `BenchRun` struct instance (lines 10-25) containing raw timing data and token counts. The `BenchSummary::from_runs()` method (lines 62-92) computes statistical summaries including:

- Minimum, average, and maximum tokens-per-second (TPS)
- Average TTFT when available (Ollama native metrics only)
- Average latency and token count across runs

The finalized data wraps in a `BenchResult` struct (lines 30-35) and renders via `BenchResult::display()` (lines 298-329) for formatted terminal output.

Persistence occurs automatically at:

```text
~/.local/share/llmfit/benchmarks/pending/

```

When invoked with the `--share` flag, the system bundles stored results and creates a pull request against the `llmfit-core/data/community/` directory. This process uses **GitHub's device-flow authentication** mechanism as documented in [`docs/cli.md`](https://github.com/AlexsJones/llmfit/blob/main/docs/cli.md) (lines 44-53), requiring no manual token management.

## CLI Usage Examples

Execute benchmarks against automatically detected servers:

```sh

# Benchmark the first available model (default 3 runs)

llmfit bench

# Target a specific provider and model

llmfit bench --provider vllm --model tinyllama-1b

# Increase statistical confidence with more runs

llmfit bench --runs 5

# Benchmark and contribute results to the community repository

llmfit bench --share

```

Command-line flag parsing occurs in [`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs) around line 851, handling options including `--provider`, `--model`, `--runs`, `--share`, and `--yes` for non-interactive execution.

## Summary

- **Auto-detection**: The `auto_detect_target()` function in [`llmfit-core/src/bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/bench.rs) probes vLLM, Ferrum, Ollama, llama-cpp, and MLX endpoints without requiring manual configuration.
- **Provider-specific metrics**: Ollama benchmarks extract native `eval_duration` fields while OpenAI-compatible providers use wall-clock timing calculations.
- **Statistical rigor**: The `BenchSummary::from_runs()` method aggregates min/avg/max TPS, latency percentiles, and TTFT metrics across multiple runs.
- **Community sharing**: Results persist to `~/.local/share/llmfit/benchmarks/pending/` and submit via `--share` using automated GitHub pull request workflows.

## Frequently Asked Questions

### Which inference providers does llmfit bench support?

The benchmark command supports Ollama, vLLM, Ferrum, MLX, and llama-cpp. The `auto_detect_target()` function in [`llmfit-core/src/bench.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/bench.rs) probes each provider's standard endpoints automatically, eliminating the need to specify the provider type manually unless desired.

### How does llmfit calculate tokens-per-second for different providers?

For Ollama, the system extracts `eval_duration` directly from the native API response (lines 124-158). For OpenAI-compatible providers (vLLM, Ferrum, MLX, llama-cpp), `bench_openai_compat()` calculates TPS by dividing the total generated tokens by the wall-clock elapsed time, as these endpoints do not expose native timing fields in the standard chat completions format.

### Where are benchmark results stored locally?

Successful benchmarks write JSON data to `~/.local/share/llmfit/benchmarks/pending/` following the XDG Base Directory specification. These files persist until explicitly submitted via `--share` or manually removed by the user.

### What does the --share flag do when benchmarking?

The `--share` flag triggers a GitHub device-flow authentication sequence and creates a pull request adding your benchmark results to the `llmfit-core/data/community/` directory. This contributes performance data to the public community database for model comparison and hardware optimization research, as documented in [`docs/cli.md`](https://github.com/AlexsJones/llmfit/blob/main/docs/cli.md) (lines 98-108).