How the llmfit Benchmark Command Works: Automated LLM Inference Testing

The llmfit bench command automatically detects running LLM inference servers, measures tokens-per-second throughput across configurable runs, and optionally submits results to the community repository via GitHub pull requests.

The llmfit CLI tool from the AlexsJones/llmfit repository provides a comprehensive benchmarking subsystem that eliminates manual configuration when testing large language model performance. The llmfit bench command orchestrates a complete workflow from server detection through statistical aggregation, supporting major inference runtimes including Ollama, vLLM, Ferrum, MLX, and llama-cpp without requiring explicit endpoint setup.

Overview of the llmfit Benchmark Architecture

The benchmarking system implements a three-phase pipeline defined in llmfit-core/src/bench.rs. First, it discovers which inference provider is active and which model it serves. Next, it executes a warm-up request followed by multiple measurement runs to ensure stable results. Finally, it aggregates performance statistics and persists results locally or submits them to the community database when requested.

Phase 1: Automatic Provider Detection

The detection logic centers on the auto_detect_target() function (lines 730-779 in llmfit-core/src/bench.rs). This probe sequence checks known ports and endpoints for active inference servers in a specific priority order, returning a concrete BenchTarget struct.

Detection proceeds through provider-specific helper functions:

  • detect_ollama_model() queries the /api/tags endpoint to enumerate available models
  • detect_vllm_model() and detect_ferrum_model() target OpenAI-compatible /v1/models endpoints
  • detect_llamacpp_model() checks for llama.cpp server instances on standard ports
  • detect_openai_model() serves as a generic fallback for compatible servers

Each function returns the first model matching an optional user-supplied hint from the --model flag. If no specific provider is indicated, the system probes vLLM/Ferrum first, followed by Ollama and llama-cpp, ultimately falling back to MLX.

Phase 2: Executing Performance Benchmarks

Once detection returns a concrete target, the benchmark_target() dispatcher (lines 661-682) routes execution to provider-specific implementations. The system performs a warm-up request to stabilize the runtime before collecting timing data across the configured number of runs (defaulting to 3, configurable via --runs).

Ollama-Specific Benchmarking

For Ollama deployments, the bench_ollama() function (lines 124-158) sends prompts to the /api/generate endpoint. This implementation extracts native timing metrics including eval_duration and prompt_eval_duration directly from the response JSON. If these fields are absent, it falls back to wall-clock timing calculations to ensure compatibility across Ollama versions.

OpenAI-Compatible Providers

The bench_openai_compat() function (lines 261-295) handles vLLM, Ferrum, MLX, and llama-cpp through their shared OpenAI-compatible API surface. It posts requests to /v1/chat/completions and calculates tokens-per-second from wall-clock elapsed time. Note that Time-To-First-Token (TTFT) metrics are unavailable for these providers because the current implementation does not use streaming responses.

Phase 3: Result Aggregation and Storage

Each benchmark run produces a BenchRun struct instance (lines 10-25) containing raw timing data and token counts. The BenchSummary::from_runs() method (lines 62-92) computes statistical summaries including:

  • Minimum, average, and maximum tokens-per-second (TPS)
  • Average TTFT when available (Ollama native metrics only)
  • Average latency and token count across runs

The finalized data wraps in a BenchResult struct (lines 30-35) and renders via BenchResult::display() (lines 298-329) for formatted terminal output.

Persistence occurs automatically at:

~/.local/share/llmfit/benchmarks/pending/

When invoked with the --share flag, the system bundles stored results and creates a pull request against the llmfit-core/data/community/ directory. This process uses GitHub's device-flow authentication mechanism as documented in docs/cli.md (lines 44-53), requiring no manual token management.

CLI Usage Examples

Execute benchmarks against automatically detected servers:


# Benchmark the first available model (default 3 runs)

llmfit bench

# Target a specific provider and model

llmfit bench --provider vllm --model tinyllama-1b

# Increase statistical confidence with more runs

llmfit bench --runs 5

# Benchmark and contribute results to the community repository

llmfit bench --share

Command-line flag parsing occurs in llmfit-tui/src/main.rs around line 851, handling options including --provider, --model, --runs, --share, and --yes for non-interactive execution.

Summary

  • Auto-detection: The auto_detect_target() function in llmfit-core/src/bench.rs probes vLLM, Ferrum, Ollama, llama-cpp, and MLX endpoints without requiring manual configuration.
  • Provider-specific metrics: Ollama benchmarks extract native eval_duration fields while OpenAI-compatible providers use wall-clock timing calculations.
  • Statistical rigor: The BenchSummary::from_runs() method aggregates min/avg/max TPS, latency percentiles, and TTFT metrics across multiple runs.
  • Community sharing: Results persist to ~/.local/share/llmfit/benchmarks/pending/ and submit via --share using automated GitHub pull request workflows.

Frequently Asked Questions

Which inference providers does llmfit bench support?

The benchmark command supports Ollama, vLLM, Ferrum, MLX, and llama-cpp. The auto_detect_target() function in llmfit-core/src/bench.rs probes each provider's standard endpoints automatically, eliminating the need to specify the provider type manually unless desired.

How does llmfit calculate tokens-per-second for different providers?

For Ollama, the system extracts eval_duration directly from the native API response (lines 124-158). For OpenAI-compatible providers (vLLM, Ferrum, MLX, llama-cpp), bench_openai_compat() calculates TPS by dividing the total generated tokens by the wall-clock elapsed time, as these endpoints do not expose native timing fields in the standard chat completions format.

Where are benchmark results stored locally?

Successful benchmarks write JSON data to ~/.local/share/llmfit/benchmarks/pending/ following the XDG Base Directory specification. These files persist until explicitly submitted via --share or manually removed by the user.

What does the --share flag do when benchmarking?

The --share flag triggers a GitHub device-flow authentication sequence and creates a pull request adding your benchmark results to the llmfit-core/data/community/ directory. This contributes performance data to the public community database for model comparison and hardware optimization research, as documented in docs/cli.md (lines 98-108).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →