How LLMFIT's Benchmarking Subsystem Measures Real tok/s and Submits Community Data

LLMFIT measures real tokens‑per‑second by issuing live inference requests to running providers like Ollama, vLLM, MLX, and llama‑cpp, calculating TPS from native timing fields or wall‑clock duration, then aggregates results and submits them via automated GitHub pull requests to enrich the public leaderboard.

LLMFIT is an open‑source framework designed to match large language models with optimal hardware configurations. Understanding how the benchmarking subsystem measures real tok/s against running providers and submits community data requires examining the provider‑specific timing logic in llmfit-core/src/bench.rs and the submission pipeline in llmfit-core/src/share.rs.

Measuring Real Tokens‑per‑Second Against Live Providers

The core benchmarking logic resides in llmfit-core/src/bench.rs, which implements distinct measurement strategies depending on the provider's API capabilities.

Ollama Native Timing (TTFT and TPS)

For Ollama endpoints, the bench_ollama function leverages native timing fields exposed by the generate API to calculate precise throughput metrics.

The implementation sends a request to /{base_url}/api/generate and extracts three key measurements from the response:

  • Time‑to‑first‑token (TTFT): Directly obtained from the prompt_eval_duration field
  • Tokens‑per‑second (TPS): Computed as eval_count / eval_duration using Ollama's native timing, falling back to wall‑clock calculation (output_tokens / total_wall) when native data is unavailable
  • Total latency: The complete round‑trip duration in milliseconds

The single‑run logic lives in the ollama_generate helper (line 144). The function performs a mandatory warmup request to stabilize the inference engine, then executes the configured number of benchmark runs, aggregating results through BenchSummary::from_runs (lines 47‑78).

OpenAI‑Compatible Providers (vLLM, MLX, llama‑cpp)

For vLLM, MLX, and llama‑cpp (OpenAI‑compatible endpoints), bench_openai_compat sends requests to /v1/chat/completions.

Since these providers do not expose granular token timing fields in non‑streaming responses, LLMFIT estimates TPS strictly from wall‑clock measurement: output_tokens / total_wall. TTFT is recorded as None with an internal note that streaming would be required for true first‑token latency measurement. The request construction resides in openai_chat (line 82).

Both provider handlers execute a warmup request (excluded from statistics) before running the configured iteration count, ensuring measurements reflect steady‑state performance rather than cold‑start overhead.

Aggregating Benchmark Statistics

After completing the run series, each provider returns a structured BenchResult containing the model identifier, provider tag, individual BenchRun instances, and a statistical summary. The BenchSummary computes aggregate metrics including mean, median, and standard deviation across all successful iterations, filtering out warmup data automatically.

Storing and Submitting Community Data

Once benchmarking completes, the --share flag triggers the community submission workflow implemented in llmfit-core/src/share.rs.

Local Storage with Hardware Detection

The store_local function (line 397) persists benchmark results to a local JSON store before upstream submission. This process:

  1. Detects system specifications via SystemSpecs::detect() from llmfit-core/src/hardware.rs, capturing RAM, CPU, and GPU configurations
  2. Constructs a payload containing the benchmark summary, hardware identification, and tool version metadata
  3. Writes the data to a pending directory for batch processing

Automated PR Submission Workflow

The share_all_pending function (line 96) orchestrates community contributions:

  1. Enumerates all pending benchmark files in the local store
  2. Optionally displays a dry‑run preview for user verification
  3. Triggers submit_stored (line 81) for each payload

The submission process creates a deterministic hardware slug via build_submission (line 41), generates a stable content hash to prevent duplicate uploads, and manages the GitHub workflow:

  • Creates a fork of the upstream repository if the user lacks direct write access
  • Generates a stable branch name using short_hash of the content
  • Opens a new pull request or reuses an existing one for the hardware configuration
  • Uploads the JSON payload via put_file

Upon merge, the benchmark data populates llmfit-core/data/community in the upstream repository, enriching the public leaderboard and enabling future LLMFIT runs to utilize locally‑measured TPS values rather than estimates.

Code Example: Running a Full Benchmark Cycle

The following Rust snippet demonstrates the complete workflow from live measurement through community submission:

// Run Ollama benchmark with 3 iterations
let ollama_result = bench_ollama(
    "http://localhost:11434",
    "llama3.1:8b",
    3,
    &|run, total| println!("Ollama run {}/{}", run, total)
)?;

// Run vLLM benchmark with 4 iterations  
let vllm_result = bench_openai_compat(
    "http://localhost:8000",
    "llama3.1:8b",
    "vllm",
    4,
    &|run, total| println!("vLLM run {}/{}", run, total)
)?;

// Detect hardware and store locally
let specs = SystemSpecs::detect()?;
store_local(&[ollama_result, vllm_result], &specs)?;

// Submit to upstream repository (creates PR)
let share_opts = ShareOptions { dry_run: false, assume_yes: true };
share_all_pending(&share_opts, None)?;

This pattern executes real inference against running providers, aggregates statistically valid throughput data, and contributes findings to the community dataset according to the LLMFIT source code architecture.

Summary

  • Real measurements: bench_ollama and bench_openai_compat in llmfit-core/src/bench.rs issue live requests to calculate TPS using native timing (Ollama) or wall‑clock division (OpenAI‑compatible providers).
  • Statistical rigor: Warmup requests precede timed runs, with results aggregated through BenchSummary::from_runs to eliminate cold‑start variance.
  • Hardware attribution: SystemSpecs::detect() in llmfit-core/src/hardware.rs captures RAM, CPU, and GPU details for accurate performance context.
  • Community contribution: store_local and share_all_pending in llmfit-core/src/share.rs automate GitHub PR creation to populate llmfit-core/data/community with validated benchmarks.

Frequently Asked Questions

How does LLMFIT calculate tokens‑per‑second for providers without native timing APIs?

For OpenAI‑compatible providers like vLLM, MLX, and llama‑cpp, LLMFIT calculates TPS by dividing the total output token count by the wall‑clock duration of the request. This method lacks granularity for time‑to‑first‑token measurement, which remains unreported unless the provider exposes streaming metadata.

What is the purpose of the warmup request in LLMFIT benchmarks?

The warmup request stabilizes the inference engine's caches and GPU memory allocations before measurement begins. This ensures that subsequent timed runs reflect steady‑state performance rather than initialization overhead, producing more consistent and reproducible tok/s measurements.

Where does LLMFIT store benchmark data before submitting to the community?

Pending benchmark data resides in a local JSON store managed by store_local in llmfit-core/src/share.rs (line 397). The system groups results by hardware configuration, creating a deterministic payload_slug that prevents duplicate submissions while allowing batch uploads via share_all_pending.

How does LLMFIT prevent duplicate benchmark submissions?

The build_submission function (line 41) generates a stable content hash for each benchmark payload. When submit_stored processes uploads, it checks for existing files with matching hashes in the target repository, skipping redundant entries and ensuring the community dataset contains only unique hardware‑model‑performance combinations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →