How LLMFIT's Benchmarking Subsystem Measures Real tok/s and Submits Community Data
LLMFIT measures real tokens‑per‑second by issuing live inference requests to running providers like Ollama, vLLM, MLX, and llama‑cpp, calculating TPS from native timing fields or wall‑clock duration, then aggregates results and submits them via automated GitHub pull requests to enrich the public leaderboard.
LLMFIT is an open‑source framework designed to match large language models with optimal hardware configurations. Understanding how the benchmarking subsystem measures real tok/s against running providers and submits community data requires examining the provider‑specific timing logic in llmfit-core/src/bench.rs and the submission pipeline in llmfit-core/src/share.rs.
Measuring Real Tokens‑per‑Second Against Live Providers
The core benchmarking logic resides in llmfit-core/src/bench.rs, which implements distinct measurement strategies depending on the provider's API capabilities.
Ollama Native Timing (TTFT and TPS)
For Ollama endpoints, the bench_ollama function leverages native timing fields exposed by the generate API to calculate precise throughput metrics.
The implementation sends a request to /{base_url}/api/generate and extracts three key measurements from the response:
- Time‑to‑first‑token (TTFT): Directly obtained from the
prompt_eval_durationfield - Tokens‑per‑second (TPS): Computed as
eval_count / eval_durationusing Ollama's native timing, falling back to wall‑clock calculation (output_tokens / total_wall) when native data is unavailable - Total latency: The complete round‑trip duration in milliseconds
The single‑run logic lives in the ollama_generate helper (line 144). The function performs a mandatory warmup request to stabilize the inference engine, then executes the configured number of benchmark runs, aggregating results through BenchSummary::from_runs (lines 47‑78).
OpenAI‑Compatible Providers (vLLM, MLX, llama‑cpp)
For vLLM, MLX, and llama‑cpp (OpenAI‑compatible endpoints), bench_openai_compat sends requests to /v1/chat/completions.
Since these providers do not expose granular token timing fields in non‑streaming responses, LLMFIT estimates TPS strictly from wall‑clock measurement: output_tokens / total_wall. TTFT is recorded as None with an internal note that streaming would be required for true first‑token latency measurement. The request construction resides in openai_chat (line 82).
Both provider handlers execute a warmup request (excluded from statistics) before running the configured iteration count, ensuring measurements reflect steady‑state performance rather than cold‑start overhead.
Aggregating Benchmark Statistics
After completing the run series, each provider returns a structured BenchResult containing the model identifier, provider tag, individual BenchRun instances, and a statistical summary. The BenchSummary computes aggregate metrics including mean, median, and standard deviation across all successful iterations, filtering out warmup data automatically.
Storing and Submitting Community Data
Once benchmarking completes, the --share flag triggers the community submission workflow implemented in llmfit-core/src/share.rs.
Local Storage with Hardware Detection
The store_local function (line 397) persists benchmark results to a local JSON store before upstream submission. This process:
- Detects system specifications via
SystemSpecs::detect()fromllmfit-core/src/hardware.rs, capturing RAM, CPU, and GPU configurations - Constructs a payload containing the benchmark summary, hardware identification, and tool version metadata
- Writes the data to a pending directory for batch processing
Automated PR Submission Workflow
The share_all_pending function (line 96) orchestrates community contributions:
- Enumerates all pending benchmark files in the local store
- Optionally displays a dry‑run preview for user verification
- Triggers
submit_stored(line 81) for each payload
The submission process creates a deterministic hardware slug via build_submission (line 41), generates a stable content hash to prevent duplicate uploads, and manages the GitHub workflow:
- Creates a fork of the upstream repository if the user lacks direct write access
- Generates a stable branch name using
short_hashof the content - Opens a new pull request or reuses an existing one for the hardware configuration
- Uploads the JSON payload via
put_file
Upon merge, the benchmark data populates llmfit-core/data/community in the upstream repository, enriching the public leaderboard and enabling future LLMFIT runs to utilize locally‑measured TPS values rather than estimates.
Code Example: Running a Full Benchmark Cycle
The following Rust snippet demonstrates the complete workflow from live measurement through community submission:
// Run Ollama benchmark with 3 iterations
let ollama_result = bench_ollama(
"http://localhost:11434",
"llama3.1:8b",
3,
&|run, total| println!("Ollama run {}/{}", run, total)
)?;
// Run vLLM benchmark with 4 iterations
let vllm_result = bench_openai_compat(
"http://localhost:8000",
"llama3.1:8b",
"vllm",
4,
&|run, total| println!("vLLM run {}/{}", run, total)
)?;
// Detect hardware and store locally
let specs = SystemSpecs::detect()?;
store_local(&[ollama_result, vllm_result], &specs)?;
// Submit to upstream repository (creates PR)
let share_opts = ShareOptions { dry_run: false, assume_yes: true };
share_all_pending(&share_opts, None)?;
This pattern executes real inference against running providers, aggregates statistically valid throughput data, and contributes findings to the community dataset according to the LLMFIT source code architecture.
Summary
- Real measurements:
bench_ollamaandbench_openai_compatinllmfit-core/src/bench.rsissue live requests to calculate TPS using native timing (Ollama) or wall‑clock division (OpenAI‑compatible providers). - Statistical rigor: Warmup requests precede timed runs, with results aggregated through
BenchSummary::from_runsto eliminate cold‑start variance. - Hardware attribution:
SystemSpecs::detect()inllmfit-core/src/hardware.rscaptures RAM, CPU, and GPU details for accurate performance context. - Community contribution:
store_localandshare_all_pendinginllmfit-core/src/share.rsautomate GitHub PR creation to populatellmfit-core/data/communitywith validated benchmarks.
Frequently Asked Questions
How does LLMFIT calculate tokens‑per‑second for providers without native timing APIs?
For OpenAI‑compatible providers like vLLM, MLX, and llama‑cpp, LLMFIT calculates TPS by dividing the total output token count by the wall‑clock duration of the request. This method lacks granularity for time‑to‑first‑token measurement, which remains unreported unless the provider exposes streaming metadata.
What is the purpose of the warmup request in LLMFIT benchmarks?
The warmup request stabilizes the inference engine's caches and GPU memory allocations before measurement begins. This ensures that subsequent timed runs reflect steady‑state performance rather than initialization overhead, producing more consistent and reproducible tok/s measurements.
Where does LLMFIT store benchmark data before submitting to the community?
Pending benchmark data resides in a local JSON store managed by store_local in llmfit-core/src/share.rs (line 397). The system groups results by hardware configuration, creating a deterministic payload_slug that prevents duplicate submissions while allowing batch uploads via share_all_pending.
How does LLMFIT prevent duplicate benchmark submissions?
The build_submission function (line 41) generates a stable content hash for each benchmark payload. When submit_stored processes uploads, it checks for existing files with matching hashes in the target repository, skipping redundant entries and ensuring the community dataset contains only unique hardware‑model‑performance combinations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →