How the Inference Benchmark Package Measures tok/s in Magnitude
The inference benchmark package calculates tokens per second by dividing the total generated tokens by the elapsed time in milliseconds, multiplying by 1000 to convert to seconds, then reporting median throughput rates for both prefill and decode phases.
The inference-benchmark package within the magnitudedev/magnitude repository provides a systematic approach to evaluating language model performance through automated trial execution and statistical analysis. By capturing precise timing data and token counts from each request, it derives accurate tok/s measurements that enable performance comparisons across different model configurations.
Executing Requests and Capturing Timestamps
The measurement process begins when the benchmark executes individual requests against a specified target.
The runRequest Function
In packages/inference-benchmark/src/benchmark.ts, the runRequest function (lines 58-71) orchestrates the request lifecycle. This function:
- Records the exact moment a request is submitted using the
submittedAtMstimestamp - Awaits the model's complete response
- Returns a structured
RequestObservationobject containing the raw response data and metadata
This observation becomes the foundational data point for all subsequent throughput calculations.
Extracting Token Usage from Observations
Each RequestObservation object captured during the trial contains critical usage statistics in its terminal usage field:
promptTokens: The number of tokens in the input promptcompletionTokens: Tokens generated in the responsegeneratedTokens: The total count of tokens produced during the inference phase
These values represent the actual computational workload processed by the model, serving as the numerator in all tok/s calculations.
Computing Tokens per Second
After collecting observations across all requests in a trial, the benchmark aggregates the data to produce throughput metrics.
The analyzeTrial Aggregation Logic
The analyzeTrial helper, defined in packages/inference-benchmark/src/analysis.ts, handles the statistical aggregation. It computes two distinct throughput metrics using the following formulas:
- Prefill tok/s =
generatedTokens * 1000 / prefillMs - Decode tok/s =
generatedTokens * 1000 / decodeMs
The multiplication by 1000 converts the elapsed time from milliseconds to seconds, yielding a per-second token rate. The benchmark separates prefill (initial processing) and decode (generation) phases to provide granular insight into where computational bottlenecks occur.
Reporting Benchmark Results
Once calculations are complete, packages/inference-benchmark/src/report.ts (lines 48-67) formats the results into a readable markdown table. This output displays the median values across all observations for:
- Prefill tok/s: Throughput during the prompt processing phase
- Decode tok/s: Throughput during token generation
- Achieved completion (tok/s): The effective completion rate realized during the trial
Using median values rather than averages ensures that outlier requests do not skew the performance characterization.
Practical Implementation Example
The following TypeScript example demonstrates how to run a benchmark and extract tok/s metrics:
import { evaluate, compare } from "@magnitudedev/inference-benchmark";
import { readFileSync } from "fs";
// 1️⃣ Load a trial plan (JSON generated by the CLI)
const plan = JSON.parse(readFileSync("plan.json", "utf-8"));
// 2️⃣ Define the target configuration (e.g., a local model)
const target = {
id: "local-qwen3",
kind: "local",
model: "Qwen3-Coder",
// …other required fields
};
// 3️⃣ Run the benchmark
const result = await evaluate(plan, target).pipe(
Effect.runPromise,
);
// 4️⃣ Inspect token-per-second metrics
result.analyses.forEach((a) => {
console.log(
`Trial ${a.trialId}:`,
`prefill ${a.prefillTokensPerSecond?.median?.toFixed(1)} tok/s`,
`decode ${a.decodeTokensPerSecond?.median?.toFixed(1)} tok/s`,
`completion ${a.achievedCompletionTokensPerSecond?.toFixed(1)} tok/s`,
);
});
Summary
- The inference benchmark package measures tok/s by combining token counts from
RequestObservationobjects with precise timing data captured inrunRequest. - Prefill throughput and decode throughput are calculated separately using the formula:
generatedTokens * 1000 / durationInMs. - The
analyzeTrialfunction inanalysis.tsaggregates these rates and computes median values to ensure statistically robust reporting. - Results are formatted in
report.ts(lines 48-67) as a markdown table showing median tok/s for each phase.
Frequently Asked Questions
What is the difference between prefill tok/s and decode tok/s?
Prefill tok/s measures the throughput during the initial prompt processing phase, where the model ingests and processes the input tokens. Decode tok/s measures the generation speed once the model begins producing new tokens. Separating these metrics allows developers to identify whether latency stems from prompt complexity (prefill) or generation efficiency (decode).
Where does the token count data originate?
The token counts originate from the RequestObservation object's terminal usage field, which records promptTokens, completionTokens, and generatedTokens as reported by the model or inference engine. These values represent the actual tokens processed during the request lifecycle managed by runRequest in benchmark.ts.
How does the benchmark calculate time durations?
The benchmark records submittedAtMs when the request initiates and derives phase-specific durations (prefillMs and decodeMs) from the response stream timing. These millisecond values are converted to seconds by multiplying the token count by 1000 before division, as implemented in the analyzeTrial function.
Can this benchmark measure tok/s for remote API-based models?
Yes, the inference benchmark package supports both local and remote targets. As long as the target returns the required usage statistics (prompt tokens, completion tokens, and timing data), the analyzeTrial logic computes tok/s identically regardless of whether the model runs locally or via a remote API endpoint.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →