# How the Inference Benchmark Package Measures tok/s in Magnitude

> Discover how the inference benchmark package in magnitudedev/magnitude measures tok/s by dividing tokens by time, reporting median throughput for prefill and decode phases.

- Repository: [Magnitude/magnitude](https://github.com/magnitudedev/magnitude)
- Tags: performance
- Published: 2026-09-08

---

**The inference benchmark package calculates tokens per second by dividing the total generated tokens by the elapsed time in milliseconds, multiplying by 1000 to convert to seconds, then reporting median throughput rates for both prefill and decode phases.**

The `inference-benchmark` package within the magnitudedev/magnitude repository provides a systematic approach to evaluating language model performance through automated trial execution and statistical analysis. By capturing precise timing data and token counts from each request, it derives accurate tok/s measurements that enable performance comparisons across different model configurations.

## Executing Requests and Capturing Timestamps

The measurement process begins when the benchmark executes individual requests against a specified target.

### The runRequest Function

In [`packages/inference-benchmark/src/benchmark.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/inference-benchmark/src/benchmark.ts), the `runRequest` function (lines 58-71) orchestrates the request lifecycle. This function:

- Records the exact moment a request is submitted using the `submittedAtMs` timestamp
- Awaits the model's complete response
- Returns a structured `RequestObservation` object containing the raw response data and metadata

This observation becomes the foundational data point for all subsequent throughput calculations.

## Extracting Token Usage from Observations

Each `RequestObservation` object captured during the trial contains critical usage statistics in its terminal usage field:

- **`promptTokens`**: The number of tokens in the input prompt
- **`completionTokens`**: Tokens generated in the response
- **`generatedTokens`**: The total count of tokens produced during the inference phase

These values represent the actual computational workload processed by the model, serving as the numerator in all tok/s calculations.

## Computing Tokens per Second

After collecting observations across all requests in a trial, the benchmark aggregates the data to produce throughput metrics.

### The analyzeTrial Aggregation Logic

The `analyzeTrial` helper, defined in [`packages/inference-benchmark/src/analysis.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/inference-benchmark/src/analysis.ts), handles the statistical aggregation. It computes two distinct throughput metrics using the following formulas:

- **Prefill tok/s** = `generatedTokens * 1000 / prefillMs`
- **Decode tok/s** = `generatedTokens * 1000 / decodeMs`

The multiplication by 1000 converts the elapsed time from milliseconds to seconds, yielding a per-second token rate. The benchmark separates prefill (initial processing) and decode (generation) phases to provide granular insight into where computational bottlenecks occur.

## Reporting Benchmark Results

Once calculations are complete, [`packages/inference-benchmark/src/report.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/inference-benchmark/src/report.ts) (lines 48-67) formats the results into a readable markdown table. This output displays the median values across all observations for:

- **Prefill tok/s**: Throughput during the prompt processing phase
- **Decode tok/s**: Throughput during token generation
- **Achieved completion (tok/s)**: The effective completion rate realized during the trial

Using median values rather than averages ensures that outlier requests do not skew the performance characterization.

## Practical Implementation Example

The following TypeScript example demonstrates how to run a benchmark and extract tok/s metrics:

```typescript
import { evaluate, compare } from "@magnitudedev/inference-benchmark";
import { readFileSync } from "fs";

// 1️⃣ Load a trial plan (JSON generated by the CLI)
const plan = JSON.parse(readFileSync("plan.json", "utf-8"));

// 2️⃣ Define the target configuration (e.g., a local model)
const target = {
  id: "local-qwen3",
  kind: "local",
  model: "Qwen3-Coder",
  // …other required fields
};

// 3️⃣ Run the benchmark
const result = await evaluate(plan, target).pipe(
  Effect.runPromise,
);

// 4️⃣ Inspect token-per-second metrics
result.analyses.forEach((a) => {
  console.log(
    `Trial ${a.trialId}:`,
    `prefill ${a.prefillTokensPerSecond?.median?.toFixed(1)} tok/s`,
    `decode ${a.decodeTokensPerSecond?.median?.toFixed(1)} tok/s`,
    `completion ${a.achievedCompletionTokensPerSecond?.toFixed(1)} tok/s`,
  );
});

```

## Summary

- The inference benchmark package measures tok/s by combining token counts from `RequestObservation` objects with precise timing data captured in `runRequest`.
- **Prefill throughput** and **decode throughput** are calculated separately using the formula: `generatedTokens * 1000 / durationInMs`.
- The `analyzeTrial` function in [`analysis.ts`](https://github.com/magnitudedev/magnitude/blob/main/analysis.ts) aggregates these rates and computes median values to ensure statistically robust reporting.
- Results are formatted in [`report.ts`](https://github.com/magnitudedev/magnitude/blob/main/report.ts) (lines 48-67) as a markdown table showing median tok/s for each phase.

## Frequently Asked Questions

### What is the difference between prefill tok/s and decode tok/s?

**Prefill tok/s** measures the throughput during the initial prompt processing phase, where the model ingests and processes the input tokens. **Decode tok/s** measures the generation speed once the model begins producing new tokens. Separating these metrics allows developers to identify whether latency stems from prompt complexity (prefill) or generation efficiency (decode).

### Where does the token count data originate?

The token counts originate from the `RequestObservation` object's terminal usage field, which records `promptTokens`, `completionTokens`, and `generatedTokens` as reported by the model or inference engine. These values represent the actual tokens processed during the request lifecycle managed by `runRequest` in [`benchmark.ts`](https://github.com/magnitudedev/magnitude/blob/main/benchmark.ts).

### How does the benchmark calculate time durations?

The benchmark records `submittedAtMs` when the request initiates and derives phase-specific durations (prefillMs and decodeMs) from the response stream timing. These millisecond values are converted to seconds by multiplying the token count by 1000 before division, as implemented in the `analyzeTrial` function.

### Can this benchmark measure tok/s for remote API-based models?

Yes, the inference benchmark package supports both local and remote targets. As long as the target returns the required usage statistics (prompt tokens, completion tokens, and timing data), the `analyzeTrial` logic computes tok/s identically regardless of whether the model runs locally or via a remote API endpoint.