# How to Run Performance Benchmarks Using the Twinkle Eval Benchmark Command

> Learn to run performance benchmarks with the Twinkle Eval benchmark command. Measure LLM latency throughput TTFT and TPOT easily. Explore the BenchmarkRunner class today.

- Repository: [Twinkle AI/eval](https://github.com/ai-twinkle/eval)
- Tags: how-to-guide
- Published: 2026-02-23

---

**Run comprehensive LLM performance benchmarks by executing `twinkle-eval --benchmark`, which measures latency, throughput, time-to-first-token (TTFT), and time-per-output-token (TPOT) through the `BenchmarkRunner` class in [`twinkle_eval/benchmark.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/benchmark.py).**

The `ai-twinkle/eval` repository provides a built-in benchmarking suite that evaluates Large Language Model (LLM) performance without requiring external tools. Whether you need to measure baseline latency or stress-test throughput under concurrent load, the Twinkle Eval benchmark command offers both burst and rate-limited execution modes. This guide explains how to run performance benchmarks using the Twinkle Eval benchmark command, configure test parameters, and interpret the resulting metrics.

## What the Twinkle Eval Benchmark Measures

The benchmark engine captures four critical performance indicators for LLM inference:

- **Latency**: End-to-end request duration from submission to final token
- **Throughput**: Requests per second and tokens per second processed
- **Time-to-First-Token (TTFT)**: Duration from request start until the first response token arrives
- **Time-per-Output-Token (TPOT)**: Average generation time for each subsequent token after the first

These metrics are calculated in [`twinkle_eval/benchmark.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/benchmark.py) within the `_calculate_metrics()` method (lines 252-444), which processes raw `RequestResult` objects into a `BenchmarkMetrics` dataclass.

## Prerequisites and Configuration

Before running benchmarks, ensure you have a valid [`config.yaml`](https://github.com/ai-twinkle/eval/blob/main/config.yaml) file that specifies the LLM provider and model parameters. The benchmark relies on [`twinkle_eval/config.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/config.py) to load settings and [`twinkle_eval/models.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/models.py) (specifically `LLMFactory.create_llm()`) to instantiate the client.

Minimal configuration example:

```yaml
llm:
  provider: openai
  model: gpt-3.5-turbo
  api_key: ${OPENAI_API_KEY}
  temperature: 0.7

```

## Running the Benchmark from the Command Line

### Basic Benchmark Execution

Execute the simplest benchmark using default parameters:

```bash
twinkle-eval --benchmark

```

This command:
1. Loads [`config.yaml`](https://github.com/ai-twinkle/eval/blob/main/config.yaml) via `load_config()` in [`twinkle_eval/main.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/main.py)
2. Initializes `BenchmarkRunner` with default values (100 requests, 10 concurrent workers)
3. Uses the default prompt: "台灣的首都是哪裡？" (What is the capital of Taiwan?)
4. Outputs results to [`benchmark_results_YYYYMMDD_HHMMSS.json`](https://github.com/ai-twinkle/eval/blob/main/benchmark_results_YYYYMMDD_HHMMSS.json)

### Advanced Options and Rate Limiting

For controlled load testing, specify concurrency limits, request rates, and custom prompts:

```bash
twinkle-eval \
    --benchmark \
    --benchmark-requests 50 \
    --benchmark-concurrency 5 \
    --benchmark-rate 2 \
    --benchmark-duration 30 \
    --benchmark-prompt "Explain quantum entanglement in one sentence."

```

**Parameter explanations:**
- `--benchmark-requests`: Total number of requests to issue (default: 100)
- `--benchmark-concurrency`: Maximum simultaneous HTTP connections (default: 10)
- `--benchmark-rate`: Requests-per-second ceiling (enables rate-limited mode when specified)
- `--benchmark-duration`: Maximum wall-clock time in seconds (runner stops early if exceeded)
- `--benchmark-prompt`: Input text sent to the LLM for each request

When `--benchmark-rate` is provided, `BenchmarkRunner.run_benchmark()` (lines 106-145 in [`twinkle_eval/benchmark.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/benchmark.py)) invokes `_run_rate_limited_benchmark()` (lines 181-213), which uses a `ThreadPoolExecutor` combined with a `Semaphore` to enforce the rate limit. Without a rate limit, `_run_burst_benchmark()` (lines 145-179) executes requests at maximum speed.

## Understanding the Benchmark Architecture

### CLI Entry Point and Argument Parsing

The command-line interface is defined in [`twinkle_eval/main.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/main.py). Lines 73-100 define the argument parser, adding the `--benchmark` flag and its associated parameters. When `--benchmark` is detected, the `main()` function (around lines 604-646) branches to benchmark execution:

```python
if args.benchmark:
    config = load_config(args.config)
    runner = BenchmarkRunner(
        config,
        prompt=args.benchmark_prompt,
        num_requests=args.benchmark_requests,
        concurrent_requests=args.benchmark_concurrency,
        request_rate=args.benchmark_rate,
        duration=args.benchmark_duration
    )
    metrics = runner.run_benchmark()

```

### The BenchmarkRunner Class

Located in [`twinkle_eval/benchmark.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/benchmark.py), the `BenchmarkRunner` class orchestrates the load test. The `__init__` method (lines 96-100) stores configuration parameters, while `_setup_llm()` (lines 101-105) initializes the LLM client via `LLMFactory.create_llm()`.

The core execution methods are:
- `run_benchmark()`: Entry point that selects burst or rate-limited mode
- `_run_burst_benchmark()`: Maximum throughput testing without rate constraints
- `_run_rate_limited_benchmark()`: Controlled load with `Semaphore`-based throttling
- `_send_request()`: Individual request execution with timing instrumentation (lines 214-242)

### Metrics Calculation and Results Export

After request completion, `_calculate_metrics()` (lines 252-444) aggregates raw results into statistical summaries. It computes:
- Throughput statistics (requests/second, tokens/second)
- Latency percentiles (mean, median, p95, p99)
- TTFT and TPOT distributions

The [`main.py`](https://github.com/ai-twinkle/eval/blob/main/main.py) module then handles persistence via `save_benchmark_results()`, which writes the sanitized `BenchmarkMetrics` dataclass to a timestamped JSON file, stripping live objects like `llm_instance` before serialization.

## Interpreting Benchmark Results

### JSON Output Structure

The benchmark generates a JSON file with the following structure:

```json
{
  "timestamp": "20240223_143200",
  "config": { "provider": "openai", "model": "gpt-3.5-turbo" },
  "metrics": {
    "throughput": {
      "requests_per_second": 9.85,
      "tokens_per_second": 58.3
    },
    "latency": {
      "mean": 0.103,
      "median": 0.098,
      "p95": 0.215,
      "p99": 0.298
    },
    "time_to_first_token": { "mean": 0.042, "median": 0.039 },
    "time_per_output_token": { "mean": 0.006, "median": 0.005 }
  },
  "summary": {
    "total_requests": 100,
    "successful_requests": 100,
    "total_tokens": 587,
    "total_duration": 10.14
  }
}

```

### Generating HTML Reports

Convert JSON results to a formatted HTML report using:

```bash
twinkle-eval --convert-to-html benchmark_results_20240223_143200.json

```

This command invokes `convert_json_to_html()` in [`twinkle_eval/main.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/main.py) (lines 19-58) and produces a standalone HTML file with visualized metrics tables and charts.

## Summary

- Execute `twinkle-eval --benchmark` to run performance benchmarks using the Twinkle Eval benchmark command with default settings (100 requests, 10 concurrent workers).
- Customize load testing using flags like `--benchmark-concurrency`, `--benchmark-rate`, and `--benchmark-duration` to control traffic patterns.
- The `BenchmarkRunner` class in [`twinkle_eval/benchmark.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/benchmark.py) orchestrates execution, while `_calculate_metrics()` computes latency percentiles, TTFT, TPOT, and throughput statistics.
- Results are automatically saved as timestamped JSON files and can be converted to HTML reports using `--convert-to-html`.

## Frequently Asked Questions

### What metrics does the Twinkle Eval benchmark command measure?

The benchmark captures **latency** (total request duration), **throughput** (requests and tokens per second), **Time-to-First-Token (TTFT)**, and **Time-per-Output-Token (TPOT)**. These metrics are calculated in [`twinkle_eval/benchmark.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/benchmark.py) by the `_calculate_metrics()` method after raw request data is collected by `_send_request()`.

### How do I limit the request rate when running benchmarks?

Use the `--benchmark-rate` flag to enable rate-limited mode. For example, `twinkle-eval --benchmark --benchmark-rate 2 --benchmark-concurrency 5` limits the test to 2 requests per second with 5 concurrent workers. According to the source code in [`twinkle_eval/benchmark.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/benchmark.py), this triggers `_run_rate_limited_benchmark()`, which uses a `Semaphore` to throttle requests.

### Where are benchmark results saved?

Results are automatically persisted to a JSON file named [`benchmark_results_YYYYMMDD_HHMMSS.json`](https://github.com/ai-twinkle/eval/blob/main/benchmark_results_YYYYMMDD_HHMMSS.json) in the current working directory. The `save_benchmark_results()` function in [`twinkle_eval/main.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/main.py) handles serialization, stripping live objects like `llm_instance` from the configuration before writing. You can convert these JSON files to HTML reports using `twinkle-eval --convert-to-html`.

### Can I use a custom prompt for benchmarking?

Yes, specify the `--benchmark-prompt` flag followed by your desired text. If omitted, the benchmark defaults to the Chinese prompt "台灣的首都是哪裡？" (What is the capital of Taiwan?). The prompt is passed to the `BenchmarkRunner` constructor in [`twinkle_eval/main.py`](https://github.com/ai-twinkle/eval/blob/main/twinkle_eval/main.py) and sent with every request during the load test.