# Evaluating AI Agent Performance Using Pass@k Metrics: A Complete Guide

> Master Pass@k metrics to evaluate AI agent performance. Learn how this guide demystifies probabilistic success for k independent attempts.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: how-to-guide
- Published: 2026-08-18

---

**Pass@k is a probabilistic success metric that measures whether an AI agent succeeds at least once when given k independent attempts at the same task, calculated as 1 minus (1 minus p) raised to the k power, where p represents the single-shot success probability.**

Pass@k has become the standard benchmark for evaluating AI agent reliability in both research and production settings. This metric originated in code generation [challenges](https://bojieli/ai-agent-book) like HumanEval but now extends across autonomous systems, robotics, and multi-step reasoning tasks. In the bojieli/ai-agent-book repository, Pass@k serves as the foundational evaluation principle throughout Chapter 7's metrics framework and Chapter 9's self-evolution experiments.

## Understanding the Pass@k Formula

The mathematical definition of Pass@k captures the probability of at least one success across k independent trials.

```

Pass@k = 1 - (1 - p)^k

```

Where **p** = single-attempt success probability and **k** = number of independent attempts.

Key properties of this formula:

- **Interpretability**: Pass@k directly answers "will my agent get this right if I let it try k times?"
- **Diminishing returns**: Each additional attempt yields smaller marginal gains as k increases
- **Scaling behavior**: Agents with modest p (0.3-0.5) can achieve high Pass@k at k=5 or k=10

In [`book/chapter7.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/chapter7.md) (lines 75-98), the repository's *Metrics Dictionary* explains this intuition alongside the closed-form derivation.

## Variants: Best@k and Pass Consecutive@k

The repository identifies two specialized variants for different evaluation goals.

### Best@k: Maximum Performance Ceiling

When an agent produces continuous scores rather than binary outcomes, **Best@k** reports the highest score achieved across k runs. This variant captures an agent's *potential* rather than its consistency—useful for capability studies where you want to know "how good can this agent get?"

### Pass⁽ᵏ⁾ (Pass Consecutive@k): Reliability Guarantee

The complementary metric **Pass⁽ᵏ⁾** measures k *consecutive* successes, calculated as p^k. This represents a far stricter reliability standard appropriate for safety-critical applications where every attempt must succeed. The probability drops exponentially with k: even an agent with p=0.9 has only 59% chance of 5 consecutive successes.

## Computing Pass@k in Practice: The Repository Implementation

Real-world evaluations in bojieli/ai-agent-book compute empirical Pass@k through batch task evaluation rather than theoretical probability estimation.

### Core Computation in Chapter 9

The file [`chapter9/hermes-self-evolution/run_downstream_ablation.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/hermes-self-evolution/run_downstream_ablation.py) (lines 274-281) implements the concrete evaluation logic:

```python

# From run_downstream_ablation.py - lines 274-281

baseline_pass_rate = baseline_passed / baseline_total if baseline_total > 0 else 0.0
evolved_pass_rate = evolved_passed / evolved_total if evolved_total > 0 else 0.0

pass_rate_uplift = evolved_pass_rate - baseline_pass_rate

results = {
    "baseline_pass_rate": baseline_pass_rate,
    "evolved_pass_rate": evolved_pass_rate,
    "pass_rate_uplift": pass_rate_uplift,
    # ... additional diagnostics

}

```

This code pattern treats `pass_rate` as the empirical analogue of Pass@k for a fixed k (the number of sampled trajectories per task).

### Validation Through Unit Tests

The repository validates these calculations in [`tests/test_ch9_hermes_downstream_ablation.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch9_hermes_downstream_ablation.py) (lines 48-51):

```python

# From test_ch9_hermes_downstream_ablation.py - validation assertions

assert 0.0 <= result["baseline_pass_rate"] <= 1.0
assert 0.0 <= result["evolved_pass_rate"] <= 1.0
assert result["pass_rate_uplift"] == (
    result["evolved_pass_rate"] - result["baseline_pass_rate"]
)

```

These assertions ensure pass rates remain in [0, 1] and that uplift arithmetic is correct.

## The Complete Evaluation Pipeline

According to the bojieli/ai-agent-book source code, Pass@k evaluation follows a five-stage pipeline:

1. **Task Sampling** — A benchmark (e.g., τ-bench, Hermes downstream) defines N independent tasks with ground-truth success criteria
2. **Repeated Execution** — For each task, the agent runs up to k times (k=1 for standard reporting, k>1 for "ability ceiling" studies)
3. **Pass Determination** — A task passes if any k run meets success criteria (safety rubric, correctness, or output validation)
4. **Aggregate Pass Rate** — `passed / N` yields empirical Pass@k for that agent configuration
5. **Comparative Reporting** — Baseline and evolved agents compared via `baseline_pass_rate`, `evolved_pass_rate`, and `pass_rate_uplift`

## Practical Python Implementation

### Theoretical Pass@k from Single-Shot Estimate

```python
def pass_at_k(p: float, k: int) -> float:
    """
    Return the theoretical Pass@k given per-try success probability p
    and number of attempts k.
    
    Args:
        p: Single-shot success probability (0 ≤ p ≤ 1)
        k: Number of independent attempts (k ≥ 1)
    
    Returns:
        Probability of at least one success in k attempts
    
    Raises:
        ValueError: If p outside [0,1] or k < 1
    
    Example:
        >>> pass_at_k(0.6, 5)
        0.9904
        >>> pass_at_k(0.3, 10)
        0.9718
    """
    if not (0.0 <= p <= 1.0):
        raise ValueError("p must be between 0 and 1")
    if k < 1:
        raise ValueError("k must be positive integer")
    
    return 1 - (1 - p) ** k

```

### Empirical Pass@k from Task Results

```python
from typing import List, Dict

def compute_pass_at_k(
    task_results: List[List[bool]], 
    k: int
) -> Dict[str, float]:
    """
    Compute empirical Pass@k from raw trajectory outcomes.
    
    Args:
        task_results: List of tasks, each containing list of 
                     success indicators for k attempts
        k: Maximum attempts to consider per task
    
    Returns:
        Dictionary with pass_rate, total_tasks, and attempts_used
    """
    if not task_results or k < 1:
        return {"pass_rate": 0.0, "total_tasks": 0, "attempts_used": 0}
    
    passed = sum(
        any(attempts[:k])  # Task passes if any attempt in first k succeeds

        for attempts in task_results
    )
    total = len(task_results)
    
    return {
        "pass_rate": passed / total,
        "total_tasks": total,
        "attempts_used": min(k, max(len(r) for r in task_results)),
        "tasks_passed": passed
    }


# Example usage with synthetic data

trajectories = [
    [False, True, False],   # Task 1: passed on 2nd attempt

    [False, False, True],   # Task 2: passed on 3rd attempt  

    [True, False, False],   # Task 3: passed on 1st attempt

    [False, False, False],  # Task 4: never passed

]

print(compute_pass_at_k(trajectories, k=1))  # Pass@1 = 0.25

print(compute_pass_at_k(trajectories, k=2))  # Pass@2 = 0.50

print(compute_pass_at_k(trajectories, k=3))  # Pass@3 = 0.75

```

## Choosing the Right k Value

| k | Use Case | Interpretation |
|---|----------|---------------|
| 1 | Production monitoring | "What fraction of tasks succeed on first try?" |
| 3-5 | Capability benchmarking | "Can the agent reliably solve this with moderate retries?" |
| 10+ | Ability ceiling studies | "What is the maximum success rate with generous attempts?" |
| 20+ | Safety validation | "Does this approach 100% with enough samples?" (asymptotic check) |

Higher k values reduce variance in estimates but increase evaluation cost linearly. The bojieli/ai-agent-book codebase typically uses k=1 for downstream ablation comparisons and k=5-10 for evolved agent "ceiling" evaluation.

## Summary

- **Pass@k = 1 - (1-p)^k** provides interpretable, probabilistic success measurement for AI agents given multiple attempts
- The **bojieli/ai-agent-book** repository implements Pass@k conceptually in [`book/chapter7.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/chapter7.md) and practically in [`chapter9/hermes-self-evolution/run_downstream_ablation.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/hermes-self-evolution/run_downstream_ablation.py)
- **Best@k** captures maximum performance potential; **Pass⁽ᵏ⁾** demands consecutive success for reliability-critical applications
- Empirical Pass@k computes as `tasks_passed / total_tasks` where each task passes if any of k attempts succeed
- Unit tests in [`tests/test_ch9_hermes_downstream_ablation.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch9_hermes_downstream_ablation.py) validate pass rate bounds and uplift calculations

## Frequently Asked Questions

### How does Pass@k differ from simple accuracy?

Pass@k specifically accounts for multiple independent attempts at identical tasks, while accuracy typically measures single-attempt performance. An agent with 40% single-shot accuracy achieves 87% Pass@5, making Pass@k more informative for retry-tolerant applications. The gap between accuracy and Pass@k reveals an agent's consistency—narrow gaps indicate reliable agents, wide gaps suggest high variance in capability.

### What k value should I use for production AI agents?

Use **k=1** for production latency-sensitive systems where retries are expensive or impractical. Use **k=3-5** for offline batch processing where quality matters more than speed. Research evaluation often employs **k=10 or higher** to estimate theoretical capability ceilings. The choice depends on your retry budget and the cost structure of agent execution versus failure consequences.

### Can Pass@k be applied beyond code generation tasks?

Yes. Pass@k generalizes to any task with verifiable success criteria: robotic manipulation (success = goal state reached), dialogue systems (success = user satisfaction threshold), theorem proving (success = proof verified), and multi-step planning (success = plan executable). The repository's Hermes self-evolution experiments demonstrate this flexibility across diverse downstream benchmarks.