Evaluating AI Agent Performance Using Pass@k Metrics: A Complete Guide

Pass@k is a probabilistic success metric that measures whether an AI agent succeeds at least once when given k independent attempts at the same task, calculated as 1 minus (1 minus p) raised to the k power, where p represents the single-shot success probability.

Pass@k has become the standard benchmark for evaluating AI agent reliability in both research and production settings. This metric originated in code generation challenges like HumanEval but now extends across autonomous systems, robotics, and multi-step reasoning tasks. In the bojieli/ai-agent-book repository, Pass@k serves as the foundational evaluation principle throughout Chapter 7's metrics framework and Chapter 9's self-evolution experiments.

Understanding the Pass@k Formula

The mathematical definition of Pass@k captures the probability of at least one success across k independent trials.


Pass@k = 1 - (1 - p)^k

Where p = single-attempt success probability and k = number of independent attempts.

Key properties of this formula:

  • Interpretability: Pass@k directly answers "will my agent get this right if I let it try k times?"
  • Diminishing returns: Each additional attempt yields smaller marginal gains as k increases
  • Scaling behavior: Agents with modest p (0.3-0.5) can achieve high Pass@k at k=5 or k=10

In book/chapter7.md (lines 75-98), the repository's Metrics Dictionary explains this intuition alongside the closed-form derivation.

Variants: Best@k and Pass Consecutive@k

The repository identifies two specialized variants for different evaluation goals.

Best@k: Maximum Performance Ceiling

When an agent produces continuous scores rather than binary outcomes, Best@k reports the highest score achieved across k runs. This variant captures an agent's potential rather than its consistency—useful for capability studies where you want to know "how good can this agent get?"

Pass⁽ᵏ⁾ (Pass Consecutive@k): Reliability Guarantee

The complementary metric Pass⁽ᵏ⁾ measures k consecutive successes, calculated as p^k. This represents a far stricter reliability standard appropriate for safety-critical applications where every attempt must succeed. The probability drops exponentially with k: even an agent with p=0.9 has only 59% chance of 5 consecutive successes.

Computing Pass@k in Practice: The Repository Implementation

Real-world evaluations in bojieli/ai-agent-book compute empirical Pass@k through batch task evaluation rather than theoretical probability estimation.

Core Computation in Chapter 9

The file chapter9/hermes-self-evolution/run_downstream_ablation.py (lines 274-281) implements the concrete evaluation logic:


# From run_downstream_ablation.py - lines 274-281

baseline_pass_rate = baseline_passed / baseline_total if baseline_total > 0 else 0.0
evolved_pass_rate = evolved_passed / evolved_total if evolved_total > 0 else 0.0

pass_rate_uplift = evolved_pass_rate - baseline_pass_rate

results = {
    "baseline_pass_rate": baseline_pass_rate,
    "evolved_pass_rate": evolved_pass_rate,
    "pass_rate_uplift": pass_rate_uplift,
    # ... additional diagnostics

}

This code pattern treats pass_rate as the empirical analogue of Pass@k for a fixed k (the number of sampled trajectories per task).

Validation Through Unit Tests

The repository validates these calculations in tests/test_ch9_hermes_downstream_ablation.py (lines 48-51):


# From test_ch9_hermes_downstream_ablation.py - validation assertions

assert 0.0 <= result["baseline_pass_rate"] <= 1.0
assert 0.0 <= result["evolved_pass_rate"] <= 1.0
assert result["pass_rate_uplift"] == (
    result["evolved_pass_rate"] - result["baseline_pass_rate"]
)

These assertions ensure pass rates remain in [0, 1] and that uplift arithmetic is correct.

The Complete Evaluation Pipeline

According to the bojieli/ai-agent-book source code, Pass@k evaluation follows a five-stage pipeline:

  1. Task Sampling — A benchmark (e.g., τ-bench, Hermes downstream) defines N independent tasks with ground-truth success criteria
  2. Repeated Execution — For each task, the agent runs up to k times (k=1 for standard reporting, k>1 for "ability ceiling" studies)
  3. Pass Determination — A task passes if any k run meets success criteria (safety rubric, correctness, or output validation)
  4. Aggregate Pass Rate — passed / N yields empirical Pass@k for that agent configuration
  5. Comparative Reporting — Baseline and evolved agents compared via baseline_pass_rate, evolved_pass_rate, and pass_rate_uplift

Practical Python Implementation

Theoretical Pass@k from Single-Shot Estimate

def pass_at_k(p: float, k: int) -> float:
    """
    Return the theoretical Pass@k given per-try success probability p
    and number of attempts k.
    
    Args:
        p: Single-shot success probability (0 ≤ p ≤ 1)
        k: Number of independent attempts (k ≥ 1)
    
    Returns:
        Probability of at least one success in k attempts
    
    Raises:
        ValueError: If p outside [0,1] or k < 1
    
    Example:
        >>> pass_at_k(0.6, 5)
        0.9904
        >>> pass_at_k(0.3, 10)
        0.9718
    """
    if not (0.0 <= p <= 1.0):
        raise ValueError("p must be between 0 and 1")
    if k < 1:
        raise ValueError("k must be positive integer")
    
    return 1 - (1 - p) ** k

Empirical Pass@k from Task Results

from typing import List, Dict

def compute_pass_at_k(
    task_results: List[List[bool]], 
    k: int
) -> Dict[str, float]:
    """
    Compute empirical Pass@k from raw trajectory outcomes.
    
    Args:
        task_results: List of tasks, each containing list of 
                     success indicators for k attempts
        k: Maximum attempts to consider per task
    
    Returns:
        Dictionary with pass_rate, total_tasks, and attempts_used
    """
    if not task_results or k < 1:
        return {"pass_rate": 0.0, "total_tasks": 0, "attempts_used": 0}
    
    passed = sum(
        any(attempts[:k])  # Task passes if any attempt in first k succeeds

        for attempts in task_results
    )
    total = len(task_results)
    
    return {
        "pass_rate": passed / total,
        "total_tasks": total,
        "attempts_used": min(k, max(len(r) for r in task_results)),
        "tasks_passed": passed
    }


# Example usage with synthetic data

trajectories = [
    [False, True, False],   # Task 1: passed on 2nd attempt

    [False, False, True],   # Task 2: passed on 3rd attempt  

    [True, False, False],   # Task 3: passed on 1st attempt

    [False, False, False],  # Task 4: never passed

]

print(compute_pass_at_k(trajectories, k=1))  # Pass@1 = 0.25

print(compute_pass_at_k(trajectories, k=2))  # Pass@2 = 0.50

print(compute_pass_at_k(trajectories, k=3))  # Pass@3 = 0.75

Choosing the Right k Value

k Use Case Interpretation
1 Production monitoring "What fraction of tasks succeed on first try?"
3-5 Capability benchmarking "Can the agent reliably solve this with moderate retries?"
10+ Ability ceiling studies "What is the maximum success rate with generous attempts?"
20+ Safety validation "Does this approach 100% with enough samples?" (asymptotic check)

Higher k values reduce variance in estimates but increase evaluation cost linearly. The bojieli/ai-agent-book codebase typically uses k=1 for downstream ablation comparisons and k=5-10 for evolved agent "ceiling" evaluation.

Summary

  • Pass@k = 1 - (1-p)^k provides interpretable, probabilistic success measurement for AI agents given multiple attempts
  • The bojieli/ai-agent-book repository implements Pass@k conceptually in book/chapter7.md and practically in chapter9/hermes-self-evolution/run_downstream_ablation.py
  • Best@k captures maximum performance potential; Pass⁽ᵏ⁾ demands consecutive success for reliability-critical applications
  • Empirical Pass@k computes as tasks_passed / total_tasks where each task passes if any of k attempts succeed
  • Unit tests in tests/test_ch9_hermes_downstream_ablation.py validate pass rate bounds and uplift calculations

Frequently Asked Questions

How does Pass@k differ from simple accuracy?

Pass@k specifically accounts for multiple independent attempts at identical tasks, while accuracy typically measures single-attempt performance. An agent with 40% single-shot accuracy achieves 87% Pass@5, making Pass@k more informative for retry-tolerant applications. The gap between accuracy and Pass@k reveals an agent's consistency—narrow gaps indicate reliable agents, wide gaps suggest high variance in capability.

What k value should I use for production AI agents?

Use k=1 for production latency-sensitive systems where retries are expensive or impractical. Use k=3-5 for offline batch processing where quality matters more than speed. Research evaluation often employs k=10 or higher to estimate theoretical capability ceilings. The choice depends on your retry budget and the cost structure of agent execution versus failure consequences.

Can Pass@k be applied beyond code generation tasks?

Yes. Pass@k generalizes to any task with verifiable success criteria: robotic manipulation (success = goal state reached), dialogue systems (success = user satisfaction threshold), theorem proving (success = proof verified), and multi-step planning (success = plan executable). The repository's Hermes self-evolution experiments demonstrate this flexibility across diverse downstream benchmarks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →