# How to Evaluate LLM Outputs with the OpenAI Evals Framework

> Learn how to evaluate LLM outputs using the OpenAI Evals framework. This modular Python architecture offers custom classes, automated runners, and pluggable metrics for rigorous assessment.

- Repository: [Owain Lewis/awesome-artificial-intelligence](https://github.com/owainlewis/awesome-artificial-intelligence)
- Tags: how-to-guide
- Published: 2026-06-20

---

**The OpenAI Evals framework provides a modular, extensible Python architecture for evaluating large language model outputs through custom evaluation classes, automated runners, and pluggable metrics.**

The OpenAI Evals framework offers a systematic approach to evaluate LLM outputs with reproducible, automated testing pipelines. Listed in the `owainlewis/awesome-artificial-intelligence` repository at line 82 of [`README.md`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/README.md) as a core resource for AI engineering, this open-source toolkit separates evaluation logic from execution mechanics. Whether you are validating question-answer accuracy or measuring semantic similarity, the framework enables continuous monitoring of model performance across any OpenAI API-compatible endpoint.

## Understanding the OpenAI Evals Architecture

The framework separates three distinct concerns to ensure clean, maintainable evaluation code.

### Eval Definition

An **Eval** is a Python class that encodes the task definition, including prompt templates, input data, and scoring logic. To create an evaluation, you subclass `Eval` from `openai.evals` and implement the `run` and `evaluate` methods. These methods define how inputs are processed, how the LLM is queried, and how responses are scored against ground truth.

### Runner and Execution

The **Runner** handles the operational complexity of executing LLM calls at scale. The `run_eval` function from `openai.evals.runner` manages batching, parallelism, and API interaction, allowing you to process entire datasets without manual loop management or rate-limit handling.

### Results Storage and Reporting

The **Results Store** persists scores and metadata for longitudinal analysis. The framework supports CSV files, JSON, SQLite, or cloud storage backends. It records the exact request payload—including model name, temperature, and token limits—alongside the raw completion and computed score, ensuring full reproducibility.

## Building a Custom Evaluation Class

To evaluate LLM outputs with OpenAI Evals, define a concrete evaluation class that specifies your task logic. The following example implements exact-match accuracy for a question-answering task using the `gpt-4o` model:

```python

# example_eval.py – a minimal OpenAI Evals definition

from openai import OpenAI
from openai.evals import Eval, run_eval, Accuracy

class SimpleQAEval(Eval):
    """Evaluate a QA pair using exact-match accuracy."""
    def __init__(self):
        self.client = OpenAI()
        self.metric = Accuracy()

    def run(self, prompt: str, expected: str):
        # Call the LLM (e.g., gpt-4o)

        response = self.client.chat.completions.create(
            model="gpt-4o",
            messages=[{"role": "user", "content": prompt}],
            temperature=0.0,
        )
        answer = response.choices[0].message.content.strip()
        # Compute the metric

        score = self.metric.compare(answer, expected)
        return {"prompt": prompt, "answer": answer, "expected": expected, "score": score}

# Execute the evaluation on a list of QA items

if __name__ == "__main__":
    eval = SimpleQAEval()
    dataset = [
        {"prompt": "What is the capital of France?", "expected": "Paris"},
        {"prompt": "Who wrote \"Pride and Prejudice\"?", "expected": "Jane Austen"},
    ]
    results = run_eval(eval, dataset, output_path="results.csv")
    print("Saved results to results.csv")

```

## Implementing Custom Metrics

While the framework provides built-in metrics like `Accuracy` and `BLEU`, production evaluations often require domain-specific scoring logic. You can extend the base `Metric` class to implement bespoke evaluation criteria, such as semantic similarity using sentence embeddings:

```python

# custom_metric.py – defining a bespoke similarity metric

from openai.evals import Metric
import numpy as np
from sentence_transformers import SentenceTransformer, util

class CosineSimilarityMetric(Metric):
    def __init__(self, model_name="all-MiniLM-L6-v2"):
        self.embedder = SentenceTransformer(model_name)

    def compare(self, answer: str, reference: str) -> float:
        a_emb = self.embedder.encode(answer, convert_to_tensor=True)
        r_emb = self.embedder.encode(reference, convert_to_tensor=True)
        return util.cos_sim(a_emb, r_emb).item()

# Use the custom metric in an eval

from example_eval import SimpleQAEval
from custom_metric import CosineSimilarityMetric

class SimilarityQAEval(SimpleQAEval):
    def __init__(self):
        self.client = OpenAI()
        self.metric = CosineSimilarityMetric()

```

## Running Evaluations at Scale

The framework supports reproducible evaluation workflows by capturing complete request metadata. When integrating into CI pipelines, you can regression-test LLM behavior by comparing current scores against historical baselines stored in the Results Store. Because the framework is open-source, you can also swap the underlying model provider or add new data loaders to evaluate LLM outputs against proprietary datasets.

## Summary

- The OpenAI Evals framework provides a modular architecture separating **Eval definitions**, **Runners**, and **Reporting** components.
- Create evaluation tasks by subclassing `Eval` and implementing `run` and `evaluate` methods to query models and compute scores.
- Use `run_eval` from `openai.evals.runner` to handle batching, parallelism, and API execution.
- Implement custom metrics by extending the `Metric` class and defining the `compare` method for domain-specific scoring.
- Evaluations are fully reproducible because the framework captures complete request payloads, raw completions, and computed scores in the Results Store.

## Frequently Asked Questions

### What is the OpenAI Evals framework?

The OpenAI Evals framework is an open-source Python toolkit for systematically evaluating large language model outputs. As referenced in the `owainlewis/awesome-artificial-intelligence` repository, it provides a structured way to define evaluation tasks, execute them against OpenAI API endpoints, and score results using built-in or custom metrics.

### How do I create a custom metric in OpenAI Evals?

Create a custom metric by subclassing `Metric` from `openai.evals` and implementing the `compare` method. This method accepts the model's generated answer and a reference text, returning a numerical score. You then instantiate this metric class within your Eval's `__init__` method to evaluate LLM outputs with domain-specific criteria.

### Can I use OpenAI Evals with models other than GPT-4?

Yes. The framework works with any model supported by the OpenAI API client. By modifying the `model` parameter in your API call within the `run` method, you can evaluate LLM outputs from GPT-3.5, GPT-4o, or other compatible endpoints, making the framework provider-agnostic within the OpenAI ecosystem.

### How does OpenAI Evals ensure reproducibility?

The framework ensures reproducibility by recording the complete request payload—including model parameters like `temperature`, `max_tokens`, and the model name—alongside the raw completion and computed score. This metadata is stored in the Results Store (CSV, JSON, or database), allowing you to reproduce exact evaluation conditions and track performance changes across model versions or prompt updates.