# How to Use OpenAI Evals for AI Model Evaluation: A Complete Guide

> Learn how to use OpenAI Evals a powerful Python framework for AI model evaluation. Build custom suites run evaluations and generate detailed JSON reports with ease. Master AI model assessment today.

- Repository: [Owain Lewis/awesome-artificial-intelligence](https://github.com/owainlewis/awesome-artificial-intelligence)
- Tags: how-to-guide
- Published: 2026-06-22

---

**OpenAI Evals is a lightweight Python framework that lets you define custom evaluation suites by extending the `Eval` base class, running them via CLI or Python API to generate detailed JSON reports with quantitative scores.**

OpenAI Evals provides a structured, extensible pipeline for benchmarking large language models against custom datasets and metrics. According to the OpenAI Evals repository, this framework separates model adapters, datasets, and evaluation metrics into modular components, enabling reproducible testing of AI systems through simple Python classes and command-line tools.

## Installation and Configuration

### Install the Package

Install the library with the optional extras required for running evaluations:

```bash
pip install "openai[evals]"

```

This command pulls the core OpenAI library along with the evaluation-specific dependencies defined in the package configuration.

### Configure Authentication

The framework reads your API key automatically from the environment. In [`evals/__init__.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/__init__.py), the package loads the `OPENAI_API_KEY` environment variable to authenticate requests:

```bash
export OPENAI_API_KEY="sk-..."

```

You can also pass keys explicitly to model constructors when working programmatically, though the environment variable approach is the standard method used by the CLI runner.

## Defining a Custom Eval

An evaluation is a regular Python module that extends `evals.base.Eval`. The [`evals/base.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/base.py) file defines the abstract base class that provides `self.model`, `self.dataset`, and the `record()` helper method for logging results.

Create a minimal eval by subclassing `Eval` and implementing the `run()` method:

```python

# my_eval.py

from evals.base import Eval
from evals.record import Record

class MyEval(Eval):
    def __init__(self, model, dataset):
        super().__init__(model=model, dataset=dataset)

    def run(self):
        for item in self.dataset:
            # 1. Prompt the model

            response = self.model.generate(item["prompt"])
            
            # 2. Score the response

            score = self.metric(response, item["expected"])
            
            # 3. Record the result

            self.record(Record(
                prompt=item["prompt"],
                completion=response,
                expected=item["expected"],
                score=score
            ))

    @staticmethod
    def metric(completion, expected):
        # Simple exact-match metric; replace with any custom logic

        return 1.0 if completion.strip() == expected.strip() else 0.0

```

The [`evals/registry.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/registry.py) module handles automatic discovery of eval classes from dotted paths like `my_eval.MyEval`, while [`evals/metrics.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/metrics.py) provides ready-made implementations for accuracy, BLEU, ROUGE, and other standard metrics. For dataset loading, [`evals/utils.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/utils.py) contains helpers that read JSONL, CSV, and other common formats.

## Running Evaluations via CLI and Python API

### Command-Line Interface

The [`evals/run.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/run.py) module provides a CLI wrapper that constructs models, loads datasets, and executes evals. From the terminal:

```bash
python -m evals.run \
    --model openai:gpt-4o \
    --eval my_eval.MyEval \
    --dataset path/to/my_dataset.jsonl \
    --output results/my_eval_report.json

```

Internally, this command:
1. Builds a model wrapper via [`evals/models/openai.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/models/openai.py) (the `OpenAIModel` class)
2. Loads your dataset using `evals.utils.load_dataset`
3. Executes the `run()` method of your eval class
4. Serializes results to JSON following the schema defined in [`evals/report.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/report.py)

### Python API

You can also execute evaluations programmatically for integration into larger workflows:

```python
from evals.run import run_eval
from evals.models.openai import OpenAIModel
from my_eval import MyEval

model = OpenAIModel(name="gpt-4o")
eval_instance = MyEval(model=model, dataset="path/to/my_dataset.jsonl")
run_eval(eval_instance, output_path="results/report.json")

```

The resulting JSON report contains all prompts, completions, individual scores, and aggregate statistics, making it easy to ingest into dashboards or analysis pipelines.

## Advanced Features and Implementation

### Parallel Execution

For large datasets, enable parallel processing by setting the `--workers` flag in the CLI or passing `workers=N` to `run_eval()`. The implementation in [`evals/run.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/run.py) handles process pooling and result aggregation automatically.

### Built-In Metrics

Instead of writing custom scoring logic, import metrics directly from [`evals/metrics.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/metrics.py). For example, use the BLEU implementation for translation tasks:

```python

# translation_eval.py

from evals.base import Eval
from evals.metrics import BLEU
from evals.record import Record

bleu = BLEU()

class TranslationEval(Eval):
    def run(self):
        for ex in self.dataset:
            output = self.model.generate(ex["src"])
            score = bleu(output, ex["ref"])
            self.record(Record(
                prompt=ex["src"],
                completion=output,
                expected=ex["ref"],
                score=score,
            ))

```

Run this eval with:

```bash
python -m evals.run \
    --model openai:gpt-4o \
    --eval translation_eval.TranslationEval \
    --dataset data/wmt_en_de.jsonl \
    --output reports/translation.json

```

### Question-Answering Example

For zero-shot QA tasks, implement a simple accuracy-based eval:

```python

# qa_eval.py

from evals.base import Eval
from evals.record import Record

class QAEval(Eval):
    def run(self):
        for qa in self.dataset:
            answer = self.model.generate(qa["question"])
            acc = 1.0 if answer.strip().lower() == qa["answer"].strip().lower() else 0.0
            self.record(Record(
                prompt=qa["question"],
                completion=answer,
                expected=qa["answer"],
                score=acc,
            ))

```

### Additional Capabilities

- **Custom Metrics**: Subclass `evals.metrics.BaseMetric` or pass any callable scoring function to support domain-specific evaluation criteria
- **Logging**: Control verbosity via `--log-level` flags; [`evals/logging.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/logging.py) structures output for debugging
- **Embeddings**: Use [`evals/embeddings.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/embeddings.py) utilities to compute cosine similarity or RAG retrieval metrics for vector-based evaluation scenarios

## Summary

- **OpenAI Evals** provides a modular framework for AI model evaluation through the `evals.base.Eval` class in [`evals/base.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/base.py)
- Install via `pip install "openai[evals]"` and configure authentication through the `OPENAI_API_KEY` environment variable
- Define evaluations by subclassing `Eval`, implementing the `run()` method, and using `self.record()` to log results
- Execute evaluations using `python -m evals.run` (CLI) or `run_eval()` (Python API), with outputs serialized to JSON per [`evals/report.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/report.py)
- Leverage built-in metrics from [`evals/metrics.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/metrics.py) (BLEU, ROUGE, accuracy) and parallel execution via `--workers` for scalable benchmarking

## Frequently Asked Questions

### What is OpenAI Evals used for?

OpenAI Evals is a framework for systematically evaluating language models, embeddings, or any AI system accessible via API. It standardizes the process of prompting models, scoring outputs, and generating structured reports, making it suitable for regression testing, benchmarking, and research.

### How do I define a custom metric in OpenAI Evals?

You can define a custom metric by creating a callable function (like the `metric` static method in the examples above) or by subclassing `evals.metrics.BaseMetric` in [`evals/metrics.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/metrics.py). Pass your metric function to the `score` parameter when recording results, or instantiate built-in metrics like `BLEU()` directly in your eval class.

### Can I run evaluations without using the OpenAI API?

Yes. While the default model wrapper in [`evals/models/openai.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/models/openai.py) targets OpenAI's API, the `Eval` base class accepts any object with a `generate()` method. You can implement custom model adapters for other providers (Anthropic, local models, etc.) and pass them to your eval's constructor when running programmatically.

### How do I parallelize evaluation runs?

Set the `--workers N` flag when using the CLI, or pass `workers=N` to the `run_eval()` function in Python. The framework in [`evals/run.py`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/evals/run.py) manages process pools to distribute dataset items across multiple workers, significantly speeding up evaluation for large datasets while maintaining thread-safe result collection.