How to Use OpenAI Evals for AI Model Evaluation: A Complete Guide

OpenAI Evals is a lightweight Python framework that lets you define custom evaluation suites by extending the Eval base class, running them via CLI or Python API to generate detailed JSON reports with quantitative scores.

OpenAI Evals provides a structured, extensible pipeline for benchmarking large language models against custom datasets and metrics. According to the OpenAI Evals repository, this framework separates model adapters, datasets, and evaluation metrics into modular components, enabling reproducible testing of AI systems through simple Python classes and command-line tools.

Installation and Configuration

Install the Package

Install the library with the optional extras required for running evaluations:

pip install "openai[evals]"

This command pulls the core OpenAI library along with the evaluation-specific dependencies defined in the package configuration.

Configure Authentication

The framework reads your API key automatically from the environment. In evals/__init__.py, the package loads the OPENAI_API_KEY environment variable to authenticate requests:

export OPENAI_API_KEY="sk-..."

You can also pass keys explicitly to model constructors when working programmatically, though the environment variable approach is the standard method used by the CLI runner.

Defining a Custom Eval

An evaluation is a regular Python module that extends evals.base.Eval. The evals/base.py file defines the abstract base class that provides self.model, self.dataset, and the record() helper method for logging results.

Create a minimal eval by subclassing Eval and implementing the run() method:


# my_eval.py

from evals.base import Eval
from evals.record import Record

class MyEval(Eval):
    def __init__(self, model, dataset):
        super().__init__(model=model, dataset=dataset)

    def run(self):
        for item in self.dataset:
            # 1. Prompt the model

            response = self.model.generate(item["prompt"])
            
            # 2. Score the response

            score = self.metric(response, item["expected"])
            
            # 3. Record the result

            self.record(Record(
                prompt=item["prompt"],
                completion=response,
                expected=item["expected"],
                score=score
            ))

    @staticmethod
    def metric(completion, expected):
        # Simple exact-match metric; replace with any custom logic

        return 1.0 if completion.strip() == expected.strip() else 0.0

The evals/registry.py module handles automatic discovery of eval classes from dotted paths like my_eval.MyEval, while evals/metrics.py provides ready-made implementations for accuracy, BLEU, ROUGE, and other standard metrics. For dataset loading, evals/utils.py contains helpers that read JSONL, CSV, and other common formats.

Running Evaluations via CLI and Python API

Command-Line Interface

The evals/run.py module provides a CLI wrapper that constructs models, loads datasets, and executes evals. From the terminal:

python -m evals.run \
    --model openai:gpt-4o \
    --eval my_eval.MyEval \
    --dataset path/to/my_dataset.jsonl \
    --output results/my_eval_report.json

Internally, this command:

  1. Builds a model wrapper via evals/models/openai.py (the OpenAIModel class)
  2. Loads your dataset using evals.utils.load_dataset
  3. Executes the run() method of your eval class
  4. Serializes results to JSON following the schema defined in evals/report.py

Python API

You can also execute evaluations programmatically for integration into larger workflows:

from evals.run import run_eval
from evals.models.openai import OpenAIModel
from my_eval import MyEval

model = OpenAIModel(name="gpt-4o")
eval_instance = MyEval(model=model, dataset="path/to/my_dataset.jsonl")
run_eval(eval_instance, output_path="results/report.json")

The resulting JSON report contains all prompts, completions, individual scores, and aggregate statistics, making it easy to ingest into dashboards or analysis pipelines.

Advanced Features and Implementation

Parallel Execution

For large datasets, enable parallel processing by setting the --workers flag in the CLI or passing workers=N to run_eval(). The implementation in evals/run.py handles process pooling and result aggregation automatically.

Built-In Metrics

Instead of writing custom scoring logic, import metrics directly from evals/metrics.py. For example, use the BLEU implementation for translation tasks:


# translation_eval.py

from evals.base import Eval
from evals.metrics import BLEU
from evals.record import Record

bleu = BLEU()

class TranslationEval(Eval):
    def run(self):
        for ex in self.dataset:
            output = self.model.generate(ex["src"])
            score = bleu(output, ex["ref"])
            self.record(Record(
                prompt=ex["src"],
                completion=output,
                expected=ex["ref"],
                score=score,
            ))

Run this eval with:

python -m evals.run \
    --model openai:gpt-4o \
    --eval translation_eval.TranslationEval \
    --dataset data/wmt_en_de.jsonl \
    --output reports/translation.json

Question-Answering Example

For zero-shot QA tasks, implement a simple accuracy-based eval:


# qa_eval.py

from evals.base import Eval
from evals.record import Record

class QAEval(Eval):
    def run(self):
        for qa in self.dataset:
            answer = self.model.generate(qa["question"])
            acc = 1.0 if answer.strip().lower() == qa["answer"].strip().lower() else 0.0
            self.record(Record(
                prompt=qa["question"],
                completion=answer,
                expected=qa["answer"],
                score=acc,
            ))

Additional Capabilities

  • Custom Metrics: Subclass evals.metrics.BaseMetric or pass any callable scoring function to support domain-specific evaluation criteria
  • Logging: Control verbosity via --log-level flags; evals/logging.py structures output for debugging
  • Embeddings: Use evals/embeddings.py utilities to compute cosine similarity or RAG retrieval metrics for vector-based evaluation scenarios

Summary

  • OpenAI Evals provides a modular framework for AI model evaluation through the evals.base.Eval class in evals/base.py
  • Install via pip install "openai[evals]" and configure authentication through the OPENAI_API_KEY environment variable
  • Define evaluations by subclassing Eval, implementing the run() method, and using self.record() to log results
  • Execute evaluations using python -m evals.run (CLI) or run_eval() (Python API), with outputs serialized to JSON per evals/report.py
  • Leverage built-in metrics from evals/metrics.py (BLEU, ROUGE, accuracy) and parallel execution via --workers for scalable benchmarking

Frequently Asked Questions

What is OpenAI Evals used for?

OpenAI Evals is a framework for systematically evaluating language models, embeddings, or any AI system accessible via API. It standardizes the process of prompting models, scoring outputs, and generating structured reports, making it suitable for regression testing, benchmarking, and research.

How do I define a custom metric in OpenAI Evals?

You can define a custom metric by creating a callable function (like the metric static method in the examples above) or by subclassing evals.metrics.BaseMetric in evals/metrics.py. Pass your metric function to the score parameter when recording results, or instantiate built-in metrics like BLEU() directly in your eval class.

Can I run evaluations without using the OpenAI API?

Yes. While the default model wrapper in evals/models/openai.py targets OpenAI's API, the Eval base class accepts any object with a generate() method. You can implement custom model adapters for other providers (Anthropic, local models, etc.) and pass them to your eval's constructor when running programmatically.

How do I parallelize evaluation runs?

Set the --workers N flag when using the CLI, or pass workers=N to the run_eval() function in Python. The framework in evals/run.py manages process pools to distribute dataset items across multiple workers, significantly speeding up evaluation for large datasets while maintaining thread-safe result collection.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →