How to Use OpenAI Evals for AI Model Evaluation: A Complete Guide
OpenAI Evals is a lightweight Python framework that lets you define custom evaluation suites by extending the Eval base class, running them via CLI or Python API to generate detailed JSON reports with quantitative scores.
OpenAI Evals provides a structured, extensible pipeline for benchmarking large language models against custom datasets and metrics. According to the OpenAI Evals repository, this framework separates model adapters, datasets, and evaluation metrics into modular components, enabling reproducible testing of AI systems through simple Python classes and command-line tools.
Installation and Configuration
Install the Package
Install the library with the optional extras required for running evaluations:
pip install "openai[evals]"
This command pulls the core OpenAI library along with the evaluation-specific dependencies defined in the package configuration.
Configure Authentication
The framework reads your API key automatically from the environment. In evals/__init__.py, the package loads the OPENAI_API_KEY environment variable to authenticate requests:
export OPENAI_API_KEY="sk-..."
You can also pass keys explicitly to model constructors when working programmatically, though the environment variable approach is the standard method used by the CLI runner.
Defining a Custom Eval
An evaluation is a regular Python module that extends evals.base.Eval. The evals/base.py file defines the abstract base class that provides self.model, self.dataset, and the record() helper method for logging results.
Create a minimal eval by subclassing Eval and implementing the run() method:
# my_eval.py
from evals.base import Eval
from evals.record import Record
class MyEval(Eval):
def __init__(self, model, dataset):
super().__init__(model=model, dataset=dataset)
def run(self):
for item in self.dataset:
# 1. Prompt the model
response = self.model.generate(item["prompt"])
# 2. Score the response
score = self.metric(response, item["expected"])
# 3. Record the result
self.record(Record(
prompt=item["prompt"],
completion=response,
expected=item["expected"],
score=score
))
@staticmethod
def metric(completion, expected):
# Simple exact-match metric; replace with any custom logic
return 1.0 if completion.strip() == expected.strip() else 0.0
The evals/registry.py module handles automatic discovery of eval classes from dotted paths like my_eval.MyEval, while evals/metrics.py provides ready-made implementations for accuracy, BLEU, ROUGE, and other standard metrics. For dataset loading, evals/utils.py contains helpers that read JSONL, CSV, and other common formats.
Running Evaluations via CLI and Python API
Command-Line Interface
The evals/run.py module provides a CLI wrapper that constructs models, loads datasets, and executes evals. From the terminal:
python -m evals.run \
--model openai:gpt-4o \
--eval my_eval.MyEval \
--dataset path/to/my_dataset.jsonl \
--output results/my_eval_report.json
Internally, this command:
- Builds a model wrapper via
evals/models/openai.py(theOpenAIModelclass) - Loads your dataset using
evals.utils.load_dataset - Executes the
run()method of your eval class - Serializes results to JSON following the schema defined in
evals/report.py
Python API
You can also execute evaluations programmatically for integration into larger workflows:
from evals.run import run_eval
from evals.models.openai import OpenAIModel
from my_eval import MyEval
model = OpenAIModel(name="gpt-4o")
eval_instance = MyEval(model=model, dataset="path/to/my_dataset.jsonl")
run_eval(eval_instance, output_path="results/report.json")
The resulting JSON report contains all prompts, completions, individual scores, and aggregate statistics, making it easy to ingest into dashboards or analysis pipelines.
Advanced Features and Implementation
Parallel Execution
For large datasets, enable parallel processing by setting the --workers flag in the CLI or passing workers=N to run_eval(). The implementation in evals/run.py handles process pooling and result aggregation automatically.
Built-In Metrics
Instead of writing custom scoring logic, import metrics directly from evals/metrics.py. For example, use the BLEU implementation for translation tasks:
# translation_eval.py
from evals.base import Eval
from evals.metrics import BLEU
from evals.record import Record
bleu = BLEU()
class TranslationEval(Eval):
def run(self):
for ex in self.dataset:
output = self.model.generate(ex["src"])
score = bleu(output, ex["ref"])
self.record(Record(
prompt=ex["src"],
completion=output,
expected=ex["ref"],
score=score,
))
Run this eval with:
python -m evals.run \
--model openai:gpt-4o \
--eval translation_eval.TranslationEval \
--dataset data/wmt_en_de.jsonl \
--output reports/translation.json
Question-Answering Example
For zero-shot QA tasks, implement a simple accuracy-based eval:
# qa_eval.py
from evals.base import Eval
from evals.record import Record
class QAEval(Eval):
def run(self):
for qa in self.dataset:
answer = self.model.generate(qa["question"])
acc = 1.0 if answer.strip().lower() == qa["answer"].strip().lower() else 0.0
self.record(Record(
prompt=qa["question"],
completion=answer,
expected=qa["answer"],
score=acc,
))
Additional Capabilities
- Custom Metrics: Subclass
evals.metrics.BaseMetricor pass any callable scoring function to support domain-specific evaluation criteria - Logging: Control verbosity via
--log-levelflags;evals/logging.pystructures output for debugging - Embeddings: Use
evals/embeddings.pyutilities to compute cosine similarity or RAG retrieval metrics for vector-based evaluation scenarios
Summary
- OpenAI Evals provides a modular framework for AI model evaluation through the
evals.base.Evalclass inevals/base.py - Install via
pip install "openai[evals]"and configure authentication through theOPENAI_API_KEYenvironment variable - Define evaluations by subclassing
Eval, implementing therun()method, and usingself.record()to log results - Execute evaluations using
python -m evals.run(CLI) orrun_eval()(Python API), with outputs serialized to JSON perevals/report.py - Leverage built-in metrics from
evals/metrics.py(BLEU, ROUGE, accuracy) and parallel execution via--workersfor scalable benchmarking
Frequently Asked Questions
What is OpenAI Evals used for?
OpenAI Evals is a framework for systematically evaluating language models, embeddings, or any AI system accessible via API. It standardizes the process of prompting models, scoring outputs, and generating structured reports, making it suitable for regression testing, benchmarking, and research.
How do I define a custom metric in OpenAI Evals?
You can define a custom metric by creating a callable function (like the metric static method in the examples above) or by subclassing evals.metrics.BaseMetric in evals/metrics.py. Pass your metric function to the score parameter when recording results, or instantiate built-in metrics like BLEU() directly in your eval class.
Can I run evaluations without using the OpenAI API?
Yes. While the default model wrapper in evals/models/openai.py targets OpenAI's API, the Eval base class accepts any object with a generate() method. You can implement custom model adapters for other providers (Anthropic, local models, etc.) and pass them to your eval's constructor when running programmatically.
How do I parallelize evaluation runs?
Set the --workers N flag when using the CLI, or pass workers=N to the run_eval() function in Python. The framework in evals/run.py manages process pools to distribute dataset items across multiple workers, significantly speeding up evaluation for large datasets while maintaining thread-safe result collection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →