How to Evaluate LLM Outputs with the OpenAI Evals Framework
The OpenAI Evals framework provides a modular, extensible Python architecture for evaluating large language model outputs through custom evaluation classes, automated runners, and pluggable metrics.
The OpenAI Evals framework offers a systematic approach to evaluate LLM outputs with reproducible, automated testing pipelines. Listed in the owainlewis/awesome-artificial-intelligence repository at line 82 of README.md as a core resource for AI engineering, this open-source toolkit separates evaluation logic from execution mechanics. Whether you are validating question-answer accuracy or measuring semantic similarity, the framework enables continuous monitoring of model performance across any OpenAI API-compatible endpoint.
Understanding the OpenAI Evals Architecture
The framework separates three distinct concerns to ensure clean, maintainable evaluation code.
Eval Definition
An Eval is a Python class that encodes the task definition, including prompt templates, input data, and scoring logic. To create an evaluation, you subclass Eval from openai.evals and implement the run and evaluate methods. These methods define how inputs are processed, how the LLM is queried, and how responses are scored against ground truth.
Runner and Execution
The Runner handles the operational complexity of executing LLM calls at scale. The run_eval function from openai.evals.runner manages batching, parallelism, and API interaction, allowing you to process entire datasets without manual loop management or rate-limit handling.
Results Storage and Reporting
The Results Store persists scores and metadata for longitudinal analysis. The framework supports CSV files, JSON, SQLite, or cloud storage backends. It records the exact request payload—including model name, temperature, and token limits—alongside the raw completion and computed score, ensuring full reproducibility.
Building a Custom Evaluation Class
To evaluate LLM outputs with OpenAI Evals, define a concrete evaluation class that specifies your task logic. The following example implements exact-match accuracy for a question-answering task using the gpt-4o model:
# example_eval.py – a minimal OpenAI Evals definition
from openai import OpenAI
from openai.evals import Eval, run_eval, Accuracy
class SimpleQAEval(Eval):
"""Evaluate a QA pair using exact-match accuracy."""
def __init__(self):
self.client = OpenAI()
self.metric = Accuracy()
def run(self, prompt: str, expected: str):
# Call the LLM (e.g., gpt-4o)
response = self.client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)
answer = response.choices[0].message.content.strip()
# Compute the metric
score = self.metric.compare(answer, expected)
return {"prompt": prompt, "answer": answer, "expected": expected, "score": score}
# Execute the evaluation on a list of QA items
if __name__ == "__main__":
eval = SimpleQAEval()
dataset = [
{"prompt": "What is the capital of France?", "expected": "Paris"},
{"prompt": "Who wrote \"Pride and Prejudice\"?", "expected": "Jane Austen"},
]
results = run_eval(eval, dataset, output_path="results.csv")
print("Saved results to results.csv")
Implementing Custom Metrics
While the framework provides built-in metrics like Accuracy and BLEU, production evaluations often require domain-specific scoring logic. You can extend the base Metric class to implement bespoke evaluation criteria, such as semantic similarity using sentence embeddings:
# custom_metric.py – defining a bespoke similarity metric
from openai.evals import Metric
import numpy as np
from sentence_transformers import SentenceTransformer, util
class CosineSimilarityMetric(Metric):
def __init__(self, model_name="all-MiniLM-L6-v2"):
self.embedder = SentenceTransformer(model_name)
def compare(self, answer: str, reference: str) -> float:
a_emb = self.embedder.encode(answer, convert_to_tensor=True)
r_emb = self.embedder.encode(reference, convert_to_tensor=True)
return util.cos_sim(a_emb, r_emb).item()
# Use the custom metric in an eval
from example_eval import SimpleQAEval
from custom_metric import CosineSimilarityMetric
class SimilarityQAEval(SimpleQAEval):
def __init__(self):
self.client = OpenAI()
self.metric = CosineSimilarityMetric()
Running Evaluations at Scale
The framework supports reproducible evaluation workflows by capturing complete request metadata. When integrating into CI pipelines, you can regression-test LLM behavior by comparing current scores against historical baselines stored in the Results Store. Because the framework is open-source, you can also swap the underlying model provider or add new data loaders to evaluate LLM outputs against proprietary datasets.
Summary
- The OpenAI Evals framework provides a modular architecture separating Eval definitions, Runners, and Reporting components.
- Create evaluation tasks by subclassing
Evaland implementingrunandevaluatemethods to query models and compute scores. - Use
run_evalfromopenai.evals.runnerto handle batching, parallelism, and API execution. - Implement custom metrics by extending the
Metricclass and defining thecomparemethod for domain-specific scoring. - Evaluations are fully reproducible because the framework captures complete request payloads, raw completions, and computed scores in the Results Store.
Frequently Asked Questions
What is the OpenAI Evals framework?
The OpenAI Evals framework is an open-source Python toolkit for systematically evaluating large language model outputs. As referenced in the owainlewis/awesome-artificial-intelligence repository, it provides a structured way to define evaluation tasks, execute them against OpenAI API endpoints, and score results using built-in or custom metrics.
How do I create a custom metric in OpenAI Evals?
Create a custom metric by subclassing Metric from openai.evals and implementing the compare method. This method accepts the model's generated answer and a reference text, returning a numerical score. You then instantiate this metric class within your Eval's __init__ method to evaluate LLM outputs with domain-specific criteria.
Can I use OpenAI Evals with models other than GPT-4?
Yes. The framework works with any model supported by the OpenAI API client. By modifying the model parameter in your API call within the run method, you can evaluate LLM outputs from GPT-3.5, GPT-4o, or other compatible endpoints, making the framework provider-agnostic within the OpenAI ecosystem.
How does OpenAI Evals ensure reproducibility?
The framework ensures reproducibility by recording the complete request payload—including model parameters like temperature, max_tokens, and the model name—alongside the raw completion and computed score. This metadata is stored in the Results Store (CSV, JSON, or database), allowing you to reproduce exact evaluation conditions and track performance changes across model versions or prompt updates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →