# Standard LLM Evaluation Benchmarks: The Complete Guide for AI Engineers

> Discover standard LLM evaluation benchmarks like MMLU and HumanEval. AI engineers gain critical insights into model performance for knowledge, code, and safety.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: deep-dive
- Published: 2026-07-19

---

**Standard LLM evaluation benchmarks like MMLU, HumanEval, and SWE-Bench provide standardized metrics for assessing knowledge recall, code generation, reasoning, and safety in large language models.**

The `rohitg00/ai-engineering-from-scratch` curriculum catalogs the industry-standard benchmarks used to validate large language model capabilities. These benchmarks span academic knowledge, coding proficiency, mathematical reasoning, and safety alignment, forming the foundation of modern LLM evaluation pipelines.

## Knowledge and Reasoning Benchmarks

### MMLU (Massive Multitask Language Understanding)

**MMLU** serves as the primary benchmark for general factual knowledge, testing models across 57 subjects with approximately 15,000 multiple-choice questions. According to the source materials in `/phases/10-llms-from-scratch/10-evaluation/docs/en.md#L503`, this benchmark represents the "bread-and-butter" capability test for foundation models.

The curriculum emphasizes MMLU as a critical contamination-check target during fine-tuning pipelines, as noted in `/phases/19-capstone-projects/07-end-to-end-fine-tuning-pipeline/docs/en.md#L15`, where MinHash algorithms verify that training data does not leak benchmark content.

### GSM8K and MATH

For numerical reasoning assessment, **GSM8K** provides 8,000 grade-school math word problems that test arithmetic and chain-of-thought competence. **MATH** offers competition-style problems (approximately 5,000) targeting deep reasoning and symbolic manipulation. Both are referenced in `/phases/11-llm-engineering/10-evaluation/docs/en.md#L68` and frequently paired with code benchmarks in multi-agent debate contexts (`/phases/14-agent-engineering/25-multi-agent-debate/docs/en.md#L30`).

### ARC and HellaSwag

The **ARC** (AI2 Reasoning Challenge) delivers multiple-choice science questions in easy and challenge sets to measure scientific reasoning under controlled settings. **HellaSwag** tests commonsense inference with roughly 100,000 multiple-choice items designed to probe robustness against distractors. Both benchmarks are integrated into the lm-evaluation-harness suite as documented in `/phases/11-llm-engineering/10-evaluation/docs/en.md#L66`.

## Code Generation Benchmarks

### HumanEval

**HumanEval** consists of 164 Python coding problems with unit-test verification, making it the standard metric for functional code generation quality. The benchmark is detailed in `/phases/10-llms-from-scratch/10-evaluation/docs/en.md#L144` and measures whether models can synthesize syntactically correct, executable code that passes hidden test suites.

### SWE-Bench (Software Engineering Bench)

**SWE-Bench** elevates evaluation to real-world programming tasks, requiring models to resolve actual GitHub issues with unit-test validation. As implemented in `/phases/14-agent-engineering/19-benchmarks-swebench-gaia/code/main.py#L141`, this benchmark assesses end-to-end software development capabilities including debugging, code comprehension, and repository navigation.

## Chat and Instruction-Following Benchmarks

### MT-Bench

**MT-Bench** (Multi-Turn Benchmark) utilizes a "Model-to-Human" pairwise chat format with approximately 80,000 turns to evaluate chat helpfulness and safety in long-form dialogue. The curriculum cites this in `/phases/10-llms-from-scratch/10-evaluation/docs/en.md#L190` as the standard metric for conversational AI systems.

### BIG-Bench

**BIG-Bench** encompasses 204 diverse tasks covering reasoning, common sense, and linguistic capabilities for broad capability profiling. The repository lists this among the 200+ benchmarks supported by lm-evaluation-harness in `/phases/11-llm-engineering/10-evaluation/docs/en.md#L111`.

## Safety and Truthfulness Benchmarks

### TruthfulQA

**TruthfulQA** presents 817 adversarial questions specifically designed to probe factuality and susceptibility to misinformation. This benchmark appears in `/phases/11-llm-engineering/10-evaluation/docs/en.md#L66` as part of the standard harness suite for hallucination detection.

### WMDP (Dual-Use Benchmark)

**WMDP** (Weapons of Mass Destruction Proxy) contains 4,157 multiple-choice questions across biology, cybersecurity, and chemistry domains. Documented in `/phases/18-ethics-safety-alignment/17-wmdp-dual-use-evaluation/docs/en.md#L3`, this benchmark measures dual-use safety risks and supports unlearning research to prevent models from generating harmful knowledge.

## Implementing Standard Benchmarks in Production

### Using lm-evaluation-harness

The `rohitg00/ai-engineering-from-scratch` repository recommends **lm-evaluation-harness** (version 0.4.0) as the canonical tool for running standard LLM evaluation benchmarks. This open-source framework ships with definitions for MMLU, GSM8K, HellaSwag, ARC, TruthfulQA, and 200+ additional tasks.

```python

# Install the official harness

# pip install lm-eval==0.4.0

from lm_eval import evaluator, tasks

# Choose a HuggingFace checkpoint or local model wrapper

model = "EleutherAI/gpt-neo-125M"

# Load multiple standard benchmarks

task_names = ["mmlu", "gsm8k", "hellaswag", "arc_easy", "truthfulqa_mc"]

# Run evaluation on full test sets

results = evaluator.evaluate(
    model=model,
    tasks=task_names,
    limit=0,               # 0 = no limit

    device="cpu",          # use "cuda" for GPU acceleration

)

print(results)   # Returns: {'mmlu': 0.70, 'gsm8k': 0.52, ...}

```

The "Evaluation & Testing LLM Applications" lesson in `/phases/11-llm-engineering/10-evaluation/docs/en.md#L78` provides detailed walkthroughs of this implementation pattern.

### Custom Evaluation with Promptfoo

For CI/CD integration of standard benchmarks, the repository demonstrates **Promptfoo** configuration for mixed evaluation suites:

```yaml

# promptfoo.yaml

prompts:
  - "Answer the following question: {{question}}"
providers:
  - openai:gpt-4o
tests:
  - vars:
      question: "What is the capital of France?"
    assert:
      - type: contains
        value: "Paris"
      - type: llm-rubric
        value: "Factually correct and concise"
  - vars:
      question: "Compute 123 * 456."
    assert:
      - type: llm-rubric
        value: "Accurate numeric answer with correct reasoning"
  - vars:
      question: "Write a Python function to reverse a string."
    assert:
      - type: llm-rubric
        value: "Correct syntax, passes unit test"

```

Executing `promptfoo eval` automates scoring against these rubrics, providing lightweight surrogates for HumanEval-style validation as shown in [`/phases/11-llm-engineering/10-evaluation/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//phases/11-llm-engineering/10-evaluation/docs/en.md).

## Avoiding Data Contamination

Running a diverse suite of standard LLM evaluation benchmarks prevents **Goodhart's Law** effects, where over-optimizing a single metric masks regressions in other capabilities. The curriculum stresses contamination checks—specifically MinHash deduplication against MMLU-Pro and MT-Bench-v2—to ensure training data does not leak benchmark content, as implemented in `/phases/19-capstone-projects/07-end-to-end-fine-tuning-pipeline/docs/en.md#L15`.

## Summary

- **MMLU** provides the standard 57-subject knowledge test for general capability assessment, referenced in `/phases/10-llms-from-scratch/10-evaluation/docs/en.md#L503`.
- **HumanEval** and **SWE-Bench** measure code generation through 164 Python problems and real-world software engineering tasks respectively.
- **GSM8K**, **MATH**, **ARC**, and **HellaSwag** target specific reasoning domains including arithmetic, competition mathematics, and commonsense inference.
- **TruthfulQA** and **WMDP** evaluate safety-critical capabilities including hallucination resistance and dual-use knowledge.
- **lm-evaluation-harness** (version 0.4.0) serves as the primary automation framework, supporting 200+ benchmarks including all standard academic suites.
- Contamination checks using MinHash against benchmarks like MMLU-Pro are mandatory for valid evaluation results.

## Frequently Asked Questions

### What is the most important benchmark for general LLM capability?

**MMLU (Massive Multitask Language Understanding)** remains the industry standard for assessing general knowledge, covering 57 subjects with approximately 15,000 multiple-choice questions. According to the `ai-engineering-from-scratch` curriculum in `/phases/10-llms-from-scratch/10-evaluation/docs/en.md#L503`, MMLU serves as the primary "bread-and-butter" capability test, though it should always be paired with specialized benchmarks like HumanEval or GSM8K to avoid overfitting to a single metric.

### How do I run multiple standard benchmarks automatically?

Use the **lm-evaluation-harness** library (version 0.4.0), which provides standardized implementations for over 200 benchmarks including MMLU, GSM8K, and TruthfulQA. As documented in `/phases/11-llm-engineering/10-evaluation/docs/en.md#L111`, you can batch-evaluate models by passing a list of task names to the `evaluator.evaluate()` function, which handles prompt formatting, generation, and metric calculation consistently across all supported benchmarks.

### What benchmark should I use for coding ability assessment?

**HumanEval** provides the standard 164-problem Python test for basic code generation, while **SWE-Bench** offers more rigorous evaluation on real-world GitHub issues. The curriculum in `/phases/14-agent-engineering/19-benchmarks-swebench-gaia/code/main.py#L141` highlights SWE-Bench as the preferred metric for end-to-end software engineering capabilities, as it requires models to navigate actual repositories and resolve bugs using unit-test verification.

### How do I prevent benchmark data contamination in my training set?

Implement **MinHash deduplication** against standard benchmarks like MMLU-Pro and MT-Bench-v2 before training. The fine-tuning pipeline documentation in `/phases/19-capstone-projects/07-end-to-end-fine-tuning-pipeline/docs/en.md#L15` specifically recommends this approach to ensure that evaluation scores reflect genuine model capability rather than memorization of test questions.