# How to Evaluate an LLM's Performance on the GSM8K Benchmark: A Complete Implementation Guide

> Learn how to evaluate an LLM's performance on the GSM8K benchmark with this complete implementation guide. Discover the steps to accurately measure your model's math problem-solving capabilities.

- Repository: [Fareed Khan/train-llm-from-scratch](https://github.com/FareedKhan-dev/train-llm-from-scratch)
- Tags: how-to-guide
- Published: 2026-06-11

---

**You evaluate an LLM on GSM8K by loading a checkpoint into a Transformer model, running greedy decoding on the dataset's math questions, extracting numeric answers from the "#### N" format, and calculating accuracy through tolerant numeric comparison against gold labels.**

The GSM8K benchmark is the industry standard for measuring arithmetic reasoning capabilities in large language models. This guide explains the exact implementation used in the `FareedKhan-dev/train-llm-from-scratch` repository to evaluate an LLM's performance on the GSM8K benchmark, covering checkpoint reconstruction, greedy inference, and answer verification.

## Loading the Model Checkpoint

The evaluation pipeline begins by reconstructing the model architecture from a saved checkpoint. In [`scripts/eval_post_training.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/eval_post_training.py), the `model_from_ckpt` function (lines 28-42) loads the `.pt` file and preserves only the backbone language model weights.

This design allows the same evaluation code to work for both pure language model checkpoints and reward model checkpoints, as the function automatically strips any reward-specific heads while keeping the transformer backbone intact.

```python
from src.post_training.inference import load_model_from_ckpt

# Load checkpoint - works for both LM and reward model checkpoints

model = load_model_from_ckpt("/ephemeral/ckpts/sft.pt", device="cpu")

```

## Preparing the GSM8K Dataset

The `load_gsm8k_eval` function in [`src/post_training/evaluation.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/evaluation.py) (lines 107-116) pulls the official OpenAI GSM8K dataset from the Hugging Face datasets hub. It returns a list of question-answer pairs and supports an optional `limit` parameter for rapid testing on subsets of the data.

```python
from src.post_training.evaluation import load_gsm8k_eval

# Load test split with optional limit for fast iteration

qa_pairs = load_gsm8k_eval(split="test", limit=200)

```

## Generating Responses with Greedy Decoding

For each question, the pipeline constructs a prompt using the chat template via `encode_prompt`, then generates responses using `batched_generate` in [`src/post_training/evaluation.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/evaluation.py) (lines 23-44). Setting `greedy=True` forces argmax decoding (temperature 0) to ensure deterministic, reproducible results.

The batching mechanism respects the model's context length while processing multiple prompts in parallel, maximizing GPU utilization during evaluation.

```python
from src.post_training.evaluation import batched_generate

# Generate with greedy decoding for deterministic evaluation

responses = batched_generate(
    model, 
    prompts, 
    greedy=True,
    max_new_tokens=300,
    device="cpu"
)

```

## Scoring and Computing Accuracy

The scoring layer extracts numeric answers and performs tolerant comparison:

1. **Gold Answer Extraction**: `gsm8k_gold_answer` in [`src/post_training/rewards/parsing.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/rewards/parsing.py) (lines 71-92) parses the final "#### N" annotation from the GSM8K answer field to extract the reference number.

2. **Answer Verification**: `is_correct` in [`src/post_training/rewards/verifiers.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/rewards/verifiers.py) (lines 34-41) implements tolerant numeric comparison, checking for exact matches after rounding and supporting optional format bonuses.

3. **Accuracy Aggregation**: The `gsm8k_accuracy` function in [`src/post_training/evaluation.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/evaluation.py) (lines 76-104) orchestrates the full pipeline, returning the accuracy percentage along with raw counts of correct answers.

```python
from src.post_training.evaluation import gsm8k_accuracy

# Run full evaluation with sample inspection

result = gsm8k_accuracy(
    model,
    qa_pairs,
    device="cpu",
    max_new_tokens=300,
    greedy=True,
    return_samples=5  # Include 5 examples for manual inspection

)

print(f"Accuracy: {result['accuracy']*100:.1f}% ({result['correct']}/{result['n']})")

```

## Running the Evaluation

### Command-Line Interface

The recommended approach uses the evaluation script in [`scripts/eval_post_training.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/eval_post_training.py):

```bash
PYTHONPATH=. python scripts/eval_post_training.py \
    --ckpt /ephemeral/ckpts/sft.pt \
    --label my_model \
    --limit 200

```

Output format:

```

[my_model] GSM8K test accuracy: 71.0%  (142/200)

```

### Programmatic Python API

For custom evaluation workflows, import the evaluation utilities directly:

```python
from src.post_training.evaluation import gsm8k_accuracy, load_gsm8k_eval
from src.post_training.inference import load_model_from_ckpt

model = load_model_from_ckpt("/ephemeral/ckpts/sft.pt", device="cpu")
qa_pairs = load_gsm8k_eval(split="test", limit=100)

result = gsm8k_accuracy(
    model,
    qa_pairs,
    device="cpu",
    max_new_tokens=300,
    greedy=True,
    return_samples=5,
)

print(f"GSM8K accuracy: {result['accuracy']*100:.1f}%")
for samp in result["samples"]:
    print(f"Q: {samp['q']}")
    print(f"Gold: {samp['gold']} | Correct: {samp['correct']}")
    print(f"Model: {samp['response']}\n")

```

### Interactive Streamlit UI

The repository includes a web interface at [`ui/pages/8_Evaluate.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/pages/8_Evaluate.py) (lines 25-33). Launch with `streamlit run ui/app.py`, navigate to the **Evaluate** page, and select your checkpoint from `/ephemeral/ckpts/` to run the evaluation interactively. The UI exposes the same `gsm8k_accuracy` function while providing visual inspection of sample generations.

## Summary

- **Checkpoint Loading**: Use `model_from_ckpt` in [`scripts/eval_post_training.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/eval_post_training.py) to load transformer weights while automatically handling both LM and reward model checkpoints.
- **Dataset Preparation**: `load_gsm8k_eval` fetches the official GSM8K dataset with optional subset limiting for rapid iteration.
- **Greedy Decoding**: The `batched_generate` function with `greedy=True` ensures deterministic, reproducible evaluation by using argmax sampling.
- **Answer Parsing**: `gsm8k_gold_answer` extracts the "#### N" format from GSM8K answers, while `is_correct` performs tolerant numeric comparison.

- **Evaluation Entry Points**: Choose between CLI scripts ([`scripts/eval_post_training.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/eval_post_training.py)), Python API (`gsm8k_accuracy`), or the Streamlit UI ([`ui/pages/8_Evaluate.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/pages/8_Evaluate.py)).

## Frequently Asked Questions

### What is greedy decoding and why is it used for GSM8K evaluation?

**Greedy decoding generates text by always selecting the highest probability token (argmax) at each step**, setting temperature to zero. This eliminates randomness in the evaluation process, ensuring that running the same checkpoint twice produces identical results. The `batched_generate` function in [`src/post_training/evaluation.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/evaluation.py) implements this when `greedy=True`, making scores comparable across different evaluation runs and hardware configurations.

### How does the repository extract numeric answers from GSM8K's formatted output?

**The `gsm8k_gold_answer` function in [`src/post_training/rewards/parsing.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/rewards/parsing.py) (lines 71-92) parses the special "#### N" delimiter** that appears at the end of every GSM8K answer. This extracts the gold numeric value for comparison. The generated text is similarly processed to extract numeric candidates, which `is_correct` then compares using tolerant rounding logic to handle formatting variations like decimal places or thousand separators.

### Can I evaluate reward model checkpoints using the same evaluation code?

**Yes, the `model_from_ckpt` function automatically strips reward-specific heads** and keeps only the language model backbone. According to the implementation in [`scripts/eval_post_training.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/eval_post_training.py) (lines 28-42), this allows the same evaluation pipeline to work seamlessly with both pure language model checkpoints (`.pt` files) and reward model checkpoints, without requiring separate loading logic or architecture modifications.

### What is the difference between the CLI and programmatic evaluation approaches?

**The CLI in [`scripts/eval_post_training.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/eval_post_training.py) provides a quick, table-formatted summary** suitable for standard benchmarking, while the programmatic API (`gsm8k_accuracy`) offers granular control over parameters like `return_samples` for debugging specific failure modes. The CLI handles PYTHONPATH setup and device management automatically, whereas the Python API requires manual import of `load_model_from_ckpt` and `load_gsm8k_eval`, making it ideal for integration into custom training loops or hyperparameter sweeps.