# What Does the `src/evaluation.py` Script Handle in CreativeMath?

> Discover how the src/evaluation.py script drives CreativeMath evaluations. It uses LLM judges for three stages of mathematical correctness and novelty assessment.

- Repository: [Junyi Ye/creativemath](https://github.com/junyiye/creativemath)
- Tags: how-to-guide
- Published: 2026-03-05

---

**The [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) script serves as the core evaluation driver for CreativeMath, orchestrating a three-stage pipeline that uses multiple LLM judges to verify mathematical correctness and assess solution novelty at both coarse and fine-grained levels.**

The [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) file is the primary entry point for quality assessment within the CreativeMath framework. This script implements a rigorous evaluation protocol that leverages Claude-3-Opus, Gemini-1.5-Pro, and GPT-4 as independent evaluators to judge both the accuracy and creative merit of generated mathematical solutions.

## The Three-Stage Evaluation Pipeline

The script processes every generated solution through a sequential three-stage assessment, with each stage building upon the previous results.

### Stage 1: Correctness Verification

The first stage determines whether a generated solution is mathematically valid. The script calls `load_correctness_evaluation_prompt` (lines 89-91) to construct a prompt containing the original problem and candidate solution. It then queries each evaluator LLM via `ModelWrapper.generate_response` (line 92) and extracts a binary Yes/No decision using `extract_yes_no` (line 93). The majority vote across all evaluators determines the final correctness stored in `sample["correctness"]`.

### Stage 2: Coarse-Grained Novelty Assessment

Only solutions that pass the correctness check proceed to novelty evaluation. This stage uses `load_coarse_grained_novelty_evaluation_prompt` (lines 34-36) to build a prompt comparing the new solution against all known reference solutions. The script collects Yes/No votes from each evaluator to determine if the solution is significantly different from existing approaches. Results are stored in `sample["coarse_grained_novelty"]` with a majority `final_decision`.

### Stage 3: Fine-Grained Novelty Assessment

The final stage examines whether a novel solution introduces new reasoning steps or techniques. Running only on samples that passed coarse-grained assessment where the number of known solutions *k* is less than total solutions *n*, the script calls `load_fine_grained_novelty_evaluation_prompt` (lines 88-91). This stage determines if the solution contributes genuinely new mathematical insights beyond surface-level variations.

## Workflow and Implementation Details

### Command-Line Interface and Configuration

The script accepts a `--model_to_evaluate` argument (defaulting to `"Deepseek-math-7b-rl"`) to specify which generation model's outputs should be assessed. Configuration management relies on [`src/config.py`](https://github.com/junyiye/creativemath/blob/main/src/config.py), while logging initialization occurs through [`src/logger.py`](https://github.com/junyiye/creativemath/blob/main/src/logger.py) (lines 44-49), creating timestamped log files in the `logs/` directory.

### Multi-Model Evaluation Strategy

Rather than relying on a single judge, [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) distributes evaluation across three distinct LLMs: `claude-3-opus`, `gemini-1.5-pro`, and `gpt-4`. The `ModelWrapper` class from [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py) standardizes API calls across these different providers, enabling consistent prompt execution and response handling.

### Incremental Result Persistence

To safeguard against interruptions during long-running evaluations, the script implements checkpointing logic that saves intermediate results every `save_interval` samples (default 20). Progress is written to `output/evaluation/{model_name}.json`, allowing the script to resume from existing progress rather than restarting from scratch if interrupted.

### Metric Computation and Reporting

Upon completion, the script calculates aggregate statistics including `correctness_ratio` and `novelty_ratio` (lines 30-46). These metrics quantify the percentage of generated solutions that are mathematically valid and genuinely novel, respectively. The script also computes derived ratios such as novelty-to-correctness rates, logging final statistics before exiting.

## Running the Evaluation Script

Execute the evaluation pipeline from the repository root:

```bash
python -m src.evaluation --model_to_evaluate Deepseek-math-7b-rl

```

This command initiates the three-stage assessment for the specified model, logging progress to `logs/generation_Deepseek-math-7b-rl_<timestamp>.log`.

To access results programmatically:

```python
import json
from pathlib import Path

eval_path = Path("output/evaluation/Deepseek-math-7b-rl.json")
with eval_path.open() as f:
    results = json.load(f)

# Check final correctness decision

print(results[0]["correctness"]["final_decision"])

# Check fine-grained novelty assessment

print(results[0]["fine_grained_novelty"]["final_decision"])

```

## Summary

- **[`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py)** drives the complete post-generation quality assessment pipeline for CreativeMath.
- **Three-stage validation**: Correctness verification, coarse-grained novelty assessment, and fine-grained novelty assessment.
- **Multi-judge architecture**: Employs Claude-3-Opus, Gemini-1.5-Pro, and GPT-4 for robust consensus evaluation.
- **Fault tolerance**: Implements incremental saving every 20 samples to prevent data loss during long runs.
- **Quantitative reporting**: Computes correctness ratios, novelty ratios, and derived metrics to benchmark model performance.

## Frequently Asked Questions

### How does the evaluation script determine if a solution is mathematically correct?

The script uses `load_correctness_evaluation_prompt` (lines 89-91) to construct a specialized prompt containing the problem statement and candidate solution. It queries each evaluator LLM via `ModelWrapper.generate_response` (line 92) and extracts a binary Yes/No decision using `extract_yes_no` (line 93). The majority vote across all three evaluators determines the final correctness stored in `sample["correctness"]`.

### What distinguishes coarse-grained from fine-grained novelty assessment?

**Coarse-grained novelty assessment** determines whether a correct solution differs significantly from all known reference solutions using `load_coarse_grained_novelty_evaluation_prompt` (lines 34-36). **Fine-grained novelty assessment** examines whether the solution introduces new reasoning steps or mathematical techniques not present in existing solutions, using `load_fine_grained_novelty_evaluation_prompt` (lines 88-91). The fine-grained stage only executes for solutions that pass coarse-grained assessment where the number of known solutions *k* is less than the total *n*.

### Can I evaluate multiple generation models in a single command?

No, the script evaluates one model at a time via the `--model_to_evaluate` parameter (defaulting to `"Deepseek-math-7b-rl"`). To evaluate multiple models, you must invoke the script separately for each model, specifying the appropriate model name for each run. Results are stored separately in `output/evaluation/{model_name}.json` to prevent conflicts.

### How does the script prevent data loss if the evaluation is interrupted?

The script implements **incremental persistence** that saves intermediate results every `save_interval` samples (default 20) to `output/evaluation/{model_name}.json`. If interrupted, the script detects existing progress upon restart and resumes from the last saved checkpoint rather than reprocessing from the beginning. This mechanism is managed within the main evaluation loop (lines 76-100) and utilizes JSON I/O utilities from [`src/utils.py`](https://github.com/junyiye/creativemath/blob/main/src/utils.py).