What Does the `src/evaluation.py` Script Handle in CreativeMath?

The src/evaluation.py script serves as the core evaluation driver for CreativeMath, orchestrating a three-stage pipeline that uses multiple LLM judges to verify mathematical correctness and assess solution novelty at both coarse and fine-grained levels.

The src/evaluation.py file is the primary entry point for quality assessment within the CreativeMath framework. This script implements a rigorous evaluation protocol that leverages Claude-3-Opus, Gemini-1.5-Pro, and GPT-4 as independent evaluators to judge both the accuracy and creative merit of generated mathematical solutions.

The Three-Stage Evaluation Pipeline

The script processes every generated solution through a sequential three-stage assessment, with each stage building upon the previous results.

Stage 1: Correctness Verification

The first stage determines whether a generated solution is mathematically valid. The script calls load_correctness_evaluation_prompt (lines 89-91) to construct a prompt containing the original problem and candidate solution. It then queries each evaluator LLM via ModelWrapper.generate_response (line 92) and extracts a binary Yes/No decision using extract_yes_no (line 93). The majority vote across all evaluators determines the final correctness stored in sample["correctness"].

Stage 2: Coarse-Grained Novelty Assessment

Only solutions that pass the correctness check proceed to novelty evaluation. This stage uses load_coarse_grained_novelty_evaluation_prompt (lines 34-36) to build a prompt comparing the new solution against all known reference solutions. The script collects Yes/No votes from each evaluator to determine if the solution is significantly different from existing approaches. Results are stored in sample["coarse_grained_novelty"] with a majority final_decision.

Stage 3: Fine-Grained Novelty Assessment

The final stage examines whether a novel solution introduces new reasoning steps or techniques. Running only on samples that passed coarse-grained assessment where the number of known solutions k is less than total solutions n, the script calls load_fine_grained_novelty_evaluation_prompt (lines 88-91). This stage determines if the solution contributes genuinely new mathematical insights beyond surface-level variations.

Workflow and Implementation Details

Command-Line Interface and Configuration

The script accepts a --model_to_evaluate argument (defaulting to "Deepseek-math-7b-rl") to specify which generation model's outputs should be assessed. Configuration management relies on src/config.py, while logging initialization occurs through src/logger.py (lines 44-49), creating timestamped log files in the logs/ directory.

Multi-Model Evaluation Strategy

Rather than relying on a single judge, src/evaluation.py distributes evaluation across three distinct LLMs: claude-3-opus, gemini-1.5-pro, and gpt-4. The ModelWrapper class from src/models/model_loader.py standardizes API calls across these different providers, enabling consistent prompt execution and response handling.

Incremental Result Persistence

To safeguard against interruptions during long-running evaluations, the script implements checkpointing logic that saves intermediate results every save_interval samples (default 20). Progress is written to output/evaluation/{model_name}.json, allowing the script to resume from existing progress rather than restarting from scratch if interrupted.

Metric Computation and Reporting

Upon completion, the script calculates aggregate statistics including correctness_ratio and novelty_ratio (lines 30-46). These metrics quantify the percentage of generated solutions that are mathematically valid and genuinely novel, respectively. The script also computes derived ratios such as novelty-to-correctness rates, logging final statistics before exiting.

Running the Evaluation Script

Execute the evaluation pipeline from the repository root:

python -m src.evaluation --model_to_evaluate Deepseek-math-7b-rl

This command initiates the three-stage assessment for the specified model, logging progress to logs/generation_Deepseek-math-7b-rl_<timestamp>.log.

To access results programmatically:

import json
from pathlib import Path

eval_path = Path("output/evaluation/Deepseek-math-7b-rl.json")
with eval_path.open() as f:
    results = json.load(f)

# Check final correctness decision

print(results[0]["correctness"]["final_decision"])

# Check fine-grained novelty assessment

print(results[0]["fine_grained_novelty"]["final_decision"])

Summary

  • src/evaluation.py drives the complete post-generation quality assessment pipeline for CreativeMath.
  • Three-stage validation: Correctness verification, coarse-grained novelty assessment, and fine-grained novelty assessment.
  • Multi-judge architecture: Employs Claude-3-Opus, Gemini-1.5-Pro, and GPT-4 for robust consensus evaluation.
  • Fault tolerance: Implements incremental saving every 20 samples to prevent data loss during long runs.
  • Quantitative reporting: Computes correctness ratios, novelty ratios, and derived metrics to benchmark model performance.

Frequently Asked Questions

How does the evaluation script determine if a solution is mathematically correct?

The script uses load_correctness_evaluation_prompt (lines 89-91) to construct a specialized prompt containing the problem statement and candidate solution. It queries each evaluator LLM via ModelWrapper.generate_response (line 92) and extracts a binary Yes/No decision using extract_yes_no (line 93). The majority vote across all three evaluators determines the final correctness stored in sample["correctness"].

What distinguishes coarse-grained from fine-grained novelty assessment?

Coarse-grained novelty assessment determines whether a correct solution differs significantly from all known reference solutions using load_coarse_grained_novelty_evaluation_prompt (lines 34-36). Fine-grained novelty assessment examines whether the solution introduces new reasoning steps or mathematical techniques not present in existing solutions, using load_fine_grained_novelty_evaluation_prompt (lines 88-91). The fine-grained stage only executes for solutions that pass coarse-grained assessment where the number of known solutions k is less than the total n.

Can I evaluate multiple generation models in a single command?

No, the script evaluates one model at a time via the --model_to_evaluate parameter (defaulting to "Deepseek-math-7b-rl"). To evaluate multiple models, you must invoke the script separately for each model, specifying the appropriate model name for each run. Results are stored separately in output/evaluation/{model_name}.json to prevent conflicts.

How does the script prevent data loss if the evaluation is interrupted?

The script implements incremental persistence that saves intermediate results every save_interval samples (default 20) to output/evaluation/{model_name}.json. If interrupted, the script detects existing progress upon restart and resumes from the last saved checkpoint rather than reprocessing from the beginning. This mechanism is managed within the main evaluation loop (lines 76-100) and utilizes JSON I/O utilities from src/utils.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →