What Are the Evaluation Metrics Used in CreativeMath? A 3-Stage Pipeline Guide
CreativeMath employs a rigorous three-stage evaluation pipeline that calculates Correctness Ratio, Novelty Ratio, and Novel-Unknown Ratio by aggregating binary YES/NO judgments from Claude-3-Opus, Gemini-1.5-Pro, and GPT-4 using unanimous and majority voting schemes.
The evaluation metrics used in CreativeMath provide a quantitative framework for assessing both the accuracy and innovative quality of LLM-generated mathematical solutions. Implemented primarily in src/evaluation.py, this system uses a cascading filter approach where each stage depends on the previous, ensuring that novelty is only measured for mathematically correct solutions. Understanding these metrics is essential for interpreting benchmark results and optimizing generative mathematical models.
Overview of the Three-Stage Pipeline
CreativeMath evaluates generated solutions through a sequential process that yields both binary judgments and aggregate performance scores. The pipeline progresses from basic correctness validation to increasingly sophisticated measures of novelty, utilizing three distinct evaluator models: claude-3-opus, gemini-1.5-pro, and gpt-4. Each stage applies specific voting rules to combine individual model assessments into final decisions.
Stage 1: Correctness Assessment and the Correctness Ratio
The first evaluation stage determines whether a generated solution is mathematically valid.
Unanimous Consensus Strategy
Each of the three evaluator models receives the candidate solution alongside a correctness prompt stored in src/prompts/prompts.py. Every model returns a binary "YES" or "NO" judgment. The final decision is "YES" only if all three models unanimously return "YES"; any dissenting vote results in a "NO" classification. This conservative approach minimizes false positives in correctness verification.
Correctness Ratio Calculation
The Correctness Ratio (correctness_ratio in the code) represents the proportion of total samples that achieve a final "YES" decision in this stage. This metric serves as the primary filter for subsequent novelty evaluations, as only correct solutions advance to Stage 2. According to the implementation in src/evaluation.py, these ratios are computed and logged at lines 44-49 of the evaluation run.
Stage 2: Coarse-Grained Novelty and the Novelty Ratio
Solutions that pass the correctness filter proceed to evaluate high-level novelty against existing known solutions.
Majority Voting Mechanism
Coarse-Grained Novelty assessment uses majority voting rather than unanimous consensus. Each evaluator compares the correct solution against existing reference solutions and returns "YES" if the approach appears novel at a high level. The final label is "YES" if at least two of the three models agree on novelty. This stage filters out solutions that merely replicate known proof techniques.
Novelty Ratio Metric
The Novelty Ratio (novelty_ratio) calculates the proportion of all evaluated samples that achieve a "YES" decision in this stage. Because this measurement only includes samples that were already deemed correct, it effectively measures what percentage of correct solutions introduce distinguishably different approaches.
Stage 3: Fine-Grained Novelty and the Novel-Unknown Ratio
The final stage conducts detailed analysis to identify genuinely innovative mathematical ideas.
Conditional Evaluation Logic
Fine-Grained Novelty evaluation occurs only when two conditions are met:
- The solution passed coarse-grained novelty (received a "YES")
- The solution's k value (number of known solutions used) is less than n (total known solutions available)
Using majority voting among the three evaluator models, this stage determines whether a novel solution introduces non-trivial, detailed differences rather than superficial variations.
Novel-Unknown Ratio Calculation
The Novel-Unknown Ratio (novel_unknown_ratio) represents the proportion of total samples that satisfy this stringent criterion. This metric identifies solutions that are not only different from known approaches but also utilize incomplete information from the reference set, indicating stronger generative capability.
Derived Performance Ratios
Beyond the three primary metrics, CreativeMath computes relational ratios that provide deeper insight into model behavior:
- Novelty-to-Correctness Ratio: Calculated as
coarse_grained_novelty_count / correctness_count, this percentage indicates how often correct solutions are also novel at the coarse level. - Novel-Unknown-to-Novelty Ratio: Computed as
fine_grained_novelty_count / coarse_grained_novelty_count, this metric reveals what percentage of coarse-novel solutions contain genuinely new ideas versus superficial variations.
These derived ratios help researchers understand the trade-off between accuracy and innovation in generative mathematical models.
Running the Evaluation Pipeline
Execute the evaluation using the command-line interface to generate these metrics for a specific model:
# Command-line execution
!python src/evaluation.py --model_to_evaluate gpt-4o
After completion, the logger outputs a structured summary:
Correctness Ratio: 68.00%
Novelty Ratio: 42.00%
Novel-Unknown Ratio: 15.00%
Novelty-to-Correctness Ratio: 62.00%
Novel-Unknown-to-Novelty Ratio: 36.00%
Access these results programmatically by loading the saved JSON output:
import json
from pathlib import Path
# Path to the saved evaluation JSON (default location)
eval_path = Path("output/evaluation/gpt-4o.json")
results = json.load(eval_path.open())
# Compute the same ratios manually if needed
N = len(results)
correct = sum(r["correctness"]["final_decision"] == "YES" for r in results)
novel = sum(r["coarse_grained_novelty"]["final_decision"] == "YES" for r in results)
unknown = sum(r["fine_grained_novelty"]["final_decision"] == "YES" for r in results)
print(f"Correctness: {correct/N:.2%}")
print(f"Novelty: {novel/N:.2%}")
print(f"Novel-Unknown: {unknown/N:.2%}")
Key Implementation Files
The evaluation metrics are implemented across four critical components:
src/evaluation.py: Contains the three-stage pipeline implementation, aggregation logic, and ratio calculations (see lines 44-49 for metric logging).src/prompts/prompts.py: Houses the text prompts that drive the YES/NO responses for correctness and novelty assessments.src/utils.py: Provides helper functions includingextract_yes_nofor parsing model outputs, plusload_jsonandsave_jsonfor data management.config.json: Defines experiment parameters includingsave_intervaland directory paths for data, generation, and evaluation outputs.
Summary
- Correctness Ratio requires unanimous agreement from all three evaluator models (Claude-3-Opus, Gemini-1.5-Pro, GPT-4) to classify a solution as mathematically valid.
- Novelty Ratio applies majority voting to correct solutions, measuring high-level distinction from known approaches.
- Novel-Unknown Ratio uses conditional majority voting to identify solutions with genuine mathematical innovation when k < n.
- Derived ratios quantify the relationship between correctness and various novelty grades, providing normalized performance indicators.
- All metrics are computed in
src/evaluation.pyand stored in JSON format within theoutput/evaluation/directory.
Frequently Asked Questions
What evaluator models does CreativeMath use?
CreativeMath utilizes three state-of-the-art language models as evaluators: claude-3-opus, gemini-1.5-pro, and gpt-4. These models independently assess each solution, and their binary responses are aggregated through either unanimous or majority voting depending on the evaluation stage.
How does CreativeMath determine if a solution is mathematically correct?
The system applies a unanimous consensus rule for correctness determination. Each of the three evaluator models generates a YES/NO judgment based on the correctness prompt defined in src/prompts/prompts.py. The solution receives a final "YES" only if all three models return "YES"; otherwise, it is classified as incorrect. This conservative approach ensures high precision in mathematical validity assessment.
What is the difference between coarse-grained and fine-grained novelty?
Coarse-grained novelty evaluates whether a correct solution differs at a high level from existing known solutions, using majority voting among the three evaluators. Fine-grained novelty conducts a more rigorous analysis to detect genuinely new mathematical ideas rather than superficial variations, but only activates when coarse-grained novelty is positive and the solution uses fewer than the total available known solutions (k < n).
Where are the evaluation metrics logged in the codebase?
All aggregate metrics are computed and logged at lines 44-49 of src/evaluation.py. The system outputs both console logs during execution and persistent JSON files in the output/evaluation/ directory (configured via config.json), which contain detailed per-sample decisions and final ratio calculations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →