How to Evaluate Generated Solutions Using CreativeMath: A Three-Stage Pipeline
CreativeMath evaluates generated mathematical solutions through a three-stage pipeline—correctness, coarse-grained novelty, and fine-grained novelty—using majority voting across Claude-3-Opus, Gemini-1.5-Pro, and GPT-4.
The junyiye/creativemath repository provides a systematic framework to evaluate generated solutions using CreativeMath's multi-stage assessment pipeline. Whether you are benchmarking large language models or validating new mathematical proofs, understanding how to evaluate generated solutions using CreativeMath ensures reproducible quality metrics across correctness and novelty dimensions.
Overview of the CreativeMath Evaluation Pipeline
CreativeMath implements a rigorous three-stage evaluation protocol to assess generated mathematical solutions. The pipeline progresses sequentially from basic verification to nuanced novelty detection, ensuring that only high-quality, genuinely innovative solutions receive positive scores.
The three stages are:
- Correctness: Verifies that the generated solution produces the correct mathematical result when compared against reference solutions.
- Coarse-Grained Novelty: Determines whether the solution differs in reasoning approach or methodology from existing references.
- Fine-Grained Novelty: Assesses deeper structural differences in assumptions, complexity, or mathematical techniques when multiple references exist.
Setting Up the Evaluation Environment
Before running evaluations, you must configure the system to handle model APIs and data paths. CreativeMath uses a centralized configuration system and a unified model wrapper to abstract away provider-specific implementation details.
Configuration and API Setup
The evaluation pipeline begins with src/config.py, which loads a JSON configuration file containing model versions, API keys, experiment parameters such as save_interval, and file locations used throughout the pipeline.
Ensure your config.json includes valid API keys for ANTHROPIC_API_KEY, GEMINI_API_KEY, and OPENAI_API_KEY to access the evaluator models.
Model Loading and Abstraction
The ModelWrapper class in src/models/model_loader.py provides a unified interface for both API-based and locally-served models. It determines whether a model requires API access, loads the appropriate client via load_api_model or load_local_model, and exposes a generate_response method.
This method builds request messages through load_messages and forwards them to either generate_api_response or generate_local_response, ensuring consistent interaction patterns regardless of the backend provider.
Stage 1: Correctness Evaluation
The first stage verifies mathematical accuracy by comparing generated solutions against reference answers. This stage acts as a gatekeeper—only correct solutions proceed to novelty assessment.
Prompt Design for Correctness
The load_correctness_evaluation_prompt function in src/prompts/prompts.py constructs a prompt that presents the new solution alongside up to two reference solutions. The prompt asks the evaluator LLM to answer YES if the generated result matches the references mathematically, or NO if it differs or contains errors.
Binary Decision Extraction
After receiving the LLM response, src/utils.py provides helper functions to extract binary "YES/NO" decisions from the model output. The evaluation engine stores this decision in sample["correctness"] for each generated solution.
Stage 2: Coarse-Grained Novelty Assessment
Once a solution passes correctness verification, CreativeMath evaluates whether it represents a genuinely different approach from existing solutions, not merely a superficial variation.
Evaluating Reasoning Diversity
The load_coarse_grained_novelty_evaluation_prompt function in src/prompts/prompts.py generates prompts that ask evaluator models to assess whether the new solution differs in reasoning, assumptions, or methodology from reference solutions. This captures high-level strategic differences in problem-solving approaches.
Majority Voting Mechanism
The evaluation engine in src/evaluation.py sends these prompts to three independent evaluator models—Claude-3-Opus, Gemini-1.5-Pro, and GPT-4. It aggregates their responses using majority voting to determine the final coarse-grained novelty decision. If a solution fails the correctness stage, the system automatically assigns "NO" to this stage without invoking the evaluators.
Stage 3: Fine-Grained Novelty Assessment
For solutions that demonstrate coarse-grained novelty, CreativeMath performs a deeper analysis to identify subtle structural differences in mathematical technique and complexity.
Deep Structural Analysis
The load_fine_grained_novelty_evaluation_prompt function in src/prompts/prompts.py crafts prompts that examine fine distinctions in assumptions, computational complexity, and mathematical machinery. This stage distinguishes between solutions that merely reorder steps versus those that introduce genuinely novel mathematical insights.
Conditional Execution Based on Reference Availability
As implemented in src/evaluation.py, the fine-grained stage only executes when additional reference solutions exist beyond those used in earlier stages (specifically when k < n). The system again employs majority voting across the three evaluator models to ensure robust judgments before recording the final decision.
Running the Full Evaluation Pipeline
Executing the evaluation requires minimal setup once configuration is complete. The pipeline automatically handles the three-stage progression and aggregation of results.
Command-Line Execution
First, install the required dependencies:
pip install -r requirements.txt
Ensure your API keys are defined in config.json (ANTHROPIC_API_KEY, GEMINI_API_KEY, OPENAI_API_KEY). Then run the evaluation for a specific model:
python -m src.evaluation --model_to_evaluate Deepseek-math-7b-rl
Output and Metrics
The script performs the following actions:
- Loads the dataset from
config["file_paths"]["dataset"]and generated solutions fromconfig["file_paths"]["generation"]/Deepseek-math-7b-rl.json - Produces an evaluation file at
config["file_paths"]["evaluation"]/Deepseek-math-7b-rl.jsoncontaining per-sample decisions for correctness, coarse-grained novelty, and fine-grained novelty - Computes aggregate metrics including Correctness Ratio and Novelty-to-Correctness Ratio
- Logs results to
config["logging"]["log_dir"]
Key Source Files in the CreativeMath Repository
Understanding the codebase structure helps customize the evaluation pipeline for specific research needs. The following files implement the core evaluation logic:
src/config.py– Loads runtime configuration including model versions, API keys, and file paths from a JSON file.src/models/model_loader.py– ImplementsModelWrapperto abstract API and local models behind a unifiedgenerate_responseinterface.src/models/api_models.py– Handles authentication and request logic for Claude, Gemini, GPT-4, and other API-based evaluators.src/prompts/prompts.py– Definesload_correctness_evaluation_prompt,load_coarse_grained_novelty_evaluation_prompt, andload_fine_grained_novelty_evaluation_promptto generate evaluation prompts.src/utils.py– Provides JSON I/O helpers and binary YES/NO extraction from LLM responses.src/evaluation.py– Main driver that orchestrates the three-stage pipeline, aggregates majority votes, and computes final metrics.
Summary
CreativeMath provides a rigorous, reproducible framework to evaluate generated solutions using a three-stage pipeline that balances correctness verification with nuanced novelty detection. Key takeaways include:
- Three-stage filtering: Solutions must pass correctness, then coarse-grained novelty, then fine-grained novelty assessments to receive high scores.
- Multi-model adjudication: Claude-3-Opus, Gemini-1.5-Pro, and GPT-4 serve as independent evaluators with majority voting determining final decisions.
- Configurable pipeline: The
src/config.pysystem andModelWrapperabstraction allow easy swapping of evaluator models and generation sources. - Binary extraction: The
src/utils.pyhelpers standardize LLM outputs into actionable YES/NO decisions for automated metric calculation.
Frequently Asked Questions
How does CreativeMath determine if a generated solution is correct?
CreativeMath determines correctness by prompting evaluator LLMs (Claude-3-Opus, Gemini-1.5-Pro, and GPT-4) with specially crafted prompts from load_correctness_evaluation_prompt in src/prompts/prompts.py. These prompts compare the generated solution against up to two reference solutions, asking the model to respond with YES if the results match mathematically, or NO if they differ. The system extracts binary decisions using helpers in src/utils.py and stores them in sample["correctness"].
What is the difference between coarse-grained and fine-grained novelty in CreativeMath?
Coarse-grained novelty assesses whether a solution differs in high-level reasoning, assumptions, or methodology from reference solutions, capturing strategic diversity in problem-solving approaches. Fine-grained novelty examines deeper structural differences in mathematical techniques, computational complexity, and specific assumptions, distinguishing between superficial step reordering versus genuinely novel mathematical insights. The fine-grained stage only executes when additional reference solutions exist (k < n) and the solution has already passed the coarse-grained check.
Can I use different evaluator models than Claude-3-Opus, Gemini-1.5-Pro, and GPT-4?
Yes, the ModelWrapper class in src/models/model_loader.py abstracts the underlying model implementation, allowing you to configure different evaluator models through src/config.py. The system supports both API-based models (via src/models/api_models.py) and locally-served models, switching between them using load_api_model or load_local_model based on the configuration. You can modify the JSON configuration file to specify alternative model versions while maintaining the same three-stage evaluation workflow.
How do I interpret the evaluation output files generated by CreativeMath?
The evaluation script produces a JSON file at config["file_paths"]["evaluation"]/[model_name].json containing per-sample decisions for each of the three stages. Each entry includes the binary correctness decision stored in sample["correctness"], the coarse-grained novelty verdict determined by majority voting across the three evaluator models, and the fine-grained novelty decision (when applicable). Additionally, the script prints aggregate metrics including Correctness Ratio and Novelty-to-Correctness Ratio to the console and saves detailed logs to config["logging"]["log_dir"] for further analysis.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →