How the CreativeMath Evaluation Pipeline Uses Multiple LLM Judges
The CreativeMath evaluation pipeline employs a panel of three LLM judges—Claude-3-Opus, Gemini-1.5-Pro, and GPT-4—to validate mathematical solutions through three sequential stages: correctness verification, coarse-grained novelty assessment, and fine-grained novelty assessment.
The CreativeMath framework, hosted in the junyiye/creativemath repository, automates the assessment of mathematically creative solutions using a consensus-based arbitration system. This CreativeMath evaluation pipeline orchestrates independent evaluations from multiple large language model architectures to ensure both rigorous accuracy and genuine originality in generated mathematical proofs.
Three-Stage Evaluation Architecture
The pipeline defined in src/evaluation.py processes every generated solution through three distinct validation layers. Each stage filters the dataset before passing qualifying samples to the next phase, creating a cascading quality assurance system.
Stage 1: Correctness Verification
The first stage validates mathematical accuracy using load_correctness_evaluation_prompt from src/prompts/prompts.py (lines 29-52). For each sample, the system loads the problem statement, up to two reference solutions, and the newly generated solution. The prompt asks each judge to respond with YES or NO, and the extract_yes_no utility parses these binary decisions.
The evaluation loop starting at line 72 in src/evaluation.py iterates through each judge in the panel, builds the evaluation context, calls model.generate_response, and stores per-judge results under sample["correctness"][model_name].
Stage 2: Coarse-Grained Novelty Assessment
Samples that achieve unanimous correctness proceed to coarse-grained novelty evaluation. This stage uses load_coarse_grained_novelty_evaluation_prompt (lines 55-84) to present the first k reference solutions to the judge panel. The judges determine whether the new solution represents a genuinely novel approach compared to these initial references.
The pipeline implements strict early-exit filtering at lines 119-124 of src/evaluation.py, ensuring only correct solutions enter this novelty assessment phase. Results are stored under sample["coarse_grained_novelty"][model_name].
Stage 3: Fine-Grained Novelty Assessment
The final stage executes only when coarse-grained novelty returns "YES" and additional reference solutions exist (k < n). Using load_fine_grained_novelty_evaluation_prompt (lines 86-115), the system presents the remaining reference solutions (from k+1 to n) to verify that the solution maintains its novelty against the complete corpus of existing approaches.
This two-tier approach filters obvious similarities before investing computational resources in detailed comparisons against the full reference set.
The LLM Judge Panel and Model Wrapping
The evaluation pipeline relies on a fixed panel of three judges defined in the evaluators list on lines 19-20 of src/evaluation.py: claude-3-opus, gemini-1.5-pro, and gpt-4.
Abstracting API and Local Models
The ModelWrapper class in src/models/model_loader.py provides a unified interface for both API-based and locally-hosted LLMs. When instantiating ModelWrapper(model_name), the system determines whether the model requires API access (lines 12-20) and routes to either load_api_model or load_local_model.
The generate_response method standardizes message formatting through load_messages before routing requests to generate_api_response (handled in src/models/api_models.py) or generate_local_response (implemented in src/models/local_models.py).
Prompt Engineering for Consistent Judging
All evaluation prompts reside in src/prompts/prompts.py as static functions returning structured strings. Each prompt includes:
- The original mathematical problem
- Reference solutions (varying by stage)
- The candidate solution
- Explicit instructions for binary (YES/NO) responses
This standardization ensures that claude-3-opus, gemini-1.5-pro, and gpt-4 receive identical evaluation contexts, enabling fair comparison and reliable consensus building across different LLM architectures.
Aggregation Logic and Final Decisions
After collecting individual judgments from the three-judge panel, the pipeline applies stage-specific aggregation rules to determine final sample outcomes.
Majority Voting Mechanisms
The correctness stage requires unanimous agreement—all three judges must return "YES" for a solution to pass (lines 102-106 in src/evaluation.py). This conservative threshold ensures high confidence in mathematical accuracy before proceeding to novelty assessment.
For novelty evaluation (both coarse and fine-grained), the pipeline uses simple majority voting—at least two of three judges must agree that the solution is novel. This balanced approach accommodates subjective interpretation of mathematical creativity while filtering clear non-novel cases.
The final per-sample decisions are stored as final_decision fields within each stage's dictionary, creating a complete audit trail of the consensus process.
Persistence and Metrics Calculation
The evaluation pipeline writes intermediate results to disk based on the save_interval parameter from config["experiment"]["save_interval"], typically every 20 samples (lines 98-100). This prevents data loss during long-running evaluations across thousands of mathematical problems.
Upon completion, the system calculates aggregate statistics including correctness ratios, coarse-grained novelty rates, and fine-grained novelty rates (lines 44-50), providing quantitative insights into model performance on creative mathematical reasoning tasks.
Running the CreativeMath Evaluation Pipeline
Execute the full evaluation workflow using the module entry point:
python -m src.evaluation
To integrate a single LLM judge into custom evaluation logic, instantiate the components directly:
from src.models import ModelWrapper
from src.prompts.prompts import load_correctness_evaluation_prompt
from src.utils import extract_yes_no
# Initialize a specific judge
model = ModelWrapper("gpt-4")
# Prepare evaluation context
prompt = load_correctness_evaluation_prompt(
problem=problem_text,
reference_solutions=existing_solutions,
new_solution=candidate_solution
)
# Execute judgment
response = model.generate_response(prompt)
decision = extract_yes_no(response) # Returns "YES" or "NO"
print(f"Judge decision: {decision}")
Summary
- The CreativeMath evaluation pipeline processes mathematical solutions through three sequential stages: correctness verification, coarse-grained novelty assessment, and fine-grained novelty assessment.
- A fixed panel of three LLM judges—Claude-3-Opus, Gemini-1.5-Pro, and GPT-4—evaluates every sample, with the
ModelWrapperclass providing unified access to both API and local models. - Unanimous voting determines correctness (all three judges must agree), while majority voting decides novelty (at least two judges).
- All prompts are standardized in
src/prompts/prompts.pyto ensure consistent evaluation contexts across different LLM architectures. - Intermediate results save every 20 samples based on
config["experiment"]["save_interval"], with final aggregate statistics calculated upon completion.
Frequently Asked Questions
What LLM judges does CreativeMath use by default?
The evaluation pipeline uses a fixed panel of three judges defined in src/evaluation.py lines 19-20: claude-3-opus, gemini-1.5-pro, and gpt-4. These models provide diverse architectural perspectives—Anthropic's Claude, Google's Gemini, and OpenAI's GPT-4—to ensure robust consensus-based evaluation of mathematical solutions.
How does the evaluation pipeline handle incorrect solutions?
The pipeline implements strict early-exit filtering. During the correctness stage, a solution must receive unanimous "YES" votes from all three judges to proceed. If any judge returns "NO", the sample fails correctness and is excluded from subsequent novelty evaluation stages (coarse-grained and fine-grained), as implemented in the filtering logic at lines 119-124 of src/evaluation.py.
What is the difference between coarse and fine-grained novelty assessment?
Coarse-grained novelty evaluation compares the candidate solution against the first k reference solutions using load_coarse_grained_novelty_evaluation_prompt (lines 55-84). Fine-grained novelty evaluation activates only when coarse-grained returns "YES" and additional references exist (k < n), comparing against the remaining solutions (k+1 to n) via load_fine_grained_novelty_evaluation_prompt (lines 86-115). This two-tier approach filters obvious similarities before detailed comparison against the full reference set.
Can I use local models instead of API-based judges?
Yes. The ModelWrapper class in src/models/model_loader.py abstracts both deployment modes. When instantiating ModelWrapper(model_name), the system detects whether the model requires API access (lines 12-20) and routes to either load_api_model or load_local_model. Local models are handled by src/models/local_models.py, allowing you to substitute the default API judges with locally-hosted alternatives like LLaMA or Mistral while maintaining the same evaluation interface.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →