# How the CreativeMath Evaluation Pipeline Uses Multiple LLM Judges

> Discover how the CreativeMath evaluation pipeline uses multiple LLM judges like Claude 3 Opus Gemini 1.5 Pro and GPT 4 to verify mathematical solutions and assess novelty across three stages.

- Repository: [Junyi Ye/creativemath](https://github.com/junyiye/creativemath)
- Tags: internals
- Published: 2026-03-05

---

**The CreativeMath evaluation pipeline employs a panel of three LLM judges—Claude-3-Opus, Gemini-1.5-Pro, and GPT-4—to validate mathematical solutions through three sequential stages: correctness verification, coarse-grained novelty assessment, and fine-grained novelty assessment.**

The CreativeMath framework, hosted in the `junyiye/creativemath` repository, automates the assessment of mathematically creative solutions using a consensus-based arbitration system. This **CreativeMath evaluation pipeline** orchestrates independent evaluations from multiple large language model architectures to ensure both rigorous accuracy and genuine originality in generated mathematical proofs.

## Three-Stage Evaluation Architecture

The pipeline defined in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) processes every generated solution through three distinct validation layers. Each stage filters the dataset before passing qualifying samples to the next phase, creating a cascading quality assurance system.

### Stage 1: Correctness Verification

The first stage validates mathematical accuracy using `load_correctness_evaluation_prompt` from [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py) (lines 29-52). For each sample, the system loads the problem statement, up to two reference solutions, and the newly generated solution. The prompt asks each judge to respond with *YES* or *NO*, and the `extract_yes_no` utility parses these binary decisions.

The evaluation loop starting at line 72 in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) iterates through each judge in the panel, builds the evaluation context, calls `model.generate_response`, and stores per-judge results under `sample["correctness"][model_name]`.

### Stage 2: Coarse-Grained Novelty Assessment

Samples that achieve unanimous correctness proceed to coarse-grained novelty evaluation. This stage uses `load_coarse_grained_novelty_evaluation_prompt` (lines 55-84) to present the first *k* reference solutions to the judge panel. The judges determine whether the new solution represents a genuinely novel approach compared to these initial references.

The pipeline implements strict early-exit filtering at lines 119-124 of [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py), ensuring only correct solutions enter this novelty assessment phase. Results are stored under `sample["coarse_grained_novelty"][model_name]`.

### Stage 3: Fine-Grained Novelty Assessment

The final stage executes only when coarse-grained novelty returns "YES" and additional reference solutions exist (*k < n*). Using `load_fine_grained_novelty_evaluation_prompt` (lines 86-115), the system presents the remaining reference solutions (from *k+1* to *n*) to verify that the solution maintains its novelty against the complete corpus of existing approaches.

This two-tier approach filters obvious similarities before investing computational resources in detailed comparisons against the full reference set.

## The LLM Judge Panel and Model Wrapping

The evaluation pipeline relies on a fixed panel of three judges defined in the `evaluators` list on lines 19-20 of [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py): `claude-3-opus`, `gemini-1.5-pro`, and `gpt-4`.

### Abstracting API and Local Models

The `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py) provides a unified interface for both API-based and locally-hosted LLMs. When instantiating `ModelWrapper(model_name)`, the system determines whether the model requires API access (lines 12-20) and routes to either `load_api_model` or `load_local_model`.

The `generate_response` method standardizes message formatting through `load_messages` before routing requests to `generate_api_response` (handled in [`src/models/api_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/api_models.py)) or `generate_local_response` (implemented in [`src/models/local_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/local_models.py)).

### Prompt Engineering for Consistent Judging

All evaluation prompts reside in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py) as static functions returning structured strings. Each prompt includes:
- The original mathematical problem
- Reference solutions (varying by stage)
- The candidate solution
- Explicit instructions for binary (*YES/NO*) responses

This standardization ensures that `claude-3-opus`, `gemini-1.5-pro`, and `gpt-4` receive identical evaluation contexts, enabling fair comparison and reliable consensus building across different LLM architectures.

## Aggregation Logic and Final Decisions

After collecting individual judgments from the three-judge panel, the pipeline applies stage-specific aggregation rules to determine final sample outcomes.

### Majority Voting Mechanisms

The correctness stage requires **unanimous** agreement—all three judges must return "YES" for a solution to pass (lines 102-106 in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py)). This conservative threshold ensures high confidence in mathematical accuracy before proceeding to novelty assessment.

For novelty evaluation (both coarse and fine-grained), the pipeline uses simple **majority voting**—at least two of three judges must agree that the solution is novel. This balanced approach accommodates subjective interpretation of mathematical creativity while filtering clear non-novel cases.

The final per-sample decisions are stored as `final_decision` fields within each stage's dictionary, creating a complete audit trail of the consensus process.

### Persistence and Metrics Calculation

The evaluation pipeline writes intermediate results to disk based on the `save_interval` parameter from `config["experiment"]["save_interval"]`, typically every 20 samples (lines 98-100). This prevents data loss during long-running evaluations across thousands of mathematical problems.

Upon completion, the system calculates aggregate statistics including correctness ratios, coarse-grained novelty rates, and fine-grained novelty rates (lines 44-50), providing quantitative insights into model performance on creative mathematical reasoning tasks.

## Running the CreativeMath Evaluation Pipeline

Execute the full evaluation workflow using the module entry point:

```bash
python -m src.evaluation

```

To integrate a single LLM judge into custom evaluation logic, instantiate the components directly:

```python
from src.models import ModelWrapper
from src.prompts.prompts import load_correctness_evaluation_prompt
from src.utils import extract_yes_no

# Initialize a specific judge

model = ModelWrapper("gpt-4")

# Prepare evaluation context

prompt = load_correctness_evaluation_prompt(
    problem=problem_text,
    reference_solutions=existing_solutions,
    new_solution=candidate_solution
)

# Execute judgment

response = model.generate_response(prompt)
decision = extract_yes_no(response)  # Returns "YES" or "NO"

print(f"Judge decision: {decision}")

```

## Summary

- The **CreativeMath evaluation pipeline** processes mathematical solutions through three sequential stages: correctness verification, coarse-grained novelty assessment, and fine-grained novelty assessment.
- A fixed panel of three LLM judges—**Claude-3-Opus**, **Gemini-1.5-Pro**, and **GPT-4**—evaluates every sample, with the `ModelWrapper` class providing unified access to both API and local models.
- **Unanimous voting** determines correctness (all three judges must agree), while **majority voting** decides novelty (at least two judges).
- All prompts are standardized in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py) to ensure consistent evaluation contexts across different LLM architectures.
- Intermediate results save every 20 samples based on `config["experiment"]["save_interval"]`, with final aggregate statistics calculated upon completion.

## Frequently Asked Questions

### What LLM judges does CreativeMath use by default?

The evaluation pipeline uses a fixed panel of three judges defined in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) lines 19-20: `claude-3-opus`, `gemini-1.5-pro`, and `gpt-4`. These models provide diverse architectural perspectives—Anthropic's Claude, Google's Gemini, and OpenAI's GPT-4—to ensure robust consensus-based evaluation of mathematical solutions.

### How does the evaluation pipeline handle incorrect solutions?

The pipeline implements strict early-exit filtering. During the correctness stage, a solution must receive unanimous "YES" votes from all three judges to proceed. If any judge returns "NO", the sample fails correctness and is excluded from subsequent novelty evaluation stages (coarse-grained and fine-grained), as implemented in the filtering logic at lines 119-124 of [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py).

### What is the difference between coarse and fine-grained novelty assessment?

Coarse-grained novelty evaluation compares the candidate solution against the first *k* reference solutions using `load_coarse_grained_novelty_evaluation_prompt` (lines 55-84). Fine-grained novelty evaluation activates only when coarse-grained returns "YES" and additional references exist (*k < n*), comparing against the remaining solutions (*k+1* to *n*) via `load_fine_grained_novelty_evaluation_prompt` (lines 86-115). This two-tier approach filters obvious similarities before detailed comparison against the full reference set.

### Can I use local models instead of API-based judges?

Yes. The `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py) abstracts both deployment modes. When instantiating `ModelWrapper(model_name)`, the system detects whether the model requires API access (lines 12-20) and routes to either `load_api_model` or `load_local_model`. Local models are handled by [`src/models/local_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/local_models.py), allowing you to substitute the default API judges with locally-hosted alternatives like LLaMA or Mistral while maintaining the same evaluation interface.