# How CreativeMath Assesses the Novelty of LLM-Generated Math Solutions: A Two-Stage Voting Pipeline

> CreativeMath assesses LLM math solution novelty using a two-stage voting pipeline. Discover how this system filters for correctness and identifies original solutions.

- Repository: [Junyi Ye/creativemath](https://github.com/junyiye/creativemath)
- Tags: deep-dive
- Published: 2026-03-05

---

**CreativeMath evaluates the novelty of LLM-generated math solutions through a rigorous three-stage pipeline that filters for correctness before applying a two-tiered novelty assessment using multiple judge models and majority voting.**

CreativeMath, developed in the `junyiye/creativemath` repository, implements a systematic approach to determine whether generated mathematical solutions offer genuinely new insights compared to existing reference solutions. The framework treats novelty assessment as a model-driven voting process that occurs only after solutions have passed strict correctness verification.

## The Three-Stage Evaluation Pipeline

The assessment architecture operates sequentially: correctness validation followed by coarse-grained then fine-grained novelty detection. This design ensures that only accurate solutions are considered for novelty analysis, preventing false positives from incorrect but superficially unique answers.

### Stage 1: Correctness Filtering with Unanimous Consensus

Before any novelty assessment begins, CreativeMath requires unanimous agreement among three high-capacity evaluator models: `claude-3-opus`, `gemini-1.5-pro`, and `gpt-4`. Only solutions receiving a unanimous "YES" verdict proceed to novelty evaluation.

This strict consensus mechanism is implemented in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) at lines 72-105, where the system queries each judge model and validates that all responses confirm correctness before marking a solution as valid for further analysis.

### Stage 2: Coarse-Grained Novelty Assessment

For each correct solution, CreativeMath initiates the first novelty tier by comparing the generated solution against the first *k* reference solutions. The system constructs a detailed evaluation prompt using `load_coarse_grained_novelty_evaluation_prompt` defined in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py) at lines 55-84.

This prompt injects the mathematical problem, the first *k* reference solutions, and the newly generated solution, asking judge models to determine whether the new approach is novel. The `extract_yes_no` utility function (from [`src/utils.py`](https://github.com/junyiye/creativemath/blob/main/src/utils.py)) normalizes model responses to binary "YES" or "NO" verdicts.

After collecting responses from all three evaluator models, CreativeMath applies majority voting to determine the final coarse-grained decision, as implemented in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) at lines 147-158.

### Stage 3: Fine-Grained Novelty Verification

If the coarse-grained vote returns "YES" and the generated solution does not appear among all *n* reference solutions (where *k* < *n*), CreativeMath proceeds to fine-grained verification. This second round evaluates the solution against the remaining reference solutions (those beyond the first *k*).

The prompt for this stage is generated by `load_fine_grained_novelty_evaluation_prompt` in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py) at lines 86-114. Following the same pattern as the coarse stage, the three evaluator models render binary judgments, and majority voting determines the final fine-grained novelty verdict at [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) lines 162-174 and 202-214.

## Implementation Details and Code Examples

CreativeMath provides direct access to its evaluation pipeline through command-line execution and modular Python imports.

To run the complete evaluation pipeline including novelty assessment for a specific model:

```bash
python -m src.evaluation --model_to_evaluate Deepseek-math-7b-rl

```

You can also reuse the novelty evaluation prompts independently. The following example demonstrates how to construct a coarse-grained novelty prompt:

```python
from src.prompts.prompts import load_coarse_grained_novelty_evaluation_prompt

problem = "Prove that the sum of the first n integers is n(n+1)/2."
reference = [
    "Solution 1: Use induction.",
    "Solution 2: Pair numbers from opposite ends."
]
new_solution = """Solution 3: Write the series forwards and backwards,
add term‑by‑term to get 2S = n(n+1), so S = n(n+1)/2."""

prompt = load_coarse_grained_novelty_evaluation_prompt(
    problem, reference, k=2, new_solution=new_solution
)

# The prompt can now be sent to any LLM evaluator

```

To extract binary verdicts from model responses, use the `extract_yes_no` utility:

```python
from src.utils import extract_yes_no

response = "YES, the solution is novel because it uses a pairing argument."
verdict = extract_yes_no(response)   # Returns "YES"

```

## Key Components in the CreativeMath Repository

The novelty assessment system is distributed across several specialized modules:

| File | Role in Novelty Assessment |
|------|---------------------------|
| [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) | Orchestrates the three-stage pipeline (correctness → coarse-grained → fine-grained) and aggregates votes via majority voting. |
| [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py) | Defines the novelty-evaluation prompts (`load_coarse_grained_novelty_evaluation_prompt`, `load_fine_grained_novelty_evaluation_prompt`) with explicit criteria for judges. |
| [`src/models/api_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/api_models.py) | Provides `ModelWrapper` to send prompts to external LLM APIs (Claude, Gemini, GPT-4) and retrieve raw judgments. |
| [`src/utils.py`](https://github.com/junyiye/creativemath/blob/main/src/utils.py) | Implements `extract_yes_no` to normalize LLM outputs into binary "YES"/"NO" verdicts for consistent voting. |

## Summary

CreativeMath assesses the novelty of LLM-generated math solutions through a rigorous, multi-stage pipeline:

- **Correctness prerequisite**: Only solutions passing unanimous validation by three high-capacity models (Claude, Gemini, GPT-4) proceed to novelty evaluation.
- **Coarse-grained screening**: Solutions are compared against the first *k* reference solutions using structured prompts and majority voting to identify potentially novel approaches.
- **Fine-grained verification**: Promising candidates undergo secondary evaluation against remaining reference solutions to confirm genuine novelty.
- **Binary aggregation**: The `extract_yes_no` utility standardizes model outputs, enabling consistent majority voting across all assessment stages.

This architecture ensures transparent, reproducible novelty assessment while mitigating individual model biases through ensemble evaluation.

## Frequently Asked Questions

### What models does CreativeMath use to assess the novelty of LLM-generated math solutions?

CreativeMath employs three high-capacity evaluator models: `claude-3-opus`, `gemini-1.5-pro`, and `gpt-4`. These models serve as judges in both the correctness filtering and novelty assessment stages, with their binary responses aggregated through majority voting to determine final verdicts.

### How does the coarse-grained novelty stage differ from the fine-grained stage?

The coarse-grained stage compares generated solutions against only the first *k* reference solutions to quickly filter obvious duplicates, using the prompt defined in `load_coarse_grained_novelty_evaluation_prompt`. The fine-grained stage activates only for solutions passing the coarse filter and compares them against the remaining *n-k* reference solutions (where *n* is the total reference count) using `load_fine_grained_novelty_evaluation_prompt` to catch subtle similarities missed in the initial screening.

### Why does CreativeMath use a majority voting mechanism instead of a single model?

Majority voting across three distinct high-capacity models (Claude, Gemini, and GPT-4) reduces reliance on any single model's biases, hallucinations, or idiosyncratic interpretations of novelty. This ensemble approach, implemented in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) through the voting logic at lines 147-158 and 202-214, provides more robust and reproducible assessments than single-model evaluation.

### Can I run the novelty assessment independently without the correctness check?

While the full pipeline in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) enforces correctness filtering as a prerequisite, you can manually invoke the novelty evaluation components independently by importing the prompt loaders from [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py) and the `extract_yes_no` utility from [`src/utils.py`](https://github.com/junyiye/creativemath/blob/main/src/utils.py). However, the standard `python -m src.evaluation` command always executes the complete three-stage pipeline to ensure methodological rigor.