How CreativeMath Assesses the Novelty of LLM-Generated Math Solutions: A Two-Stage Voting Pipeline
CreativeMath evaluates the novelty of LLM-generated math solutions through a rigorous three-stage pipeline that filters for correctness before applying a two-tiered novelty assessment using multiple judge models and majority voting.
CreativeMath, developed in the junyiye/creativemath repository, implements a systematic approach to determine whether generated mathematical solutions offer genuinely new insights compared to existing reference solutions. The framework treats novelty assessment as a model-driven voting process that occurs only after solutions have passed strict correctness verification.
The Three-Stage Evaluation Pipeline
The assessment architecture operates sequentially: correctness validation followed by coarse-grained then fine-grained novelty detection. This design ensures that only accurate solutions are considered for novelty analysis, preventing false positives from incorrect but superficially unique answers.
Stage 1: Correctness Filtering with Unanimous Consensus
Before any novelty assessment begins, CreativeMath requires unanimous agreement among three high-capacity evaluator models: claude-3-opus, gemini-1.5-pro, and gpt-4. Only solutions receiving a unanimous "YES" verdict proceed to novelty evaluation.
This strict consensus mechanism is implemented in src/evaluation.py at lines 72-105, where the system queries each judge model and validates that all responses confirm correctness before marking a solution as valid for further analysis.
Stage 2: Coarse-Grained Novelty Assessment
For each correct solution, CreativeMath initiates the first novelty tier by comparing the generated solution against the first k reference solutions. The system constructs a detailed evaluation prompt using load_coarse_grained_novelty_evaluation_prompt defined in src/prompts/prompts.py at lines 55-84.
This prompt injects the mathematical problem, the first k reference solutions, and the newly generated solution, asking judge models to determine whether the new approach is novel. The extract_yes_no utility function (from src/utils.py) normalizes model responses to binary "YES" or "NO" verdicts.
After collecting responses from all three evaluator models, CreativeMath applies majority voting to determine the final coarse-grained decision, as implemented in src/evaluation.py at lines 147-158.
Stage 3: Fine-Grained Novelty Verification
If the coarse-grained vote returns "YES" and the generated solution does not appear among all n reference solutions (where k < n), CreativeMath proceeds to fine-grained verification. This second round evaluates the solution against the remaining reference solutions (those beyond the first k).
The prompt for this stage is generated by load_fine_grained_novelty_evaluation_prompt in src/prompts/prompts.py at lines 86-114. Following the same pattern as the coarse stage, the three evaluator models render binary judgments, and majority voting determines the final fine-grained novelty verdict at src/evaluation.py lines 162-174 and 202-214.
Implementation Details and Code Examples
CreativeMath provides direct access to its evaluation pipeline through command-line execution and modular Python imports.
To run the complete evaluation pipeline including novelty assessment for a specific model:
python -m src.evaluation --model_to_evaluate Deepseek-math-7b-rl
You can also reuse the novelty evaluation prompts independently. The following example demonstrates how to construct a coarse-grained novelty prompt:
from src.prompts.prompts import load_coarse_grained_novelty_evaluation_prompt
problem = "Prove that the sum of the first n integers is n(n+1)/2."
reference = [
"Solution 1: Use induction.",
"Solution 2: Pair numbers from opposite ends."
]
new_solution = """Solution 3: Write the series forwards and backwards,
add term‑by‑term to get 2S = n(n+1), so S = n(n+1)/2."""
prompt = load_coarse_grained_novelty_evaluation_prompt(
problem, reference, k=2, new_solution=new_solution
)
# The prompt can now be sent to any LLM evaluator
To extract binary verdicts from model responses, use the extract_yes_no utility:
from src.utils import extract_yes_no
response = "YES, the solution is novel because it uses a pairing argument."
verdict = extract_yes_no(response) # Returns "YES"
Key Components in the CreativeMath Repository
The novelty assessment system is distributed across several specialized modules:
| File | Role in Novelty Assessment |
|---|---|
src/evaluation.py |
Orchestrates the three-stage pipeline (correctness → coarse-grained → fine-grained) and aggregates votes via majority voting. |
src/prompts/prompts.py |
Defines the novelty-evaluation prompts (load_coarse_grained_novelty_evaluation_prompt, load_fine_grained_novelty_evaluation_prompt) with explicit criteria for judges. |
src/models/api_models.py |
Provides ModelWrapper to send prompts to external LLM APIs (Claude, Gemini, GPT-4) and retrieve raw judgments. |
src/utils.py |
Implements extract_yes_no to normalize LLM outputs into binary "YES"/"NO" verdicts for consistent voting. |
Summary
CreativeMath assesses the novelty of LLM-generated math solutions through a rigorous, multi-stage pipeline:
- Correctness prerequisite: Only solutions passing unanimous validation by three high-capacity models (Claude, Gemini, GPT-4) proceed to novelty evaluation.
- Coarse-grained screening: Solutions are compared against the first k reference solutions using structured prompts and majority voting to identify potentially novel approaches.
- Fine-grained verification: Promising candidates undergo secondary evaluation against remaining reference solutions to confirm genuine novelty.
- Binary aggregation: The
extract_yes_noutility standardizes model outputs, enabling consistent majority voting across all assessment stages.
This architecture ensures transparent, reproducible novelty assessment while mitigating individual model biases through ensemble evaluation.
Frequently Asked Questions
What models does CreativeMath use to assess the novelty of LLM-generated math solutions?
CreativeMath employs three high-capacity evaluator models: claude-3-opus, gemini-1.5-pro, and gpt-4. These models serve as judges in both the correctness filtering and novelty assessment stages, with their binary responses aggregated through majority voting to determine final verdicts.
How does the coarse-grained novelty stage differ from the fine-grained stage?
The coarse-grained stage compares generated solutions against only the first k reference solutions to quickly filter obvious duplicates, using the prompt defined in load_coarse_grained_novelty_evaluation_prompt. The fine-grained stage activates only for solutions passing the coarse filter and compares them against the remaining n-k reference solutions (where n is the total reference count) using load_fine_grained_novelty_evaluation_prompt to catch subtle similarities missed in the initial screening.
Why does CreativeMath use a majority voting mechanism instead of a single model?
Majority voting across three distinct high-capacity models (Claude, Gemini, and GPT-4) reduces reliance on any single model's biases, hallucinations, or idiosyncratic interpretations of novelty. This ensemble approach, implemented in src/evaluation.py through the voting logic at lines 147-158 and 202-214, provides more robust and reproducible assessments than single-model evaluation.
Can I run the novelty assessment independently without the correctness check?
While the full pipeline in src/evaluation.py enforces correctness filtering as a prerequisite, you can manually invoke the novelty evaluation components independently by importing the prompt loaders from src/prompts/prompts.py and the extract_yes_no utility from src/utils.py. However, the standard python -m src.evaluation command always executes the complete three-stage pipeline to ensure methodological rigor.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →