# What Are the Evaluation Metrics Used in CreativeMath? A 3-Stage Pipeline Guide

> Discover CreativeMath evaluation metrics. Understand the 3-stage pipeline using Correctness Novelty & Novel-Unknown Ratios with AI judgments. Learn how CreativeMath ensures quality.

- Repository: [Junyi Ye/creativemath](https://github.com/junyiye/creativemath)
- Tags: how-to-guide
- Published: 2026-03-05

---

**CreativeMath employs a rigorous three-stage evaluation pipeline that calculates Correctness Ratio, Novelty Ratio, and Novel-Unknown Ratio by aggregating binary YES/NO judgments from Claude-3-Opus, Gemini-1.5-Pro, and GPT-4 using unanimous and majority voting schemes.**

The evaluation metrics used in CreativeMath provide a quantitative framework for assessing both the accuracy and innovative quality of LLM-generated mathematical solutions. Implemented primarily in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py), this system uses a cascading filter approach where each stage depends on the previous, ensuring that novelty is only measured for mathematically correct solutions. Understanding these metrics is essential for interpreting benchmark results and optimizing generative mathematical models.

## Overview of the Three-Stage Pipeline

CreativeMath evaluates generated solutions through a sequential process that yields both binary judgments and aggregate performance scores. The pipeline progresses from basic correctness validation to increasingly sophisticated measures of novelty, utilizing three distinct evaluator models: **`claude-3-opus`**, **`gemini-1.5-pro`**, and **`gpt-4`**. Each stage applies specific voting rules to combine individual model assessments into final decisions.

## Stage 1: Correctness Assessment and the Correctness Ratio

The first evaluation stage determines whether a generated solution is mathematically valid.

### Unanimous Consensus Strategy

Each of the three evaluator models receives the candidate solution alongside a correctness prompt stored in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py). Every model returns a binary "YES" or "NO" judgment. The **final decision is "YES" only if all three models unanimously return "YES"**; any dissenting vote results in a "NO" classification. This conservative approach minimizes false positives in correctness verification.

### Correctness Ratio Calculation

The **Correctness Ratio** (`correctness_ratio` in the code) represents the proportion of total samples that achieve a final "YES" decision in this stage. This metric serves as the primary filter for subsequent novelty evaluations, as only correct solutions advance to Stage 2. According to the implementation in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py), these ratios are computed and logged at lines 44-49 of the evaluation run.

## Stage 2: Coarse-Grained Novelty and the Novelty Ratio

Solutions that pass the correctness filter proceed to evaluate high-level novelty against existing known solutions.

### Majority Voting Mechanism

**Coarse-Grained Novelty** assessment uses **majority voting** rather than unanimous consensus. Each evaluator compares the correct solution against existing reference solutions and returns "YES" if the approach appears novel at a high level. The final label is "YES" if at least two of the three models agree on novelty. This stage filters out solutions that merely replicate known proof techniques.

### Novelty Ratio Metric

The **Novelty Ratio** (`novelty_ratio`) calculates the proportion of all evaluated samples that achieve a "YES" decision in this stage. Because this measurement only includes samples that were already deemed correct, it effectively measures what percentage of correct solutions introduce distinguishably different approaches.

## Stage 3: Fine-Grained Novelty and the Novel-Unknown Ratio

The final stage conducts detailed analysis to identify genuinely innovative mathematical ideas.

### Conditional Evaluation Logic

**Fine-Grained Novelty** evaluation occurs only when two conditions are met:
1. The solution passed coarse-grained novelty (received a "YES")
2. The solution's *k* value (number of known solutions used) is less than *n* (total known solutions available)

Using majority voting among the three evaluator models, this stage determines whether a novel solution introduces non-trivial, detailed differences rather than superficial variations.

### Novel-Unknown Ratio Calculation

The **Novel-Unknown Ratio** (`novel_unknown_ratio`) represents the proportion of total samples that satisfy this stringent criterion. This metric identifies solutions that are not only different from known approaches but also utilize incomplete information from the reference set, indicating stronger generative capability.

## Derived Performance Ratios

Beyond the three primary metrics, CreativeMath computes relational ratios that provide deeper insight into model behavior:

- **Novelty-to-Correctness Ratio**: Calculated as `coarse_grained_novelty_count / correctness_count`, this percentage indicates how often correct solutions are also novel at the coarse level.
- **Novel-Unknown-to-Novelty Ratio**: Computed as `fine_grained_novelty_count / coarse_grained_novelty_count`, this metric reveals what percentage of coarse-novel solutions contain genuinely new ideas versus superficial variations.

These derived ratios help researchers understand the trade-off between accuracy and innovation in generative mathematical models.

## Running the Evaluation Pipeline

Execute the evaluation using the command-line interface to generate these metrics for a specific model:

```python

# Command-line execution

!python src/evaluation.py --model_to_evaluate gpt-4o

```

After completion, the logger outputs a structured summary:

```

Correctness Ratio: 68.00%
Novelty Ratio: 42.00%
Novel-Unknown Ratio: 15.00%
Novelty-to-Correctness Ratio: 62.00%
Novel-Unknown-to-Novelty Ratio: 36.00%

```

Access these results programmatically by loading the saved JSON output:

```python
import json
from pathlib import Path

# Path to the saved evaluation JSON (default location)

eval_path = Path("output/evaluation/gpt-4o.json")
results = json.load(eval_path.open())

# Compute the same ratios manually if needed

N = len(results)
correct = sum(r["correctness"]["final_decision"] == "YES" for r in results)
novel   = sum(r["coarse_grained_novelty"]["final_decision"] == "YES" for r in results)
unknown = sum(r["fine_grained_novelty"]["final_decision"] == "YES" for r in results)

print(f"Correctness: {correct/N:.2%}")
print(f"Novelty: {novel/N:.2%}")
print(f"Novel-Unknown: {unknown/N:.2%}")

```

## Key Implementation Files

The evaluation metrics are implemented across four critical components:

- **[`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py)**: Contains the three-stage pipeline implementation, aggregation logic, and ratio calculations (see lines 44-49 for metric logging).
- **[`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py)**: Houses the text prompts that drive the YES/NO responses for correctness and novelty assessments.
- **[`src/utils.py`](https://github.com/junyiye/creativemath/blob/main/src/utils.py)**: Provides helper functions including `extract_yes_no` for parsing model outputs, plus `load_json` and `save_json` for data management.
- **[`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json)**: Defines experiment parameters including `save_interval` and directory paths for data, generation, and evaluation outputs.

## Summary

- **Correctness Ratio** requires unanimous agreement from all three evaluator models (Claude-3-Opus, Gemini-1.5-Pro, GPT-4) to classify a solution as mathematically valid.
- **Novelty Ratio** applies majority voting to correct solutions, measuring high-level distinction from known approaches.
- **Novel-Unknown Ratio** uses conditional majority voting to identify solutions with genuine mathematical innovation when *k* < *n*.
- **Derived ratios** quantify the relationship between correctness and various novelty grades, providing normalized performance indicators.
- All metrics are computed in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) and stored in JSON format within the `output/evaluation/` directory.

## Frequently Asked Questions

### What evaluator models does CreativeMath use?

CreativeMath utilizes three state-of-the-art language models as evaluators: **`claude-3-opus`**, **`gemini-1.5-pro`**, and **`gpt-4`**. These models independently assess each solution, and their binary responses are aggregated through either unanimous or majority voting depending on the evaluation stage.

### How does CreativeMath determine if a solution is mathematically correct?

The system applies a **unanimous consensus rule** for correctness determination. Each of the three evaluator models generates a YES/NO judgment based on the correctness prompt defined in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py). The solution receives a final "YES" only if **all three models** return "YES"; otherwise, it is classified as incorrect. This conservative approach ensures high precision in mathematical validity assessment.

### What is the difference between coarse-grained and fine-grained novelty?

**Coarse-grained novelty** evaluates whether a correct solution differs at a high level from existing known solutions, using majority voting among the three evaluators. **Fine-grained novelty** conducts a more rigorous analysis to detect genuinely new mathematical ideas rather than superficial variations, but only activates when coarse-grained novelty is positive and the solution uses fewer than the total available known solutions (*k* < *n*).

### Where are the evaluation metrics logged in the codebase?

All aggregate metrics are computed and logged at **lines 44-49 of [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py)**. The system outputs both console logs during execution and persistent JSON files in the `output/evaluation/` directory (configured via [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json)), which contain detailed per-sample decisions and final ratio calculations.