# What Is the CreativeMath Benchmark? A Deep Dive into LLM Mathematical Creativity

> Discover the CreativeMath benchmark for LLMs. This tool measures a model's ability to generate novel mathematical solutions, not just correct answers. Explore its impact.

- Repository: [Junyi Ye/creativemath](https://github.com/junyiye/creativemath)
- Tags: deep-dive
- Published: 2026-03-05

---

**The CreativeMath benchmark evaluates how well large language models (LLMs) can generate novel mathematical solutions rather than merely producing correct answers.**

Unlike traditional math evaluations that only verify answer accuracy, the CreativeMath benchmark—hosted in the `junyiye/creativemath` repository—measures an LLM's ability to invent alternative, creative approaches to problems after analyzing existing reference solutions. This dual focus on **correctness** and **creativity** makes it the first comprehensive framework for quantifying mathematical innovation in AI systems.

## Core Objective of the CreativeMath Benchmark

The primary goal of the CreativeMath benchmark is to move beyond binary correctness metrics. According to the repository's README, the benchmark explicitly "assesses the creativity of LLMs in proposing novel solutions to mathematical problems"【https://github.com/junyiye/creativemath/blob/main/README.md#L19-L21】.

This assessment operates through a three-stage evaluation pipeline that checks:
1. **Correctness** — Does the solution actually solve the problem?
2. **Coarse-grained novelty** — Is the approach fundamentally different from reference solutions?
3. **Fine-grained novelty** — Does the solution demonstrate specific creative attributes like different assumptions, intermediate steps, or generalization?

## How the CreativeMath Benchmark Measures Novelty

The benchmark implements a rigorous pipeline that forces models to diverge from standard approaches while maintaining mathematical validity.

### Generation Phase: Prompting for Novel Solutions

In [`src/generation.py`](https://github.com/junyiye/creativemath/blob/main/src/generation.py), the benchmark constructs prompts that explicitly demand creative divergence. Lines 20-44 of this file generate instructions requesting a "novel solution distinct from the given ones" for each mathematical problem【https://github.com/junyiye/creativemath/blob/main/src/generation.py#L20-L44】.

The generation script feeds the model:
- The problem statement
- *k* existing reference solutions (to establish what approaches to avoid)
- Strict novelty criteria requiring different methods, assumptions, or complexity levels

### Evaluation Pipeline: Correctness and Creativity

The [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) file implements the three-stage judging process (lines 72-140)【https://github.com/junyiye/creativemath/blob/main/src/evaluation.py#L72-L140】:

1. **Correctness Verification** — Uses an LLM evaluator to verify mathematical validity
2. **Coarse-Grained Novelty Assessment** — Determines if the solution represents a genuinely different approach from references
3. **Fine-Grained Novelty Scoring** — Evaluates specific creative dimensions including:
   - Different intermediate steps
   - Alternative assumptions
   - Improved generality
   - Complexity variations

### Prompt Engineering for Creativity Criteria

The novelty criteria are formally defined in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py). Lines 11-26 establish generation prompts that embed explicit creativity requirements, while lines 55-81 define evaluation rubrics for judging novelty【https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py#L11-L26】【https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py#L55-L81】.

These prompts require models to consider:
- **Methodological divergence** — Using different mathematical techniques
- **Assumption variation** — Changing premises or constraints
- **Step innovation** — Altering intermediate reasoning paths
- **Generality** — Creating solutions that apply to broader problem classes

## Running the CreativeMath Benchmark: Code Examples

You can execute the benchmark pipeline using the command-line interfaces provided in the repository.

### Generate Novel Solutions

To generate creative solutions for a specific model:

```bash
python src/generation.py --model_name gpt-4o

```

- The `--model_name` parameter accepts model identifiers like `gpt-4o` or `gemini-1.5-pro`
- The script iterates through the dataset, constructs prompts containing reference solutions, and stores outputs in `output/generation/<model_name>.json`

### Evaluate Generated Solutions

To assess correctness and novelty:

```bash
python src/evaluation.py --model_to_evaluate gpt-4o

```

This command runs three separate LLM evaluators (`claude-3-opus`, `gemini-1.5-pro`, `gpt-4`) to verify correctness and score novelty at both coarse and fine-grained levels. Results are saved to `output/evaluation/<model_name>.json`.

### Inspect Prompts Programmatically

To examine the exact prompts used for generation:

```python
from prompts import load_novel_solution_generation_prompt

prompt = load_novel_solution_generation_prompt(
    problem="Find the area of a circle with radius 3.", 
    solutions=["Solution 1: Use πr².", "Solution 2: Approximate with polygons."], 
    k=2
)
print(prompt)

```

This reveals the embedded novelty criteria and instructions that force the model to diverge from the provided reference approaches.

## Key Files in the CreativeMath Repository

Understanding the repository structure helps clarify how the benchmark achieves its creative evaluation goals:

- **[`README.md`](https://github.com/junyiye/creativemath/blob/main/README.md)** — Contains the high-level description and motivation, explicitly stating the benchmark assesses "the creativity of LLMs in proposing novel solutions"【https://github.com/junyiye/creativemath/blob/main/README.md】
- **[`src/generation.py`](https://github.com/junyiye/creativemath/blob/main/src/generation.py)** — Implements the novel solution generation pipeline, constructing prompts that demand distinct approaches from reference solutions【https://github.com/junyiye/creativemath/blob/main/src/generation.py】
- **[`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py)** — Contains the three-stage evaluation logic for correctness, coarse-grained novelty, and fine-grained novelty assessment【https://github.com/junyiye/creativemath/blob/main/src/evaluation.py】
- **[`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py)** — Defines the prompt templates that encode creativity criteria for both generation and evaluation phases【https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py】

## Summary

The CreativeMath benchmark represents a paradigm shift in evaluating large language models, moving beyond binary correctness to quantify mathematical creativity. Key takeaways include:

- **Novelty over correctness**: Unlike traditional benchmarks, CreativeMath requires models to generate alternative, inventive solutions rather than just accurate answers.
- **Three-stage evaluation**: The pipeline in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) assesses correctness, coarse-grained novelty, and fine-grained novelty to provide comprehensive creativity metrics.
- **Prompt engineering**: The benchmark uses sophisticated prompts in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py) that explicitly demand methodological divergence, assumption variation, and step innovation.
- **Practical implementation**: Users can run the full pipeline via [`src/generation.py`](https://github.com/junyiye/creativemath/blob/main/src/generation.py) and [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) to test models like GPT-4o, Claude-3 Opus, and Gemini 1.5 Pro.

## Frequently Asked Questions

### How does CreativeMath differ from standard math benchmarks like GSM8K or MATH?

Standard benchmarks like GSM8K or MATH evaluate models solely on answer correctness—whether the final numerical result matches the expected output. The CreativeMath benchmark instead requires models to produce **novel solution methods** that differ from provided reference approaches. While correctness is still verified, the primary metric is **creative divergence** measured through coarse and fine-grained novelty assessments in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py).

### What specific criteria determine if a solution is "novel" in CreativeMath?

According to the prompt definitions in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py), novelty is assessed across multiple dimensions: **different mathematical methods**, **alternative assumptions or constraints**, **varied intermediate reasoning steps**, **improved generality** (solutions applicable to broader problem classes), and **complexity variations** (simpler or more sophisticated approaches). The evaluation pipeline checks these criteria through both coarse-grained (binary distinctness) and fine-grained (scored attributes) novelty judgments.

### Can I evaluate open-source models with CreativeMath, or is it limited to API-based LLMs?

The CreativeMath benchmark supports any model accessible through a unified interface. The [`src/generation.py`](https://github.com/junyiye/creativemath/blob/main/src/generation.py) script accepts a `--model_name` parameter that can reference API-based models (like `gpt-4o`, `claude-3-opus`, `gemini-1.5-pro`) or local open-source models if the generation backend is configured to handle them. The evaluation phase in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) uses separate judge models, but the target model being evaluated can be any system capable of producing mathematical solutions in response to the structured prompts defined in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py).

### What outputs does the CreativeMath benchmark provide after evaluation?

After running [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py), the benchmark produces a JSON results file containing three key metrics for each problem: **correctness** (whether the solution accurately solves the problem), **coarse-grained novelty** (a binary judgment indicating if the approach is fundamentally distinct from reference solutions), and **fine-grained novelty** (detailed scores across specific creative dimensions like method variation, assumption differences, and step innovation). These outputs enable quantitative comparison of different models' creative mathematical reasoning capabilities.