# How to Use CreativeMath to Evaluate LLM Creativity on Math Problems

> Discover how CreativeMath evaluates LLM creativity on math problems. Learn to generate and assess novel solutions with our three-stage pipeline for correctness and multi-level novelty.

- Repository: [Junyi Ye/creativemath](https://github.com/junyiye/creativemath)
- Tags: how-to-guide
- Published: 2026-03-05

---

**CreativeMath provides a three-stage evaluation pipeline that measures how creatively large language models solve mathematical problems by generating novel solutions and assessing them for correctness and multi-level novelty.**

The `junyiye/creativemath` repository offers a fully-featured framework for benchmarking LLM creativity on math problems. Unlike standard accuracy benchmarks, CreativeMath evaluates whether models can produce *unique* solution approaches that differ from provided reference solutions while maintaining mathematical correctness.

## Understanding the CreativeMath Architecture

The framework consists of three tightly-integrated components that work sequentially to evaluate creativity:

1. **Prompt Library** ([`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py)) – Synthesizes problem statements, reference solutions, and evaluation criteria into structured natural-language prompts.

2. **Generation Engine** ([`src/generation.py`](https://github.com/junyiye/creativemath/blob/main/src/generation.py)) – Wraps your chosen LLM via `models.ModelWrapper` and executes the novel-solution generation loop for every problem and reference set combination.

3. **Evaluation Pipeline** ([`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py)) – Implements a three-stage assessment protocol using majority voting across three evaluator LLMs (Claude-3-Opus, Gemini-1.5-Pro, and GPT-4).

## Setting Up the CreativeMath Environment

Before running evaluations, configure the framework with your API keys and model preferences.

Install the required dependencies:

```bash
pip install -r requirements.txt

```

Update [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) (referenced by [`src/config.py`](https://github.com/junyiye/creativemath/blob/main/src/config.py)) with your model identifiers and API endpoints. The configuration file also specifies file paths for the benchmark data ([`data/subset.json`](https://github.com/junyiye/creativemath/blob/main/data/subset.json)) and output directories.

## Generating Novel Solutions with the Generation Engine

The generation phase produces candidate solutions under controlled information exposure. The engine in [`src/generation.py`](https://github.com/junyiye/creativemath/blob/main/src/generation.py) iterates through every problem in [`data/subset.json`](https://github.com/junyiye/creativemath/blob/main/data/subset.json), feeding the LLM a *novel-solution generation* prompt that includes the problem statement and *k* reference solutions.

Run the generation script:

```bash
python src/generation.py --model_name gpt-4o

```

Results are persisted as JSON files in [`output/generation/gpt-4o.json`](https://github.com/junyiye/creativemath/blob/main/output/generation/gpt-4o.json). Each entry contains the generated solution and metadata about the problem and reference set used.

The prompt construction logic resides in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py), specifically within `load_novel_solution_generation_prompt`, which formats the mathematical problem and reference solutions into the exact string fed to the LLM.

## Evaluating LLM Creativity Through the Three-Stage Pipeline

Once solutions are generated, [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) executes a rigorous three-stage evaluation protocol to assess creativity. The pipeline uses three distinct evaluation prompts defined in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py):

### Stage 1: Correctness Verification

The pipeline first checks whether the generated solution arrives at the same answer as the reference solutions using `load_correctness_evaluation_prompt`. This filters out mathematically invalid attempts before assessing novelty.

### Stage 2: Coarse-Grained Novelty Assessment

Solutions passing the correctness check undergo coarse-grained novelty evaluation via `load_coarse_grained_novelty_evaluation_prompt`. This stage determines whether the solution is qualitatively different from the first *k* reference solutions shown during generation.

### Stage 3: Fine-Grained Novelty Assessment

Finally, `load_fine_grained_novelty_evaluation_prompt` assesses whether the solution remains novel when compared against the *remaining* reference solutions not shown during generation. This distinguishes between superficial variation and genuine creative insight.

Each stage employs majority voting across Claude-3-Opus, Gemini-1.5-Pro, and GPT-4 to ensure robust judgments. Run the evaluation with:

```bash
python src/evaluation.py --model_to_evaluate gpt-4o

```

Results are saved to [`output/evaluation/gpt-4o.json`](https://github.com/junyiye/creativemath/blob/main/output/evaluation/gpt-4o.json), containing binary decisions for correctness, coarse novelty, and fine novelty for each generated solution.

## Interpreting CreativeMath Evaluation Results

The framework outputs quantitative creativity metrics including **correctness ratio** (percentage of valid solutions), **coarse novelty ratio** (diversity from shown examples), and **fine novelty ratio** (diversity from all examples). These metrics enable direct comparison of creativity across different LLMs or prompting strategies.

Utility functions in [`src/utils.py`](https://github.com/junyiye/creativemath/blob/main/src/utils.py) handle JSON I/O and extract YES/NO answers from evaluator responses, ensuring consistent result parsing across the pipeline.

## Summary

- CreativeMath evaluates LLM creativity through a **three-stage pipeline** (correctness, coarse novelty, fine novelty) using majority voting across three evaluator models.
- The **Generation Engine** ([`src/generation.py`](https://github.com/junyiye/creativemath/blob/main/src/generation.py)) produces candidate solutions by exposing models to *k* reference solutions via prompts defined in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py).
- The **Evaluation Pipeline** ([`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py)) applies staged assessments using `load_correctness_evaluation_prompt`, `load_coarse_grained_novelty_evaluation_prompt`, and `load_fine_grained_novelty_evaluation_prompt`.
- Configuration is managed through [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) and [`src/config.py`](https://github.com/junyiye/creativemath/blob/main/src/config.py), while [`data/subset.json`](https://github.com/junyiye/creativemath/blob/main/data/subset.json) contains the benchmark problems and reference solutions.

## Frequently Asked Questions

### What makes CreativeMath different from standard math benchmarks?

Standard benchmarks measure accuracy only, whereas CreativeMath specifically evaluates whether an LLM can produce **novel solution approaches** that differ from provided references while maintaining correctness. It uses a three-stage evaluation protocol with dual novelty checks (coarse and fine-grained) to distinguish true creativity from superficial variation.

### Which LLMs can I evaluate using CreativeMath?

You can evaluate any LLM supported by the `ModelWrapper` abstraction in `src/models/`. The framework supports OpenAI models (GPT-4, GPT-4o), Anthropic models (Claude-3-Opus), Google models (Gemini-1.5-Pro), and open-source models like DeepSeek-Math-7B-RL. Specify your chosen model using the `--model_name` flag in [`src/generation.py`](https://github.com/junyiye/creativemath/blob/main/src/generation.py).

### How does the three-stage evaluation ensure reliable creativity assessment?

The pipeline first filters for **correctness** to ensure only valid mathematics proceed to novelty checks. **Coarse-grained novelty** verifies the solution differs from the *k* references shown during generation (preventing memorization). **Fine-grained novelty** checks against the remaining unseen references (ensuring true originality). Majority voting across three independent evaluator LLMs (Claude-3-Opus, Gemini-1.5-Pro, GPT-4) at each stage eliminates individual model bias.

### Can I customize the prompt templates in CreativeMath?

Yes. All prompt templates reside in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py). You can modify `load_novel_solution_generation_prompt` to change how problems and reference solutions are presented to the target LLM, or adjust the three evaluation prompts (`load_correctness_evaluation_prompt`, `load_coarse_grained_novelty_evaluation_prompt`, `load_fine_grained_novelty_evaluation_prompt`) to alter the judging criteria. Ensure you maintain the expected output format (YES/NO responses) so that [`src/utils.py`](https://github.com/junyiye/creativemath/blob/main/src/utils.py) can parse results correctly.