What Is the CreativeMath Benchmark? A Deep Dive into LLM Mathematical Creativity

The CreativeMath benchmark evaluates how well large language models (LLMs) can generate novel mathematical solutions rather than merely producing correct answers.

Unlike traditional math evaluations that only verify answer accuracy, the CreativeMath benchmark—hosted in the junyiye/creativemath repository—measures an LLM's ability to invent alternative, creative approaches to problems after analyzing existing reference solutions. This dual focus on correctness and creativity makes it the first comprehensive framework for quantifying mathematical innovation in AI systems.

Core Objective of the CreativeMath Benchmark

The primary goal of the CreativeMath benchmark is to move beyond binary correctness metrics. According to the repository's README, the benchmark explicitly "assesses the creativity of LLMs in proposing novel solutions to mathematical problems"【https://github.com/junyiye/creativemath/blob/main/README.md#L19-L21】.

This assessment operates through a three-stage evaluation pipeline that checks:

  1. Correctness — Does the solution actually solve the problem?
  2. Coarse-grained novelty — Is the approach fundamentally different from reference solutions?
  3. Fine-grained novelty — Does the solution demonstrate specific creative attributes like different assumptions, intermediate steps, or generalization?

How the CreativeMath Benchmark Measures Novelty

The benchmark implements a rigorous pipeline that forces models to diverge from standard approaches while maintaining mathematical validity.

Generation Phase: Prompting for Novel Solutions

In src/generation.py, the benchmark constructs prompts that explicitly demand creative divergence. Lines 20-44 of this file generate instructions requesting a "novel solution distinct from the given ones" for each mathematical problem【https://github.com/junyiye/creativemath/blob/main/src/generation.py#L20-L44】.

The generation script feeds the model:

  • The problem statement
  • k existing reference solutions (to establish what approaches to avoid)
  • Strict novelty criteria requiring different methods, assumptions, or complexity levels

Evaluation Pipeline: Correctness and Creativity

The src/evaluation.py file implements the three-stage judging process (lines 72-140)【https://github.com/junyiye/creativemath/blob/main/src/evaluation.py#L72-L140】:

  1. Correctness Verification — Uses an LLM evaluator to verify mathematical validity
  2. Coarse-Grained Novelty Assessment — Determines if the solution represents a genuinely different approach from references
  3. Fine-Grained Novelty Scoring — Evaluates specific creative dimensions including:
    • Different intermediate steps
    • Alternative assumptions
    • Improved generality
    • Complexity variations

Prompt Engineering for Creativity Criteria

The novelty criteria are formally defined in src/prompts/prompts.py. Lines 11-26 establish generation prompts that embed explicit creativity requirements, while lines 55-81 define evaluation rubrics for judging novelty【https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py#L11-L26】【https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py#L55-L81】.

These prompts require models to consider:

  • Methodological divergence — Using different mathematical techniques
  • Assumption variation — Changing premises or constraints
  • Step innovation — Altering intermediate reasoning paths
  • Generality — Creating solutions that apply to broader problem classes

Running the CreativeMath Benchmark: Code Examples

You can execute the benchmark pipeline using the command-line interfaces provided in the repository.

Generate Novel Solutions

To generate creative solutions for a specific model:

python src/generation.py --model_name gpt-4o
  • The --model_name parameter accepts model identifiers like gpt-4o or gemini-1.5-pro
  • The script iterates through the dataset, constructs prompts containing reference solutions, and stores outputs in output/generation/<model_name>.json

Evaluate Generated Solutions

To assess correctness and novelty:

python src/evaluation.py --model_to_evaluate gpt-4o

This command runs three separate LLM evaluators (claude-3-opus, gemini-1.5-pro, gpt-4) to verify correctness and score novelty at both coarse and fine-grained levels. Results are saved to output/evaluation/<model_name>.json.

Inspect Prompts Programmatically

To examine the exact prompts used for generation:

from prompts import load_novel_solution_generation_prompt

prompt = load_novel_solution_generation_prompt(
    problem="Find the area of a circle with radius 3.", 
    solutions=["Solution 1: Use πr².", "Solution 2: Approximate with polygons."], 
    k=2
)
print(prompt)

This reveals the embedded novelty criteria and instructions that force the model to diverge from the provided reference approaches.

Key Files in the CreativeMath Repository

Understanding the repository structure helps clarify how the benchmark achieves its creative evaluation goals:

Summary

The CreativeMath benchmark represents a paradigm shift in evaluating large language models, moving beyond binary correctness to quantify mathematical creativity. Key takeaways include:

  • Novelty over correctness: Unlike traditional benchmarks, CreativeMath requires models to generate alternative, inventive solutions rather than just accurate answers.
  • Three-stage evaluation: The pipeline in src/evaluation.py assesses correctness, coarse-grained novelty, and fine-grained novelty to provide comprehensive creativity metrics.
  • Prompt engineering: The benchmark uses sophisticated prompts in src/prompts/prompts.py that explicitly demand methodological divergence, assumption variation, and step innovation.
  • Practical implementation: Users can run the full pipeline via src/generation.py and src/evaluation.py to test models like GPT-4o, Claude-3 Opus, and Gemini 1.5 Pro.

Frequently Asked Questions

How does CreativeMath differ from standard math benchmarks like GSM8K or MATH?

Standard benchmarks like GSM8K or MATH evaluate models solely on answer correctness—whether the final numerical result matches the expected output. The CreativeMath benchmark instead requires models to produce novel solution methods that differ from provided reference approaches. While correctness is still verified, the primary metric is creative divergence measured through coarse and fine-grained novelty assessments in src/evaluation.py.

What specific criteria determine if a solution is "novel" in CreativeMath?

According to the prompt definitions in src/prompts/prompts.py, novelty is assessed across multiple dimensions: different mathematical methods, alternative assumptions or constraints, varied intermediate reasoning steps, improved generality (solutions applicable to broader problem classes), and complexity variations (simpler or more sophisticated approaches). The evaluation pipeline checks these criteria through both coarse-grained (binary distinctness) and fine-grained (scored attributes) novelty judgments.

Can I evaluate open-source models with CreativeMath, or is it limited to API-based LLMs?

The CreativeMath benchmark supports any model accessible through a unified interface. The src/generation.py script accepts a --model_name parameter that can reference API-based models (like gpt-4o, claude-3-opus, gemini-1.5-pro) or local open-source models if the generation backend is configured to handle them. The evaluation phase in src/evaluation.py uses separate judge models, but the target model being evaluated can be any system capable of producing mathematical solutions in response to the structured prompts defined in src/prompts/prompts.py.

What outputs does the CreativeMath benchmark provide after evaluation?

After running src/evaluation.py, the benchmark produces a JSON results file containing three key metrics for each problem: correctness (whether the solution accurately solves the problem), coarse-grained novelty (a binary judgment indicating if the approach is fundamentally distinct from reference solutions), and fine-grained novelty (detailed scores across specific creative dimensions like method variation, assumption differences, and step innovation). These outputs enable quantitative comparison of different models' creative mathematical reasoning capabilities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →