How to Use CreativeMath to Evaluate LLM Creativity on Math Problems

CreativeMath provides a three-stage evaluation pipeline that measures how creatively large language models solve mathematical problems by generating novel solutions and assessing them for correctness and multi-level novelty.

The junyiye/creativemath repository offers a fully-featured framework for benchmarking LLM creativity on math problems. Unlike standard accuracy benchmarks, CreativeMath evaluates whether models can produce unique solution approaches that differ from provided reference solutions while maintaining mathematical correctness.

Understanding the CreativeMath Architecture

The framework consists of three tightly-integrated components that work sequentially to evaluate creativity:

  1. Prompt Library (src/prompts/prompts.py) – Synthesizes problem statements, reference solutions, and evaluation criteria into structured natural-language prompts.

  2. Generation Engine (src/generation.py) – Wraps your chosen LLM via models.ModelWrapper and executes the novel-solution generation loop for every problem and reference set combination.

  3. Evaluation Pipeline (src/evaluation.py) – Implements a three-stage assessment protocol using majority voting across three evaluator LLMs (Claude-3-Opus, Gemini-1.5-Pro, and GPT-4).

Setting Up the CreativeMath Environment

Before running evaluations, configure the framework with your API keys and model preferences.

Install the required dependencies:

pip install -r requirements.txt

Update config.json (referenced by src/config.py) with your model identifiers and API endpoints. The configuration file also specifies file paths for the benchmark data (data/subset.json) and output directories.

Generating Novel Solutions with the Generation Engine

The generation phase produces candidate solutions under controlled information exposure. The engine in src/generation.py iterates through every problem in data/subset.json, feeding the LLM a novel-solution generation prompt that includes the problem statement and k reference solutions.

Run the generation script:

python src/generation.py --model_name gpt-4o

Results are persisted as JSON files in output/generation/gpt-4o.json. Each entry contains the generated solution and metadata about the problem and reference set used.

The prompt construction logic resides in src/prompts/prompts.py, specifically within load_novel_solution_generation_prompt, which formats the mathematical problem and reference solutions into the exact string fed to the LLM.

Evaluating LLM Creativity Through the Three-Stage Pipeline

Once solutions are generated, src/evaluation.py executes a rigorous three-stage evaluation protocol to assess creativity. The pipeline uses three distinct evaluation prompts defined in src/prompts/prompts.py:

Stage 1: Correctness Verification

The pipeline first checks whether the generated solution arrives at the same answer as the reference solutions using load_correctness_evaluation_prompt. This filters out mathematically invalid attempts before assessing novelty.

Stage 2: Coarse-Grained Novelty Assessment

Solutions passing the correctness check undergo coarse-grained novelty evaluation via load_coarse_grained_novelty_evaluation_prompt. This stage determines whether the solution is qualitatively different from the first k reference solutions shown during generation.

Stage 3: Fine-Grained Novelty Assessment

Finally, load_fine_grained_novelty_evaluation_prompt assesses whether the solution remains novel when compared against the remaining reference solutions not shown during generation. This distinguishes between superficial variation and genuine creative insight.

Each stage employs majority voting across Claude-3-Opus, Gemini-1.5-Pro, and GPT-4 to ensure robust judgments. Run the evaluation with:

python src/evaluation.py --model_to_evaluate gpt-4o

Results are saved to output/evaluation/gpt-4o.json, containing binary decisions for correctness, coarse novelty, and fine novelty for each generated solution.

Interpreting CreativeMath Evaluation Results

The framework outputs quantitative creativity metrics including correctness ratio (percentage of valid solutions), coarse novelty ratio (diversity from shown examples), and fine novelty ratio (diversity from all examples). These metrics enable direct comparison of creativity across different LLMs or prompting strategies.

Utility functions in src/utils.py handle JSON I/O and extract YES/NO answers from evaluator responses, ensuring consistent result parsing across the pipeline.

Summary

  • CreativeMath evaluates LLM creativity through a three-stage pipeline (correctness, coarse novelty, fine novelty) using majority voting across three evaluator models.
  • The Generation Engine (src/generation.py) produces candidate solutions by exposing models to k reference solutions via prompts defined in src/prompts/prompts.py.
  • The Evaluation Pipeline (src/evaluation.py) applies staged assessments using load_correctness_evaluation_prompt, load_coarse_grained_novelty_evaluation_prompt, and load_fine_grained_novelty_evaluation_prompt.
  • Configuration is managed through config.json and src/config.py, while data/subset.json contains the benchmark problems and reference solutions.

Frequently Asked Questions

What makes CreativeMath different from standard math benchmarks?

Standard benchmarks measure accuracy only, whereas CreativeMath specifically evaluates whether an LLM can produce novel solution approaches that differ from provided references while maintaining correctness. It uses a three-stage evaluation protocol with dual novelty checks (coarse and fine-grained) to distinguish true creativity from superficial variation.

Which LLMs can I evaluate using CreativeMath?

You can evaluate any LLM supported by the ModelWrapper abstraction in src/models/. The framework supports OpenAI models (GPT-4, GPT-4o), Anthropic models (Claude-3-Opus), Google models (Gemini-1.5-Pro), and open-source models like DeepSeek-Math-7B-RL. Specify your chosen model using the --model_name flag in src/generation.py.

How does the three-stage evaluation ensure reliable creativity assessment?

The pipeline first filters for correctness to ensure only valid mathematics proceed to novelty checks. Coarse-grained novelty verifies the solution differs from the k references shown during generation (preventing memorization). Fine-grained novelty checks against the remaining unseen references (ensuring true originality). Majority voting across three independent evaluator LLMs (Claude-3-Opus, Gemini-1.5-Pro, GPT-4) at each stage eliminates individual model bias.

Can I customize the prompt templates in CreativeMath?

Yes. All prompt templates reside in src/prompts/prompts.py. You can modify load_novel_solution_generation_prompt to change how problems and reference solutions are presented to the target LLM, or adjust the three evaluation prompts (load_correctness_evaluation_prompt, load_coarse_grained_novelty_evaluation_prompt, load_fine_grained_novelty_evaluation_prompt) to alter the judging criteria. Ensure you maintain the expected output format (YES/NO responses) so that src/utils.py can parse results correctly.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →