# How to Evaluate Generated Solutions Using CreativeMath: A Three-Stage Pipeline

> Learn how to evaluate generated solutions with CreativeMath. Our three-stage pipeline uses Claude-3-Opus, Gemini-1.5-Pro, and GPT-4 for robust correctness and novelty assessment.

- Repository: [Junyi Ye/creativemath](https://github.com/junyiye/creativemath)
- Tags: tutorial
- Published: 2026-03-05

---

**CreativeMath evaluates generated mathematical solutions through a three-stage pipeline—correctness, coarse-grained novelty, and fine-grained novelty—using majority voting across Claude-3-Opus, Gemini-1.5-Pro, and GPT-4.**

The `junyiye/creativemath` repository provides a systematic framework to evaluate generated solutions using CreativeMath's multi-stage assessment pipeline. Whether you are benchmarking large language models or validating new mathematical proofs, understanding how to evaluate generated solutions using CreativeMath ensures reproducible quality metrics across correctness and novelty dimensions.

## Overview of the CreativeMath Evaluation Pipeline

CreativeMath implements a rigorous three-stage evaluation protocol to assess generated mathematical solutions. The pipeline progresses sequentially from basic verification to nuanced novelty detection, ensuring that only high-quality, genuinely innovative solutions receive positive scores.

The three stages are:

- **Correctness**: Verifies that the generated solution produces the correct mathematical result when compared against reference solutions.
- **Coarse-Grained Novelty**: Determines whether the solution differs in reasoning approach or methodology from existing references.
- **Fine-Grained Novelty**: Assesses deeper structural differences in assumptions, complexity, or mathematical techniques when multiple references exist.

## Setting Up the Evaluation Environment

Before running evaluations, you must configure the system to handle model APIs and data paths. CreativeMath uses a centralized configuration system and a unified model wrapper to abstract away provider-specific implementation details.

### Configuration and API Setup

The evaluation pipeline begins with [`src/config.py`](https://github.com/junyiye/creativemath/blob/main/src/config.py), which loads a JSON configuration file containing model versions, API keys, experiment parameters such as `save_interval`, and file locations used throughout the pipeline.

Ensure your [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) includes valid API keys for ANTHROPIC_API_KEY, GEMINI_API_KEY, and OPENAI_API_KEY to access the evaluator models.

### Model Loading and Abstraction

The `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py) provides a unified interface for both API-based and locally-served models. It determines whether a model requires API access, loads the appropriate client via `load_api_model` or `load_local_model`, and exposes a `generate_response` method.

This method builds request messages through `load_messages` and forwards them to either `generate_api_response` or `generate_local_response`, ensuring consistent interaction patterns regardless of the backend provider.

## Stage 1: Correctness Evaluation

The first stage verifies mathematical accuracy by comparing generated solutions against reference answers. This stage acts as a gatekeeper—only correct solutions proceed to novelty assessment.

### Prompt Design for Correctness

The `load_correctness_evaluation_prompt` function in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py) constructs a prompt that presents the new solution alongside up to two reference solutions. The prompt asks the evaluator LLM to answer **YES** if the generated result matches the references mathematically, or **NO** if it differs or contains errors.

### Binary Decision Extraction

After receiving the LLM response, [`src/utils.py`](https://github.com/junyiye/creativemath/blob/main/src/utils.py) provides helper functions to extract binary "YES/NO" decisions from the model output. The evaluation engine stores this decision in `sample["correctness"]` for each generated solution.

## Stage 2: Coarse-Grained Novelty Assessment

Once a solution passes correctness verification, CreativeMath evaluates whether it represents a genuinely different approach from existing solutions, not merely a superficial variation.

### Evaluating Reasoning Diversity

The `load_coarse_grained_novelty_evaluation_prompt` function in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py) generates prompts that ask evaluator models to assess whether the new solution differs in reasoning, assumptions, or methodology from reference solutions. This captures high-level strategic differences in problem-solving approaches.

### Majority Voting Mechanism

The evaluation engine in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) sends these prompts to three independent evaluator models—Claude-3-Opus, Gemini-1.5-Pro, and GPT-4. It aggregates their responses using majority voting to determine the final coarse-grained novelty decision. If a solution fails the correctness stage, the system automatically assigns "NO" to this stage without invoking the evaluators.

## Stage 3: Fine-Grained Novelty Assessment

For solutions that demonstrate coarse-grained novelty, CreativeMath performs a deeper analysis to identify subtle structural differences in mathematical technique and complexity.

### Deep Structural Analysis

The `load_fine_grained_novelty_evaluation_prompt` function in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py) crafts prompts that examine fine distinctions in assumptions, computational complexity, and mathematical machinery. This stage distinguishes between solutions that merely reorder steps versus those that introduce genuinely novel mathematical insights.

### Conditional Execution Based on Reference Availability

As implemented in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py), the fine-grained stage only executes when additional reference solutions exist beyond those used in earlier stages (specifically when `k < n`). The system again employs majority voting across the three evaluator models to ensure robust judgments before recording the final decision.

## Running the Full Evaluation Pipeline

Executing the evaluation requires minimal setup once configuration is complete. The pipeline automatically handles the three-stage progression and aggregation of results.

### Command-Line Execution

First, install the required dependencies:

```bash
pip install -r requirements.txt

```

Ensure your API keys are defined in [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) (ANTHROPIC_API_KEY, GEMINI_API_KEY, OPENAI_API_KEY). Then run the evaluation for a specific model:

```bash
python -m src.evaluation --model_to_evaluate Deepseek-math-7b-rl

```

### Output and Metrics

The script performs the following actions:

- Loads the dataset from `config["file_paths"]["dataset"]` and generated solutions from `config["file_paths"]["generation"]/Deepseek-math-7b-rl.json`
- Produces an evaluation file at `config["file_paths"]["evaluation"]/Deepseek-math-7b-rl.json` containing per-sample decisions for correctness, coarse-grained novelty, and fine-grained novelty
- Computes aggregate metrics including **Correctness Ratio** and **Novelty-to-Correctness Ratio**
- Logs results to `config["logging"]["log_dir"]`

## Key Source Files in the CreativeMath Repository

Understanding the codebase structure helps customize the evaluation pipeline for specific research needs. The following files implement the core evaluation logic:

- **[`src/config.py`](https://github.com/junyiye/creativemath/blob/main/src/config.py)** – Loads runtime configuration including model versions, API keys, and file paths from a JSON file.
- **[`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py)** – Implements `ModelWrapper` to abstract API and local models behind a unified `generate_response` interface.
- **[`src/models/api_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/api_models.py)** – Handles authentication and request logic for Claude, Gemini, GPT-4, and other API-based evaluators.
- **[`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py)** – Defines `load_correctness_evaluation_prompt`, `load_coarse_grained_novelty_evaluation_prompt`, and `load_fine_grained_novelty_evaluation_prompt` to generate evaluation prompts.
- **[`src/utils.py`](https://github.com/junyiye/creativemath/blob/main/src/utils.py)** – Provides JSON I/O helpers and binary YES/NO extraction from LLM responses.
- **[`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py)** – Main driver that orchestrates the three-stage pipeline, aggregates majority votes, and computes final metrics.

## Summary

CreativeMath provides a rigorous, reproducible framework to evaluate generated solutions using a three-stage pipeline that balances correctness verification with nuanced novelty detection. Key takeaways include:

- **Three-stage filtering**: Solutions must pass correctness, then coarse-grained novelty, then fine-grained novelty assessments to receive high scores.
- **Multi-model adjudication**: Claude-3-Opus, Gemini-1.5-Pro, and GPT-4 serve as independent evaluators with majority voting determining final decisions.
- **Configurable pipeline**: The [`src/config.py`](https://github.com/junyiye/creativemath/blob/main/src/config.py) system and `ModelWrapper` abstraction allow easy swapping of evaluator models and generation sources.
- **Binary extraction**: The [`src/utils.py`](https://github.com/junyiye/creativemath/blob/main/src/utils.py) helpers standardize LLM outputs into actionable YES/NO decisions for automated metric calculation.

## Frequently Asked Questions

### How does CreativeMath determine if a generated solution is correct?

CreativeMath determines correctness by prompting evaluator LLMs (Claude-3-Opus, Gemini-1.5-Pro, and GPT-4) with specially crafted prompts from `load_correctness_evaluation_prompt` in [`src/prompts/prompts.py`](https://github.com/junyiye/creativemath/blob/main/src/prompts/prompts.py). These prompts compare the generated solution against up to two reference solutions, asking the model to respond with **YES** if the results match mathematically, or **NO** if they differ. The system extracts binary decisions using helpers in [`src/utils.py`](https://github.com/junyiye/creativemath/blob/main/src/utils.py) and stores them in `sample["correctness"]`.

### What is the difference between coarse-grained and fine-grained novelty in CreativeMath?

**Coarse-grained novelty** assesses whether a solution differs in high-level reasoning, assumptions, or methodology from reference solutions, capturing strategic diversity in problem-solving approaches. **Fine-grained novelty** examines deeper structural differences in mathematical techniques, computational complexity, and specific assumptions, distinguishing between superficial step reordering versus genuinely novel mathematical insights. The fine-grained stage only executes when additional reference solutions exist (`k < n`) and the solution has already passed the coarse-grained check.

### Can I use different evaluator models than Claude-3-Opus, Gemini-1.5-Pro, and GPT-4?

Yes, the `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py) abstracts the underlying model implementation, allowing you to configure different evaluator models through [`src/config.py`](https://github.com/junyiye/creativemath/blob/main/src/config.py). The system supports both API-based models (via [`src/models/api_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/api_models.py)) and locally-served models, switching between them using `load_api_model` or `load_local_model` based on the configuration. You can modify the JSON configuration file to specify alternative model versions while maintaining the same three-stage evaluation workflow.

### How do I interpret the evaluation output files generated by CreativeMath?

The evaluation script produces a JSON file at `config["file_paths"]["evaluation"]/[model_name].json` containing per-sample decisions for each of the three stages. Each entry includes the binary correctness decision stored in `sample["correctness"]`, the coarse-grained novelty verdict determined by majority voting across the three evaluator models, and the fine-grained novelty decision (when applicable). Additionally, the script prints aggregate metrics including **Correctness Ratio** and **Novelty-to-Correctness Ratio** to the console and saves detailed logs to `config["logging"]["log_dir"]` for further analysis.