# How to Customize the Evaluator Models in CreativeMath's Evaluation Pipeline

> Learn to customize evaluator models in CreativeMath's evaluation pipeline. Modify evaluators in src/evaluation.py update is_api_model classification and configure version strings for tailored results.

- Repository: [Junyi Ye/creativemath](https://github.com/junyiye/creativemath)
- Tags: how-to-guide
- Published: 2026-03-05

---

**To customize the evaluator models in CreativeMath, edit the `evaluators` list in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py), update the `is_api_model` classification in [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py), and optionally configure version strings in [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json).**

CreativeMath (junyiye/creativemath) evaluates mathematical solution candidates through a three-stage pipeline assessing correctness, coarse-grained novelty, and fine-grained novelty. Each stage relies on a configurable set of evaluator LLMs defined directly in the source code. Unlike runtime configurations, customizing these evaluators requires modifying specific Python files and JSON entries to register new models or switch between API and local deployments.

## Understanding the Evaluator Architecture

### The Three Evaluation Stages

The pipeline runs every candidate solution through sequential checks: correctness verification, coarse-grained novelty assessment, and fine-grained novelty assessment. At each stage, the system queries every model listed in the `evaluators` array defined in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py).

### Core Components Overview

Four components control evaluator behavior:

- **Evaluator list**: The `evaluators` variable in [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) (line 19) declares which models participate.
- **Model type detection**: The `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py) (lines 12-20) uses `self.is_api_model` to distinguish API calls from local inference.
- **Version identifiers**: The `model_version` section in [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) maps friendly names to provider-specific model IDs.
- **Generation parameters**: The `model_config` section in [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) sets universal hyperparameters like temperature and token limits.

## Step-by-Step Customization Guide

### Step 1: Modify the Evaluator List in evaluation.py

Locate the `evaluators` list at line 19 of [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py). This hardcoded array determines which models score each solution. Replace the default `["claude-3-opus", "gemini-1.5-pro", "gpt-4"]` with your preferred combination.

### Step 2: Configure Model Type Detection in model_loader.py

Open [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py) and find the `self.is_api_model` list between lines 12-20. Add your new model name to this array if it uses a commercial API (e.g., Anthropic, OpenAI, Google). Omit it to trigger local checkpoint loading via `load_local_model`.

### Step 3: Set Model Version Identifiers in config.json

For API models requiring specific version strings, add an entry to the `model_version` section of [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json). The key must match the name used in [`evaluation.py`](https://github.com/junyiye/creativemath/blob/main/evaluation.py), and the value should be the exact provider model identifier (e.g., `"claude-3-5-sonnet": "claude-3-5-sonnet-20240620"`).

### Step 4: Adjust Generation Parameters

Modify the `model_config` section in [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) to change `max_new_tokens`, `temperature`, `top_p`, or `do_sample`. These settings apply globally to all evaluators, including any you add.

## Practical Code Examples

### Adding a New API Model (Claude 3.5 Sonnet)

To integrate `claude-3-5-sonnet`, update three files:

```python

# src/evaluation.py

evaluators = [
    "claude-3-opus",
    "gemini-1.5-pro",
    "gpt-4",
    "claude-3-5-sonnet",  # New evaluator

]

```

```python

# src/models/model_loader.py

self.is_api_model = [
    "claude-3-opus",
    "claude-3-5-sonnet",  # Recognize as API

    "deepseek-v2",
    "gemini-1.5-pro",
    "gpt-4",
    "gpt-4o",
    "gpt-4o-mini",
]

```

```json
// config.json
{
  "model_version": {
    "claude-3-5-sonnet": "claude-3-5-sonnet-20240620"
  }
}

```

### Switching to Local Models Only

To use locally hosted checkpoints like Llama or Mixtral:

```python

# src/evaluation.py

evaluators = ["Llama-3-70B", "Mixtral-8x22B"]

```

```python

# src/models/model_loader.py

self.is_api_model = [
    "claude-3-opus",
    "gemini-1.5-pro",
    "gpt-4",
    # Exclude local model names

]

```

Local models bypass the `model_version` lookup and load via `load_local_model`.

### Tuning Generation Hyperparameters

Reduce token budgets or adjust sampling for all evaluators:

```json
// config.json
{
  "model_config": {
    "max_new_tokens": 512,
    "do_sample": true,
    "top_p": 0.9,
    "temperature": 0.7
  }
}

```

## Summary

- Edit [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py) (line 19) to change which models evaluate solutions.
- Update [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py) (lines 12-20) to classify models as API or local.
- Add version mappings to [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) under `model_version` for API models.
- Adjust universal generation settings in [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) under `model_config`.
- The pipeline applies changes immediately upon the next run of [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py).

## Frequently Asked Questions

### Where is the evaluator list defined in CreativeMath?

The evaluator list is defined as a Python list named `evaluators` at line 19 of [`src/evaluation.py`](https://github.com/junyiye/creativemath/blob/main/src/evaluation.py). Unlike many configuration-driven tools, CreativeMath does not read this list from [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) at runtime.

### Can I mix API and local models in the same evaluation run?

Yes. The `ModelWrapper` class dynamically routes each evaluator to either `generate_api_response` or local inference based on the `is_api_model` check. You can combine commercial APIs like GPT-4 with local checkpoints like Llama-3-70B in the same `evaluators` array.

### Do I need to restart the pipeline after modifying the evaluator models?

Yes. Because the `evaluators` list and `is_api_model` classification are loaded when the Python module initializes, you must restart the evaluation script for changes to take effect. The pipeline does not hot-reload these configurations.

### How does CreativeMath determine if a model is API-based or local?

The `ModelWrapper` constructor in [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py) checks if the model name exists in the hardcoded `self.is_api_model` list. If present, it initializes an API client via `load_api_model`; otherwise, it calls `load_local_model` to load weights from disk.