How to Customize the Evaluator Models in CreativeMath's Evaluation Pipeline

To customize the evaluator models in CreativeMath, edit the evaluators list in src/evaluation.py, update the is_api_model classification in src/models/model_loader.py, and optionally configure version strings in config.json.

CreativeMath (junyiye/creativemath) evaluates mathematical solution candidates through a three-stage pipeline assessing correctness, coarse-grained novelty, and fine-grained novelty. Each stage relies on a configurable set of evaluator LLMs defined directly in the source code. Unlike runtime configurations, customizing these evaluators requires modifying specific Python files and JSON entries to register new models or switch between API and local deployments.

Understanding the Evaluator Architecture

The Three Evaluation Stages

The pipeline runs every candidate solution through sequential checks: correctness verification, coarse-grained novelty assessment, and fine-grained novelty assessment. At each stage, the system queries every model listed in the evaluators array defined in src/evaluation.py.

Core Components Overview

Four components control evaluator behavior:

  • Evaluator list: The evaluators variable in src/evaluation.py (line 19) declares which models participate.
  • Model type detection: The ModelWrapper class in src/models/model_loader.py (lines 12-20) uses self.is_api_model to distinguish API calls from local inference.
  • Version identifiers: The model_version section in config.json maps friendly names to provider-specific model IDs.
  • Generation parameters: The model_config section in config.json sets universal hyperparameters like temperature and token limits.

Step-by-Step Customization Guide

Step 1: Modify the Evaluator List in evaluation.py

Locate the evaluators list at line 19 of src/evaluation.py. This hardcoded array determines which models score each solution. Replace the default ["claude-3-opus", "gemini-1.5-pro", "gpt-4"] with your preferred combination.

Step 2: Configure Model Type Detection in model_loader.py

Open src/models/model_loader.py and find the self.is_api_model list between lines 12-20. Add your new model name to this array if it uses a commercial API (e.g., Anthropic, OpenAI, Google). Omit it to trigger local checkpoint loading via load_local_model.

Step 3: Set Model Version Identifiers in config.json

For API models requiring specific version strings, add an entry to the model_version section of config.json. The key must match the name used in evaluation.py, and the value should be the exact provider model identifier (e.g., "claude-3-5-sonnet": "claude-3-5-sonnet-20240620").

Step 4: Adjust Generation Parameters

Modify the model_config section in config.json to change max_new_tokens, temperature, top_p, or do_sample. These settings apply globally to all evaluators, including any you add.

Practical Code Examples

Adding a New API Model (Claude 3.5 Sonnet)

To integrate claude-3-5-sonnet, update three files:


# src/evaluation.py

evaluators = [
    "claude-3-opus",
    "gemini-1.5-pro",
    "gpt-4",
    "claude-3-5-sonnet",  # New evaluator

]

# src/models/model_loader.py

self.is_api_model = [
    "claude-3-opus",
    "claude-3-5-sonnet",  # Recognize as API

    "deepseek-v2",
    "gemini-1.5-pro",
    "gpt-4",
    "gpt-4o",
    "gpt-4o-mini",
]
// config.json
{
  "model_version": {
    "claude-3-5-sonnet": "claude-3-5-sonnet-20240620"
  }
}

Switching to Local Models Only

To use locally hosted checkpoints like Llama or Mixtral:


# src/evaluation.py

evaluators = ["Llama-3-70B", "Mixtral-8x22B"]

# src/models/model_loader.py

self.is_api_model = [
    "claude-3-opus",
    "gemini-1.5-pro",
    "gpt-4",
    # Exclude local model names

]

Local models bypass the model_version lookup and load via load_local_model.

Tuning Generation Hyperparameters

Reduce token budgets or adjust sampling for all evaluators:

// config.json
{
  "model_config": {
    "max_new_tokens": 512,
    "do_sample": true,
    "top_p": 0.9,
    "temperature": 0.7
  }
}

Summary

Frequently Asked Questions

Where is the evaluator list defined in CreativeMath?

The evaluator list is defined as a Python list named evaluators at line 19 of src/evaluation.py. Unlike many configuration-driven tools, CreativeMath does not read this list from config.json at runtime.

Can I mix API and local models in the same evaluation run?

Yes. The ModelWrapper class dynamically routes each evaluator to either generate_api_response or local inference based on the is_api_model check. You can combine commercial APIs like GPT-4 with local checkpoints like Llama-3-70B in the same evaluators array.

Do I need to restart the pipeline after modifying the evaluator models?

Yes. Because the evaluators list and is_api_model classification are loaded when the Python module initializes, you must restart the evaluation script for changes to take effect. The pipeline does not hot-reload these configurations.

How does CreativeMath determine if a model is API-based or local?

The ModelWrapper constructor in src/models/model_loader.py checks if the model name exists in the hardcoded self.is_api_model list. If present, it initializes an API client via load_api_model; otherwise, it calls load_local_model to load weights from disk.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →