How to Customize the Evaluator Models in CreativeMath's Evaluation Pipeline
To customize the evaluator models in CreativeMath, edit the evaluators list in src/evaluation.py, update the is_api_model classification in src/models/model_loader.py, and optionally configure version strings in config.json.
CreativeMath (junyiye/creativemath) evaluates mathematical solution candidates through a three-stage pipeline assessing correctness, coarse-grained novelty, and fine-grained novelty. Each stage relies on a configurable set of evaluator LLMs defined directly in the source code. Unlike runtime configurations, customizing these evaluators requires modifying specific Python files and JSON entries to register new models or switch between API and local deployments.
Understanding the Evaluator Architecture
The Three Evaluation Stages
The pipeline runs every candidate solution through sequential checks: correctness verification, coarse-grained novelty assessment, and fine-grained novelty assessment. At each stage, the system queries every model listed in the evaluators array defined in src/evaluation.py.
Core Components Overview
Four components control evaluator behavior:
- Evaluator list: The
evaluatorsvariable insrc/evaluation.py(line 19) declares which models participate. - Model type detection: The
ModelWrapperclass insrc/models/model_loader.py(lines 12-20) usesself.is_api_modelto distinguish API calls from local inference. - Version identifiers: The
model_versionsection inconfig.jsonmaps friendly names to provider-specific model IDs. - Generation parameters: The
model_configsection inconfig.jsonsets universal hyperparameters like temperature and token limits.
Step-by-Step Customization Guide
Step 1: Modify the Evaluator List in evaluation.py
Locate the evaluators list at line 19 of src/evaluation.py. This hardcoded array determines which models score each solution. Replace the default ["claude-3-opus", "gemini-1.5-pro", "gpt-4"] with your preferred combination.
Step 2: Configure Model Type Detection in model_loader.py
Open src/models/model_loader.py and find the self.is_api_model list between lines 12-20. Add your new model name to this array if it uses a commercial API (e.g., Anthropic, OpenAI, Google). Omit it to trigger local checkpoint loading via load_local_model.
Step 3: Set Model Version Identifiers in config.json
For API models requiring specific version strings, add an entry to the model_version section of config.json. The key must match the name used in evaluation.py, and the value should be the exact provider model identifier (e.g., "claude-3-5-sonnet": "claude-3-5-sonnet-20240620").
Step 4: Adjust Generation Parameters
Modify the model_config section in config.json to change max_new_tokens, temperature, top_p, or do_sample. These settings apply globally to all evaluators, including any you add.
Practical Code Examples
Adding a New API Model (Claude 3.5 Sonnet)
To integrate claude-3-5-sonnet, update three files:
# src/evaluation.py
evaluators = [
"claude-3-opus",
"gemini-1.5-pro",
"gpt-4",
"claude-3-5-sonnet", # New evaluator
]
# src/models/model_loader.py
self.is_api_model = [
"claude-3-opus",
"claude-3-5-sonnet", # Recognize as API
"deepseek-v2",
"gemini-1.5-pro",
"gpt-4",
"gpt-4o",
"gpt-4o-mini",
]
// config.json
{
"model_version": {
"claude-3-5-sonnet": "claude-3-5-sonnet-20240620"
}
}
Switching to Local Models Only
To use locally hosted checkpoints like Llama or Mixtral:
# src/evaluation.py
evaluators = ["Llama-3-70B", "Mixtral-8x22B"]
# src/models/model_loader.py
self.is_api_model = [
"claude-3-opus",
"gemini-1.5-pro",
"gpt-4",
# Exclude local model names
]
Local models bypass the model_version lookup and load via load_local_model.
Tuning Generation Hyperparameters
Reduce token budgets or adjust sampling for all evaluators:
// config.json
{
"model_config": {
"max_new_tokens": 512,
"do_sample": true,
"top_p": 0.9,
"temperature": 0.7
}
}
Summary
- Edit
src/evaluation.py(line 19) to change which models evaluate solutions. - Update
src/models/model_loader.py(lines 12-20) to classify models as API or local. - Add version mappings to
config.jsonundermodel_versionfor API models. - Adjust universal generation settings in
config.jsonundermodel_config. - The pipeline applies changes immediately upon the next run of
src/evaluation.py.
Frequently Asked Questions
Where is the evaluator list defined in CreativeMath?
The evaluator list is defined as a Python list named evaluators at line 19 of src/evaluation.py. Unlike many configuration-driven tools, CreativeMath does not read this list from config.json at runtime.
Can I mix API and local models in the same evaluation run?
Yes. The ModelWrapper class dynamically routes each evaluator to either generate_api_response or local inference based on the is_api_model check. You can combine commercial APIs like GPT-4 with local checkpoints like Llama-3-70B in the same evaluators array.
Do I need to restart the pipeline after modifying the evaluator models?
Yes. Because the evaluators list and is_api_model classification are loaded when the Python module initializes, you must restart the evaluation script for changes to take effect. The pipeline does not hot-reload these configurations.
How does CreativeMath determine if a model is API-based or local?
The ModelWrapper constructor in src/models/model_loader.py checks if the model name exists in the hardcoded self.is_api_model list. If present, it initializes an API client via load_api_model; otherwise, it calls load_local_model to load weights from disk.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →