How CreativeMath Ensures Reproducibility of Experiments: A Deep Dive into the Configuration-Driven Architecture

CreativeMath guarantees reproducibility of experiments by centralizing the random seed in a single config.json file and exposing it through a lightweight loader in src/config.py, ensuring all downstream modules share the same deterministic configuration.

The junyiye/creativemath repository implements a disciplined, configuration-first approach to scientific computing where every experiment's behavior is governed by immutable settings defined at startup. By treating the reproducibility of experiments as a core architectural requirement rather than an afterthought, the codebase eliminates hidden state and hard-coded values that typically undermine deterministic behavior in machine learning workflows.

Centralized Configuration as the Foundation for Reproducibility

At the heart of CreativeMath's reproducibility strategy lies a single source of truth: the config.json file located at the repository root. This JSON file contains an experiment object that defines the default random seed, ensuring that all stochastic operations can be made deterministic through one centralized value.

{
  "experiment": {
    "seed": 42
  }
}

The src/config.py module provides the mechanism to expose these settings to the rest of the application. It implements a load_config() function that reads config.json once at import time and exposes the contents as a module-level dictionary named config.


# src/config.py

import json

def load_config():
    with open('config.json', 'r') as f:
        return json.load(f)

config = load_config()

This design pattern ensures that any Python script importing config receives the same immutable dictionary, eliminating the risk of configuration drift between different parts of the pipeline.

Seed Propagation Across Generation and Evaluation Modules

Both src/generation.py and src/evaluation.py begin by importing the configuration object, making the experiment seed immediately available throughout the codebase.


# src/generation.py and src/evaluation.py

from config import config

# Access the reproducibility seed

seed = config["experiment"]["seed"]

While the current implementation in these files does not directly invoke random number generators, the architectural design makes it trivial to enforce deterministic behavior. When developers introduce stochastic operations—such as dataset shuffling, model initialization, or sampling—they can apply the seed consistently using the centralized value from config["experiment"]["seed"].

Deterministic Checkpointing and Structured Logging

Beyond random seed management, CreativeMath reinforces reproducibility through systematic state preservation. The config.json file specifies a save_interval parameter that guarantees intermediate results are flushed to disk every N samples.

{
  "experiment": {
    "seed": 42,
    "save_interval": 100
  }
}

This deterministic checkpointing prevents loss of experimental state if a run is interrupted, allowing researchers to resume from exact prior conditions.

Additionally, the logging configuration derived from config["logging"] ensures all runs write to timestamped files inside the logs/ directory. Coupled with the fixed seed, these logs create a complete audit trail that enables exact replay and debugging of any experimental run.

Implementing Reproducibility in Practice

To make any module deterministic, developers insert seed initialization at the entry point using the centralized configuration value. The following pattern would be placed at the top of generation.py, evaluation.py, or any new script introducing randomness:

from config import config
import random
import numpy as np

# import torch  # Uncomment if using PyTorch

# Enforce deterministic behavior using the centralized seed

random.seed(config["experiment"]["seed"])
np.random.seed(config["experiment"]["seed"])

# torch.manual_seed(config["experiment"]["seed"])

Because all modules read from the same config dictionary imported from src/config.py, changing the seed value in config.json is the only modification required to obtain different deterministic behavior across the entire pipeline. This eliminates scattered hard-coded seeds and makes reproducibility transparent to both developers and users.

Summary

  • Single source of truth: The config.json file contains the experiment.seed value that controls all randomness in the pipeline.
  • Unified configuration loader: src/config.py exposes settings via a module-level config dictionary, ensuring all scripts read identical values.
  • Deterministic checkpoints: The save_interval parameter in config.json ensures intermediate states are preserved at regular intervals.
  • Traceable execution: Timestamped logs in logs/ combined with fixed seeds create auditable experimental records.
  • Extensible design: Modules like src/generation.py and src/evaluation.py are architected to accept seed initialization without refactoring when stochastic operations are added.

Frequently Asked Questions

How do I change the random seed for a new experiment?

Change the "seed" value inside the "experiment" object in config.json. Because src/config.py loads this file at import time, any script importing config will automatically receive the updated seed value, ensuring reproducibility of experiments with the new configuration.

Which files are responsible for maintaining reproducibility in CreativeMath?

The reproducibility framework relies on four key files: config.json stores the seed and checkpointing parameters; src/config.py loads and exposes these settings; src/generation.py and src/evaluation.py consume the configuration to ensure deterministic behavior in the main pipeline stages.

What happens if my experiment is interrupted during a long run?

The save_interval parameter in config.json ensures intermediate results are written to disk at regular intervals. Combined with the deterministic seed, this allows you to resume from the last checkpoint and obtain identical results to an uninterrupted run, maintaining full reproducibility.

Can I use this configuration system with PyTorch or TensorFlow?

Yes. The config["experiment"]["seed"] value can be passed to framework-specific seed functions such as torch.manual_seed() or tf.random.set_seed(). Because src/config.py makes the seed available as a standard Python integer, it integrates seamlessly with any machine learning library's deterministic mode.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →