How Configuration Is Managed for LLM Training Jobs in train-llm-from-scratch

The repository implements a layered JSON-based configuration system that merges dataclass defaults, base JSON files, stage-specific JSON files, and CLI overrides using the custom config/loader.py module.

Training large language models requires juggling hundreds of hyperparameters across multiple stages. The FareedKhan-dev/train-llm-from-scratch repository solves this with a robust, hierarchical configuration management system that resolves settings at runtime through a four-layer precedence stack.

The Four-Layer Configuration Hierarchy

Layer 1: Dataclass Defaults (Lowest Precedence)

The foundation of the system rests on Python dataclasses defined in config/post_training_config.py (for post-training stages) and config/config.py (for pre-training). These classes declare fields like vocabulary size, context length, and device settings, providing type-safe fallback values when no JSON or CLI values are specified.

Layer 2: Base JSON Configuration

The configs/base.json file contains model-wide hyperparameters shared across all training stages, including SFT, PPO, and DPO. When a stage-specific JSON file resides in a subdirectory (e.g., configs/smoke/sft.json), the loader automatically discovers and merges a sibling base.json in that same directory (configs/smoke/base.json).

Layer 3: Stage-Specific JSON Files

Individual training stages define their own configurations in files like configs/sft.json, configs/ppo.json, and configs/reward.json. These files override any overlapping keys from the base configuration, allowing stage-specific tuning of learning rates, batch sizes, and optimization settings.

Layer 4: CLI Overrides (Highest Precedence)

When invoking training scripts, users can pass --field key=value arguments. These command-line overrides are merged last in the load_config() function, giving the highest precedence and enabling ad-hoc experimentation without modifying JSON files.

How the Configuration Loader Works

The resolution logic lives in config/loader.py, which exposes the load_config function and a _deep_merge helper for nested dictionary merging. The loader sequentially applies the four layers, with later layers overwriting earlier ones.

The system gracefully handles unknown keys by printing warnings rather than raising exceptions, preventing crashes from typos or experimental parameters while alerting users to potential issues.

from config.loader import load_config
from config.post_training_config import SFTConfig

# Resolve the full SFT configuration with all layers applied

cfg = load_config(
    SFTConfig,
    json_path="configs/sft.json",          # Stage-specific JSON

    overrides={"t_lr": 3e-4},              # CLI overrides (highest precedence)

)

Running Training Jobs with Configuration Files

Training scripts in the scripts/ directory consume these resolved configurations through the loader module.

Run a supervised fine-tuning job with default configuration:

python scripts/train_sft.py \
    --config configs/sft.json

Override a hyperparameter from the command line:

python scripts/train_sft.py \
    --config configs/sft.json \
    --field t_lr=3e-4

Use a custom configuration directory for smoke testing:

python scripts/train_sft.py \
    --config configs/smoke/sft.json

Key Configuration Files

Summary

  • The configuration system uses a four-layer precedence stack: dataclass defaults, base JSON, stage-specific JSON, and CLI overrides.
  • The config/loader.py module handles resolution via the _deep_merge helper and load_config function.
  • Automatic discovery of sibling base.json files enables organized configuration sets in subdirectories.
  • Unknown keys trigger warnings rather than crashes, providing a forgiving user experience for iterative development.

Frequently Asked Questions

What is the precedence order for configuration values?

Values cascade from lowest to highest: dataclass defaults < configs/base.json < stage-specific JSON < CLI overrides. The load_config function in config/loader.py merges these layers sequentially, with later layers overwriting earlier ones.

How does the loader handle unknown configuration keys?

The config/loader.py module prints warnings for unrecognized keys but continues execution. This prevents crashes due to typos or experimental parameters while alerting users to potential configuration issues.

Can I organize configurations into custom subdirectories?

Yes. When you place a stage JSON file in a subdirectory (e.g., configs/smoke/sft.json), the loader automatically discovers and merges a sibling base.json in that same directory (configs/smoke/base.json), allowing isolated configuration sets for different environments.

Where are the training stage dataclasses defined?

Post-training stage dataclasses like SFTConfig, PPOConfig, and DPOConfig are defined in config/post_training_config.py. Pre-training defaults reside in config/config.py. These dataclasses provide the schema and default values for the configuration system.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →