How Configuration Is Managed for LLM Training Jobs in train-llm-from-scratch
The repository implements a layered JSON-based configuration system that merges dataclass defaults, base JSON files, stage-specific JSON files, and CLI overrides using the custom config/loader.py module.
Training large language models requires juggling hundreds of hyperparameters across multiple stages. The FareedKhan-dev/train-llm-from-scratch repository solves this with a robust, hierarchical configuration management system that resolves settings at runtime through a four-layer precedence stack.
The Four-Layer Configuration Hierarchy
Layer 1: Dataclass Defaults (Lowest Precedence)
The foundation of the system rests on Python dataclasses defined in config/post_training_config.py (for post-training stages) and config/config.py (for pre-training). These classes declare fields like vocabulary size, context length, and device settings, providing type-safe fallback values when no JSON or CLI values are specified.
Layer 2: Base JSON Configuration
The configs/base.json file contains model-wide hyperparameters shared across all training stages, including SFT, PPO, and DPO. When a stage-specific JSON file resides in a subdirectory (e.g., configs/smoke/sft.json), the loader automatically discovers and merges a sibling base.json in that same directory (configs/smoke/base.json).
Layer 3: Stage-Specific JSON Files
Individual training stages define their own configurations in files like configs/sft.json, configs/ppo.json, and configs/reward.json. These files override any overlapping keys from the base configuration, allowing stage-specific tuning of learning rates, batch sizes, and optimization settings.
Layer 4: CLI Overrides (Highest Precedence)
When invoking training scripts, users can pass --field key=value arguments. These command-line overrides are merged last in the load_config() function, giving the highest precedence and enabling ad-hoc experimentation without modifying JSON files.
How the Configuration Loader Works
The resolution logic lives in config/loader.py, which exposes the load_config function and a _deep_merge helper for nested dictionary merging. The loader sequentially applies the four layers, with later layers overwriting earlier ones.
The system gracefully handles unknown keys by printing warnings rather than raising exceptions, preventing crashes from typos or experimental parameters while alerting users to potential issues.
from config.loader import load_config
from config.post_training_config import SFTConfig
# Resolve the full SFT configuration with all layers applied
cfg = load_config(
SFTConfig,
json_path="configs/sft.json", # Stage-specific JSON
overrides={"t_lr": 3e-4}, # CLI overrides (highest precedence)
)
Running Training Jobs with Configuration Files
Training scripts in the scripts/ directory consume these resolved configurations through the loader module.
Run a supervised fine-tuning job with default configuration:
python scripts/train_sft.py \
--config configs/sft.json
Override a hyperparameter from the command line:
python scripts/train_sft.py \
--config configs/sft.json \
--field t_lr=3e-4
Use a custom configuration directory for smoke testing:
python scripts/train_sft.py \
--config configs/smoke/sft.json
Key Configuration Files
config/config.py– Global defaults for pre-training transformers (vocabulary size, device, etc.)config/loader.py– Core JSON loading and deep-merge logic (_deep_merge,load_config)config/post_training_config.py– Dataclasses defining SFT, PPO, and DPO stage-specific fieldsconfigs/base.json– Shared hyperparameters across all post-training stagesconfigs/sft.json– Supervised fine-tuning specific hyperparametersconfigs/ppo.json– PPO-specific training settingsscripts/train_sft.py– Entry point demonstrating configuration loading for SFTscripts/train_ppo.py– PPO training entry point using the same loader pattern
Summary
- The configuration system uses a four-layer precedence stack: dataclass defaults, base JSON, stage-specific JSON, and CLI overrides.
- The
config/loader.pymodule handles resolution via the_deep_mergehelper andload_configfunction. - Automatic discovery of sibling
base.jsonfiles enables organized configuration sets in subdirectories. - Unknown keys trigger warnings rather than crashes, providing a forgiving user experience for iterative development.
Frequently Asked Questions
What is the precedence order for configuration values?
Values cascade from lowest to highest: dataclass defaults < configs/base.json < stage-specific JSON < CLI overrides. The load_config function in config/loader.py merges these layers sequentially, with later layers overwriting earlier ones.
How does the loader handle unknown configuration keys?
The config/loader.py module prints warnings for unrecognized keys but continues execution. This prevents crashes due to typos or experimental parameters while alerting users to potential configuration issues.
Can I organize configurations into custom subdirectories?
Yes. When you place a stage JSON file in a subdirectory (e.g., configs/smoke/sft.json), the loader automatically discovers and merges a sibling base.json in that same directory (configs/smoke/base.json), allowing isolated configuration sets for different environments.
Where are the training stage dataclasses defined?
Post-training stage dataclasses like SFTConfig, PPOConfig, and DPOConfig are defined in config/post_training_config.py. Pre-training defaults reside in config/config.py. These dataclasses provide the schema and default values for the configuration system.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →