Delphi: Marin’s Open Scaling Suite for Systematic LLM Training
Delphi is Marin’s open-source scaling suite that transforms a base LLM training recipe into a reproducible pipeline spanning 3×10¹⁸ to 10²³ FLOPs, comprising a parametric scaling recipe, an automated training suite on Google TPUs, and a learned scaling law that predicts optimal model configurations for any compute budget.
Delphi provides the systematic engine behind the marin-community/marin repository, turning abstract training recipes into concrete, reproducible scaling experiments. This open scaling suite automates the entire workflow from IsoFLOP analysis to optimal checkpoint generation, enabling researchers to map any FLOP budget to its ideal model architecture using declarative, dependency-aware execution.
Core Components of the Delphi Scaling Suite
Delphi consists of three tightly coupled components expressed as Marin ArtifactStep objects, which enable declarative, dependency-aware execution and reproducible checkpoint generation.
Scaling Recipe
The scaling recipe provides a parametric mapping from a target FLOP budget to concrete model configurations—including depth, width, and batch size parameters. According to the Marin source code in README.md (lines 31-32), this mapping bridges theoretical compute targets with practical training configurations.
Scaling Suite
The scaling suite implements the recipe as a collection of end-to-end training jobs running on the Google TPU Research Cloud. As defined in experiments/references/reference_scaling_suite.py (lines 31-32), this component produces a family of checkpoints spanning the full compute range from 3×10¹⁸ to 10²³ FLOPs.
Scaling Law
The scaling law learns to predict the optimal model size for any FLOP budget by fitting quadratic loss curves on the "isoflop" checkpoints generated by the suite. This predictive component, implemented in lib/marin/src/marin/scaling_laws/__init__.py, drives the entire optimization loop by extracting FLOP-optimal token counts and model dimensions.
Automated Workflow: From Analysis to Training
Delphi automates the complete scaling workflow through two primary mechanisms implemented in experiments/references/reference_scaling_suite.py.
IsoFLOP Analysis
The suite performs IsoFLOP analysis by reading metrics from dozens of existing checkpoints, fitting quadratic loss curves per budget, and extracting the FLOP-optimal token count. Lines 52-58 of reference_scaling_suite.py handle this curve-fitting logic to determine optimal training durations for each compute level.
Optimal Training Pipeline
For each target budget (10²¹, 10²², 10²³ FLOPs), the scaling law predicts a model configuration, which the suite then launches on TPUs with appropriate hyper-parameters including learning rate and beta₂. Lines 21-28 of reference_scaling_suite.py demonstrate this configuration prediction and job scheduling logic.
Running the Delphi Scaling Suite
All components execute as Marin ArtifactStep objects (lines 67-84 of reference_scaling_suite.py), ensuring reproducible checkpoint generation through declarative pipelines that respect dependencies between analysis and training stages.
Execute the Full Scaling Ladder
Run the complete analysis and optimal training pipeline from the repository root:
if __name__ == "__main__":
from experiments.references.reference_scaling_suite import build, StepRunner
StepRunner().run([s.lower() for s in build()])
Fit Scaling Laws from Checkpoints
Analyze existing IsoFLOP checkpoints to derive scaling relationships using the fit_scaling_laws API:
from marin.scaling_laws import fit_scaling_laws, IsoFlopRecord
records = [
IsoFlopRecord(tokens=1e12, metric=0.02, flops=1e20, params=1e9, label="runA"),
IsoFlopRecord(tokens=2e12, metric=0.018, flops=2e20, params=2e9, label="runB"),
# … more records …
]
result = fit_scaling_laws(records) # → ScalingFit objects
print(result.scaling_fits) # mapping label → (α, A)
Predict Optimal Configurations
Use learned scaling fits to configure new training runs for specific FLOP targets:
from marin.scaling_laws import ScalingFit, predict_optimal_config
from experiments.references.completed_adamh import completed_adamh_heuristic, SEQ_LEN
scaling_fits = {"nemotron": ScalingFit(alpha=0.45, A=3.2e-4)}
candidate = predict_optimal_config(
scaling_fits=scaling_fits,
target_flops=5e22,
label="nemotron",
heuristic=completed_adamh_heuristic,
seq_len=SEQ_LEN,
)
print(candidate.model_config) # model depth/width etc.
print(candidate.tokens) # optimal token count
Key Source Files
Understanding Delphi requires familiarity with these critical paths in the marin-community/marin repository:
experiments/references/reference_scaling_suite.py– End-to-end implementation of IsoFLOP analysis and optimal training steps (lines 21-88).lib/marin/src/marin/scaling_laws/__init__.py– Core APIs includingfit_scaling_lawsandpredict_optimal_configthat power the scaling law predictions.experiments/sft/delphi_chat_template.py– Defines the Delphi-v0 chat template using the Llama-3 think/tool token protocol for supervised fine-tuning runs.lib/marin/src/marin/scaling_laws/scaling_plots.py– Visualization helpers for scaling-law fits and model-budget curves.
Summary
- Delphi is Marin’s open-source scaling suite that converts training recipes into reproducible pipelines spanning 3×10¹⁸ to 10²³ FLOPs.
- The architecture comprises three components: a parametric scaling recipe, an automated scaling suite on Google TPUs, and a learned scaling law.
- IsoFLOP analysis in
reference_scaling_suite.pyfits quadratic curves to determine optimal token counts for each budget. - All steps execute as Marin
ArtifactStepobjects, ensuring dependency-aware, declarative workflow execution. - The
fit_scaling_lawsandpredict_optimal_configAPIs inlib/marin/src/marin/scaling_laws/enable programmatic scaling predictions.
Frequently Asked Questions
What hardware does the Delphi scaling suite target?
Delphi targets the Google TPU Research Cloud for its scaling suite implementation. The training jobs in reference_scaling_suite.py are specifically configured to execute on TPUs, utilizing XLA-optimized training loops appropriate for large-scale language model training.
How does Delphi determine the optimal model size for a given compute budget?
Delphi uses IsoFLOP analysis to fit quadratic loss curves across checkpoints with equal FLOP budgets but varying model sizes. The fit_scaling_laws function (defined in lib/marin/src/marin/scaling_laws/__init__.py) processes these "isoflop" records to learn the relationship between parameters, tokens, and loss, enabling predict_optimal_config to calculate the optimal depth, width, and training duration for new budgets.
Can I run the Delphi scaling suite on hardware other than Google TPUs?
While the reference implementation in experiments/references/reference_scaling_suite.py targets TPUs, the scaling laws and recipe logic are hardware-agnostic. The ScalingFit objects and configuration prediction APIs work with any training infrastructure, though you would need to adapt the StepRunner execution and device placement code for GPUs or other accelerators.
What is the Delphi chat template used for?
The Delphi chat template (defined in experiments/sft/delphi_chat_template.py) implements a Llama-3 compatible token protocol including think/tool tokens. This template standardizes the conversation format for supervised fine-tuning (SFT) runs within the scaling suite, ensuring consistent data formatting across the 3×10¹⁸ to 10²³ FLOP checkpoint family.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →