Delphi Scaling Law Methodology in Marin: Predicting Large Model Performance with IsoFLOP Analysis
Marin predicts large model performance using a two-step IsoFLOP scaling law pipeline that fits robust quadratic curves to empirical loss data, derives optimal token counts per compute budget, and extrapolates power-law relationships to forecast training configurations for any target FLOP budget.
The marin-community/marin repository implements a rigorous, statistically grounded approach to forecasting how large language models—specifically the Delphi family—will perform at scale. This Delphi scaling law methodology applies the Chinchilla/IsoFLOP framework to extrapolate from small-scale experiments to massive training runs. By analyzing the relationship between compute budget, token count, and loss metrics, Marin enables precise, data-driven decisions for model sizing and optimal training duration.
The IsoFLOP Scaling Pipeline Architecture
According to the Marin source code in lib/marin/src/marin/scaling_laws/isoflop_analysis.py, the Delphi scaling law methodology decomposes into three operational phases: empirical fitting, power-law regression, and configuration prediction.
Step 1: Fitting Empirical Optimal Token Counts
For each discrete compute bucket (FLOPs), the library fits a robust quadratic to the loss-versus-token-count curve using the robust_quad_logx function (lines 36-68). This implementation minimizes a Huber loss function to resist outliers from unstable training runs. The quadratic models the relationship as:
loss = a·log10(tokens)^2 + b·log10(tokens) + c
The empirical optimal token count D* derives from the vertex of this parabola: log_D_opt = -b/(2a), yielding D* = 10**log_D_opt. This represents the token count that minimizes loss for that specific compute budget.
Step 2: Power-Law Regression Across Budgets
The set of {(budget C, optimal tokens D*)} pairs undergoes log-log linear regression within the fit_scaling_laws function (lines 58-75). This produces the scaling law relationship D* ≈ A·C^α, where the exponent α and coefficient A characterize how optimal dataset size grows with compute investment. The intermediate ScalingFit dataclass (lines 44-57) stores these fitted parameters indexed by model label.
Step 3: Predicting Configurations for Target FLOPs
Given a target compute budget C_target, the predict_optimal_config function (lines 24-70) calculates the required token count as D_target = A·C_target**α. The system then enumerates candidate training configurations through a ScalingHeuristic implementation—such as DelphiHeuristic—and selects the most cost-effective configuration whose token budget satisfies or exceeds D_target, returning a CandidateConfig object ready for orchestration.
Core Data Structures and Robust Statistics
The IsoFlopRecord Dataclass
Experimental results populate IsoFlopRecord instances defined in isoflop_analysis.py (lines 81-103). These records encapsulate:
@dataclass
class IsoFlopRecord:
tokens: float # total tokens trained
metric: float # e.g. bits-per-byte
flops: float # training FLOPs (including 3× forward/backward)
params: float # model parameter count
label: str # experiment label (e.g. "delphi")
Robust Quadratic Implementation
The robust_quad_logx function differs from standard polynomial fitting by employing Huber loss rather than mean squared error. This statistical technique reduces the influence of outlier data points—such as those from failed training jobs or hardware interruptions—ensuring that the calculated optimal token counts reflect genuine model performance trends rather than experimental noise.
Practical Code Examples
Fitting Scaling Laws from Experiment Records
To derive scaling parameters from Delphi training runs, collect IsoFlopRecord instances and invoke the fitting pipeline:
from marin.scaling_laws.isoflop_analysis import IsoFlopRecord, fit_scaling_laws
# Records collected from small-scale Delphi experiments
records = [
IsoFlopRecord(tokens=1e10, metric=0.32, flops=1e18, params=2e9, label="delphi"),
IsoFlopRecord(tokens=2e10, metric=0.30, flops=1e18, params=2e9, label="delphi"),
# ... additional records across varying FLOP budgets
]
result = fit_scaling_laws(records)
print(result.scaling_fits) # {'delphi': ScalingFit(alpha=0.45, A=1.2e7)}
This returns a ScalingFit object containing the power-law coefficient and exponent specific to the Delphi model family.
Predicting Optimal Training Configurations
To forecast the optimal setup for a 5×10¹⁹ FLOP budget:
from marin.scaling_laws.isoflop_analysis import predict_optimal_config
from marin.scaling_laws.tpu_utils import DelphiHeuristic
scaling_fits = result.scaling_fits
heuristic = DelphiHeuristic(vocab_size=128256) # Concrete ScalingHeuristic
opt_cfg = predict_optimal_config(
scaling_fits=scaling_fits,
target_flops=5e19,
label="delphi",
heuristic=heuristic,
)
print(opt_cfg) # CandidateConfig with model_config, optimizer, batch_size, etc.
The predict_optimal_config function handles the mathematical extrapolation while the DelphiHeuristic provides hardware-aware candidate generation.
Visualizing Scaling Relationships
Marin includes plotting utilities to validate the fitted laws:
from marin.scaling_laws.scaling_plots import create_scaling_plot
fig = create_scaling_plot(
scaling_fits=result.scaling_fits,
title="Delphi Scaling Law",
output_path="delphi_scaling.html",
)
fig.show()
This generates interactive plots showing the IsoFLOP curves and the fitted power-law trend, aiding in validation of the scaling assumptions before committing to large-scale runs.
Key Files and Architecture
The Delphi scaling law methodology spans several critical files in the repository:
| File | Purpose |
|---|---|
lib/marin/src/marin/scaling_laws/isoflop_analysis.py |
Core implementation including robust_quad_logx, fit_scaling_laws, and predict_optimal_config |
lib/marin/src/marin/scaling_laws/scaling_plots.py |
Visualization utilities for IsoFLOP curves and scaling trends (lines 176-184) |
experiments/sft/configs/delphi_1e22.py |
Concrete configuration example supplying Delphi chat templates and checkpoint references |
experiments/sft/delphi_chat_template.py |
Token protocol definitions specific to Delphi experiments |
experiments/sft/launcher.py |
Orchestration layer integrating scaling-law predictions with training job deployment |
Summary
The Delphi scaling law methodology in Marin provides a systematic, statistically robust approach to extrapolating model performance:
- Robust quadratic fitting via
robust_quad_logxcalculates empirical optimal token counts per compute budget using Huber loss for outlier resistance - Power-law regression in
fit_scaling_lawsestablishes the relationshipD* ≈ A·C^αacross multiple FLOP budgets - Configuration prediction through
predict_optimal_configtranslates target compute budgets into actionable training specifications - Modular heuristics like
DelphiHeuristicenable hardware-aware candidate selection while maintaining the statistical rigor of the IsoFLOP framework
Frequently Asked Questions
What is the IsoFLOP approach in Marin?
The IsoFLOP approach in Marin involves training multiple model configurations across varying token counts while holding the total compute budget (FLOPs) constant. By measuring the loss at each token count within a fixed budget, Marin identifies the optimal token count D* that minimizes loss for that specific compute level. Collecting these optimal points across budgets generates the data necessary for power-law fitting.
How does Marin handle outliers when fitting scaling laws?
Marin uses Huber loss regression instead of standard least-squares in the robust_quad_logx function. This statistical technique reduces the weight of outlier data points—such as those from crashed training runs or hardware failures—ensuring that the fitted quadratic curves and subsequent optimal token calculations reflect genuine training dynamics rather than experimental artifacts.
Can the scaling law methodology be used for models other than Delphi?
Yes. While the delphi_1e22 configuration provides a concrete example with model-specific heuristics, the fit_scaling_laws and predict_optimal_config functions are model-agnostic. By supplying different IsoFlopRecord labels and implementing a custom ScalingHeuristic for your hardware environment, you can apply the same two-step pipeline to any transformer architecture or model family.
What compute budgets does the Delphi scaling law support?
The methodology supports arbitrary compute budgets through mathematical extrapolation. Once fitted on empirical data from smaller runs spanning multiple orders of magnitude in FLOPs, the power law D* ≈ A·C^α predicts optimal token counts for target budgets significantly larger than the training data, enabling reliable forecasting for billion-parameter scale Delphi models.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →