How the Delphi Scaling Suite Maps Compute Budgets to Model Configurations

The marin-community/marin repository implements a power-law scaling recipe that converts FLOP budgets into concrete model configurations through fitted empirical relationships.

The Delphi scaling suite in marin-community/marin automates one of the most critical decisions in large-scale training: turning a compute budget into an optimal model architecture. Rather than relying on manual tuning, the system fits scaling laws from historical training data and inverts them to generate configurations that maximize performance per FLOP. This article examines the complete mapping pipeline—from raw training records to launch-ready hyperparameters.

Fitting the Scaling Law from Empirical Data

The foundation of the compute-to-config mapping lives in marin.scaling_laws.isoflop_analysis.fit_scaling_laws at lib/marin/src/marin/scaling_laws/isoflop_analysis.py (lines 275-356). This function processes a collection of scaling records—historical training runs that pair model configurations with their final FLOP counts and performance metrics.

The fitting procedure estimates a power-law relationship:


N* = A · C^α

Where:

  • N* = optimal token count (the quantity that maximizes performance)
  • C = compute budget in FLOPs
  • α = fitted exponent (typically ~0.5 for Chinchilla-optimal scaling)
  • A = dataset-specific coefficient

Each fitted relationship is stored in a ScalingFit object, indexed by label (e.g., "delphi" for the Delphi dataset). Multiple labels can coexist, allowing the same infrastructure to serve different data distributions.

from marin.scaling_laws.isoflop_analysis import fit_scaling_laws

# Load historical training records

records = load_scaling_records()  # List[ScalingRecord]

# Fit power-law scaling relationships

fit_result = fit_scaling_laws(records)

# Access fitted parameters for a specific dataset

scaling_fit = fit_result.scaling_fits["delphi"]
print(f"Exponent α = {scaling_fit.alpha}, coefficient A = {scaling_fit.A}")

Predicting Optimal Configurations from FLOP Budgets

Once scaling laws are fitted, predict_optimal_config (lines 380-404 in isoflop_analysis.py) performs the inverse mapping: given a target compute budget, it computes the optimal token count and translates that into concrete architectural parameters.

The function signature and workflow:

from marin.scaling_laws.isoflop_analysis import predict_optimal_config

target_budget = 1e22  # 10²² FLOPs

label = "delphi"

optimal_cfg = predict_optimal_config(
    target_compute_budget=target_budget,
    label=label,
    scaling_fits=fit_result.scaling_fits,
)

# optimal_cfg contains:

# - model size (parameters)

# - hidden_dim

# - num_layers

# - num_heads

# - sequence length

# - training tokens

# - derived checkpoint name

The prediction logic:

  1. Inverts the scaling law to solve for N* given C
  2. Applies memory constraints to ensure the configuration fits on target hardware
  3. Discretizes continuous values to valid architectural choices (e.g., head dimensions that divide evenly)
  4. Validates against minimum viable model sizes for the dataset

Visualizing the Compute-to-Config Relationship

The suite includes visualization tooling to inspect how compute budgets map to model configurations. marin.scaling_laws.scaling_plots.create_scaling_plot (lines 176-227 in scaling_plots.py) renders the fitted scaling curve with the predicted optimum highlighted.

from marin.scaling_laws.scaling_plots import create_scaling_plot

fig = create_scaling_plot(
    scaling_fits=fit_result.scaling_fits,
    label="delphi",
    target_compute_budget=1e22,
)
fig.savefig("delphi_scaling_curve.png")

This visualization serves two purposes:

  • Validation — confirm the fitted law behaves sensibly across the budget range
  • Communication — demonstrate to stakeholders why a particular configuration was selected

Integration in Delphi Experiments

The complete workflow ties together in experiment configuration files. The file experiments/sft/configs/delphi_1e22.py demonstrates production usage: it imports the scaling utilities and invokes predict_optimal_config to derive checkpoint names and hyperparameters matching a 1e22 FLOP budget.

As documented in the repository README (line 31), "the scaling recipe maps compute budgets to model configurations." This integration ensures that scaling decisions are reproducible and version-controlled rather than ad-hoc.

Key Implementation Files

Path Purpose
lib/marin/src/marin/scaling_laws/isoflop_analysis.py Core scaling law fitting (fit_scaling_laws) and config prediction (predict_optimal_config)
lib/marin/src/marin/scaling_laws/scaling_plots.py Visualization of scaling relationships
experiments/sft/configs/delphi_1e22.py Example experiment config using the scaling utilities
README.md (repository root) High-level description of the Delphi suite's three components

Summary

The Delphi scaling suite maps compute budgets to model configurations through a rigorous empirical pipeline:

  • Collect training records linking model scales to FLOP counts and performance
  • Fit power-law relationships of the form N* = A·C^α per dataset label
  • Invert the fitted law to compute optimal token counts for any target budget
  • Assemble concrete architectures satisfying those token counts with hardware constraints
  • Visualize the relationship for validation and communication

This system eliminates manual configuration tuning while ensuring Chinchilla-optimal or custom-scaling tradeoffs are applied consistently across experiments.

Frequently Asked Questions

What FLOP range does the Delphi scaling suite support?

The scaling suite supports any computable budget where empirical data exists to fit the scaling law. In practice, Delphi experiments target budgets from 1e20 to 1e24 FLOPs. Extrapolation beyond fitted ranges triggers warnings, as power-law assumptions may break down at extreme scales.

How does the suite handle multiple datasets with different scaling properties?

Each dataset receives its own ScalingFit entry in the scaling_fits dictionary, indexed by string label. The predict_optimal_config function accepts a label parameter to select the appropriate fitted law. This allows the same infrastructure to serve vision-language, text-only, or code datasets with distinct optimal compute allocations.

Can I override the fitted scaling law for specific experiments?

Yes. While predict_optimal_config uses fitted coefficients by default, you can construct a custom ScalingFit object with manually specified alpha and A values. Pass this custom fit to predict_optimal_config in place of the empirically fitted one for ablation studies or theoretical comparisons.

Where are the scaling records stored and how are they updated?

The raw analysis indicates load_scaling_records() retrieves historical data, but the specific storage backend (local JSON, cloud database, or experiment tracking system) depends on your marin deployment. Records typically accumulate automatically from completed training runs that export ScalingRecord objects compatible with fit_scaling_laws.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →