Delphi: Marin’s Open Scaling Suite for Systematic LLM Training

Delphi is Marin’s open-source scaling suite that transforms a base LLM training recipe into a reproducible pipeline spanning 3×10¹⁸ to 10²³ FLOPs, comprising a parametric scaling recipe, an automated training suite on Google TPUs, and a learned scaling law that predicts optimal model configurations for any compute budget.

Delphi provides the systematic engine behind the marin-community/marin repository, turning abstract training recipes into concrete, reproducible scaling experiments. This open scaling suite automates the entire workflow from IsoFLOP analysis to optimal checkpoint generation, enabling researchers to map any FLOP budget to its ideal model architecture using declarative, dependency-aware execution.

Core Components of the Delphi Scaling Suite

Delphi consists of three tightly coupled components expressed as Marin ArtifactStep objects, which enable declarative, dependency-aware execution and reproducible checkpoint generation.

Scaling Recipe

The scaling recipe provides a parametric mapping from a target FLOP budget to concrete model configurations—including depth, width, and batch size parameters. According to the Marin source code in README.md (lines 31-32), this mapping bridges theoretical compute targets with practical training configurations.

Scaling Suite

The scaling suite implements the recipe as a collection of end-to-end training jobs running on the Google TPU Research Cloud. As defined in experiments/references/reference_scaling_suite.py (lines 31-32), this component produces a family of checkpoints spanning the full compute range from 3×10¹⁸ to 10²³ FLOPs.

Scaling Law

The scaling law learns to predict the optimal model size for any FLOP budget by fitting quadratic loss curves on the "isoflop" checkpoints generated by the suite. This predictive component, implemented in lib/marin/src/marin/scaling_laws/__init__.py, drives the entire optimization loop by extracting FLOP-optimal token counts and model dimensions.

Automated Workflow: From Analysis to Training

Delphi automates the complete scaling workflow through two primary mechanisms implemented in experiments/references/reference_scaling_suite.py.

IsoFLOP Analysis

The suite performs IsoFLOP analysis by reading metrics from dozens of existing checkpoints, fitting quadratic loss curves per budget, and extracting the FLOP-optimal token count. Lines 52-58 of reference_scaling_suite.py handle this curve-fitting logic to determine optimal training durations for each compute level.

Optimal Training Pipeline

For each target budget (10²¹, 10²², 10²³ FLOPs), the scaling law predicts a model configuration, which the suite then launches on TPUs with appropriate hyper-parameters including learning rate and beta₂. Lines 21-28 of reference_scaling_suite.py demonstrate this configuration prediction and job scheduling logic.

Running the Delphi Scaling Suite

All components execute as Marin ArtifactStep objects (lines 67-84 of reference_scaling_suite.py), ensuring reproducible checkpoint generation through declarative pipelines that respect dependencies between analysis and training stages.

Execute the Full Scaling Ladder

Run the complete analysis and optimal training pipeline from the repository root:

if __name__ == "__main__":
    from experiments.references.reference_scaling_suite import build, StepRunner
    StepRunner().run([s.lower() for s in build()])

Fit Scaling Laws from Checkpoints

Analyze existing IsoFLOP checkpoints to derive scaling relationships using the fit_scaling_laws API:

from marin.scaling_laws import fit_scaling_laws, IsoFlopRecord

records = [
    IsoFlopRecord(tokens=1e12, metric=0.02, flops=1e20, params=1e9, label="runA"),
    IsoFlopRecord(tokens=2e12, metric=0.018, flops=2e20, params=2e9, label="runB"),
    # … more records …

]

result = fit_scaling_laws(records)          # → ScalingFit objects

print(result.scaling_fits)                  # mapping label → (α, A)

Predict Optimal Configurations

Use learned scaling fits to configure new training runs for specific FLOP targets:

from marin.scaling_laws import ScalingFit, predict_optimal_config
from experiments.references.completed_adamh import completed_adamh_heuristic, SEQ_LEN

scaling_fits = {"nemotron": ScalingFit(alpha=0.45, A=3.2e-4)}
candidate = predict_optimal_config(
    scaling_fits=scaling_fits,
    target_flops=5e22,
    label="nemotron",
    heuristic=completed_adamh_heuristic,
    seq_len=SEQ_LEN,
)
print(candidate.model_config)   # model depth/width etc.

print(candidate.tokens)         # optimal token count

Key Source Files

Understanding Delphi requires familiarity with these critical paths in the marin-community/marin repository:

Summary

  • Delphi is Marin’s open-source scaling suite that converts training recipes into reproducible pipelines spanning 3×10¹⁸ to 10²³ FLOPs.
  • The architecture comprises three components: a parametric scaling recipe, an automated scaling suite on Google TPUs, and a learned scaling law.
  • IsoFLOP analysis in reference_scaling_suite.py fits quadratic curves to determine optimal token counts for each budget.
  • All steps execute as Marin ArtifactStep objects, ensuring dependency-aware, declarative workflow execution.
  • The fit_scaling_laws and predict_optimal_config APIs in lib/marin/src/marin/scaling_laws/ enable programmatic scaling predictions.

Frequently Asked Questions

What hardware does the Delphi scaling suite target?

Delphi targets the Google TPU Research Cloud for its scaling suite implementation. The training jobs in reference_scaling_suite.py are specifically configured to execute on TPUs, utilizing XLA-optimized training loops appropriate for large-scale language model training.

How does Delphi determine the optimal model size for a given compute budget?

Delphi uses IsoFLOP analysis to fit quadratic loss curves across checkpoints with equal FLOP budgets but varying model sizes. The fit_scaling_laws function (defined in lib/marin/src/marin/scaling_laws/__init__.py) processes these "isoflop" records to learn the relationship between parameters, tokens, and loss, enabling predict_optimal_config to calculate the optimal depth, width, and training duration for new budgets.

Can I run the Delphi scaling suite on hardware other than Google TPUs?

While the reference implementation in experiments/references/reference_scaling_suite.py targets TPUs, the scaling laws and recipe logic are hardware-agnostic. The ScalingFit objects and configuration prediction APIs work with any training infrastructure, though you would need to adapt the StepRunner execution and device placement code for GPUs or other accelerators.

What is the Delphi chat template used for?

The Delphi chat template (defined in experiments/sft/delphi_chat_template.py) implements a Llama-3 compatible token protocol including think/tool tokens. This template standardizes the conversation format for supervised fine-tuning (SFT) runs within the scaling suite, ensuring consistent data formatting across the 3×10¹⁸ to 10²³ FLOP checkpoint family.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →