# Delphi: Marin’s Open Scaling Suite for Systematic LLM Training

> Discover Delphi, Marin's open scaling suite for systematic LLM training. Automate reproducible pipelines from 3×10¹⁸ to 10²³ FLOPs with a learned scaling law.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: deep-dive
- Published: 2026-09-10

---

**Delphi is Marin’s open-source scaling suite that transforms a base LLM training recipe into a reproducible pipeline spanning 3×10¹⁸ to 10²³ FLOPs, comprising a parametric scaling recipe, an automated training suite on Google TPUs, and a learned scaling law that predicts optimal model configurations for any compute budget.**

Delphi provides the systematic engine behind the `marin-community/marin` repository, turning abstract training recipes into concrete, reproducible scaling experiments. This open scaling suite automates the entire workflow from IsoFLOP analysis to optimal checkpoint generation, enabling researchers to map any FLOP budget to its ideal model architecture using declarative, dependency-aware execution.

## Core Components of the Delphi Scaling Suite

Delphi consists of three tightly coupled components expressed as **Marin `ArtifactStep` objects**, which enable declarative, dependency-aware execution and reproducible checkpoint generation.

### Scaling Recipe

The **scaling recipe** provides a parametric mapping from a target FLOP budget to concrete model configurations—including depth, width, and batch size parameters. According to the Marin source code in [`README.md`](https://github.com/marin-community/marin/blob/main/README.md) (lines 31-32), this mapping bridges theoretical compute targets with practical training configurations.

### Scaling Suite

The **scaling suite** implements the recipe as a collection of end-to-end training jobs running on the **Google TPU Research Cloud**. As defined in [`experiments/references/reference_scaling_suite.py`](https://github.com/marin-community/marin/blob/main/experiments/references/reference_scaling_suite.py) (lines 31-32), this component produces a family of checkpoints spanning the full compute range from 3×10¹⁸ to 10²³ FLOPs.

### Scaling Law

The **scaling law** learns to predict the optimal model size for any FLOP budget by fitting quadratic loss curves on the "isoflop" checkpoints generated by the suite. This predictive component, implemented in [`lib/marin/src/marin/scaling_laws/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/scaling_laws/__init__.py), drives the entire optimization loop by extracting FLOP-optimal token counts and model dimensions.

## Automated Workflow: From Analysis to Training

Delphi automates the complete scaling workflow through two primary mechanisms implemented in [`experiments/references/reference_scaling_suite.py`](https://github.com/marin-community/marin/blob/main/experiments/references/reference_scaling_suite.py).

### IsoFLOP Analysis

The suite performs **IsoFLOP analysis** by reading metrics from dozens of existing checkpoints, fitting quadratic loss curves per budget, and extracting the FLOP-optimal token count. Lines 52-58 of [`reference_scaling_suite.py`](https://github.com/marin-community/marin/blob/main/reference_scaling_suite.py) handle this curve-fitting logic to determine optimal training durations for each compute level.

### Optimal Training Pipeline

For each target budget (10²¹, 10²², 10²³ FLOPs), the scaling law predicts a model configuration, which the suite then launches on TPUs with appropriate hyper-parameters including learning rate and beta₂. Lines 21-28 of [`reference_scaling_suite.py`](https://github.com/marin-community/marin/blob/main/reference_scaling_suite.py) demonstrate this configuration prediction and job scheduling logic.

## Running the Delphi Scaling Suite

All components execute as **Marin `ArtifactStep` objects** (lines 67-84 of [`reference_scaling_suite.py`](https://github.com/marin-community/marin/blob/main/reference_scaling_suite.py)), ensuring reproducible checkpoint generation through declarative pipelines that respect dependencies between analysis and training stages.

### Execute the Full Scaling Ladder

Run the complete analysis and optimal training pipeline from the repository root:

```python
if __name__ == "__main__":
    from experiments.references.reference_scaling_suite import build, StepRunner
    StepRunner().run([s.lower() for s in build()])

```

### Fit Scaling Laws from Checkpoints

Analyze existing IsoFLOP checkpoints to derive scaling relationships using the `fit_scaling_laws` API:

```python
from marin.scaling_laws import fit_scaling_laws, IsoFlopRecord

records = [
    IsoFlopRecord(tokens=1e12, metric=0.02, flops=1e20, params=1e9, label="runA"),
    IsoFlopRecord(tokens=2e12, metric=0.018, flops=2e20, params=2e9, label="runB"),
    # … more records …

]

result = fit_scaling_laws(records)          # → ScalingFit objects

print(result.scaling_fits)                  # mapping label → (α, A)

```

### Predict Optimal Configurations

Use learned scaling fits to configure new training runs for specific FLOP targets:

```python
from marin.scaling_laws import ScalingFit, predict_optimal_config
from experiments.references.completed_adamh import completed_adamh_heuristic, SEQ_LEN

scaling_fits = {"nemotron": ScalingFit(alpha=0.45, A=3.2e-4)}
candidate = predict_optimal_config(
    scaling_fits=scaling_fits,
    target_flops=5e22,
    label="nemotron",
    heuristic=completed_adamh_heuristic,
    seq_len=SEQ_LEN,
)
print(candidate.model_config)   # model depth/width etc.

print(candidate.tokens)         # optimal token count

```

## Key Source Files

Understanding Delphi requires familiarity with these critical paths in the `marin-community/marin` repository:

- **[`experiments/references/reference_scaling_suite.py`](https://github.com/marin-community/marin/blob/main/experiments/references/reference_scaling_suite.py)** – End-to-end implementation of IsoFLOP analysis and optimal training steps (lines 21-88).
- **[`lib/marin/src/marin/scaling_laws/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/scaling_laws/__init__.py)** – Core APIs including `fit_scaling_laws` and `predict_optimal_config` that power the scaling law predictions.
- **[`experiments/sft/delphi_chat_template.py`](https://github.com/marin-community/marin/blob/main/experiments/sft/delphi_chat_template.py)** – Defines the Delphi-v0 chat template using the Llama-3 think/tool token protocol for supervised fine-tuning runs.
- **[`lib/marin/src/marin/scaling_laws/scaling_plots.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/scaling_laws/scaling_plots.py)** – Visualization helpers for scaling-law fits and model-budget curves.

## Summary

- Delphi is Marin’s **open-source scaling suite** that converts training recipes into reproducible pipelines spanning 3×10¹⁸ to 10²³ FLOPs.
- The architecture comprises three components: a **parametric scaling recipe**, an automated **scaling suite** on Google TPUs, and a learned **scaling law**.
- **IsoFLOP analysis** in [`reference_scaling_suite.py`](https://github.com/marin-community/marin/blob/main/reference_scaling_suite.py) fits quadratic curves to determine optimal token counts for each budget.
- All steps execute as **Marin `ArtifactStep` objects**, ensuring dependency-aware, declarative workflow execution.
- The `fit_scaling_laws` and `predict_optimal_config` APIs in `lib/marin/src/marin/scaling_laws/` enable programmatic scaling predictions.

## Frequently Asked Questions

### What hardware does the Delphi scaling suite target?

Delphi targets the **Google TPU Research Cloud** for its scaling suite implementation. The training jobs in [`reference_scaling_suite.py`](https://github.com/marin-community/marin/blob/main/reference_scaling_suite.py) are specifically configured to execute on TPUs, utilizing XLA-optimized training loops appropriate for large-scale language model training.

### How does Delphi determine the optimal model size for a given compute budget?

Delphi uses **IsoFLOP analysis** to fit quadratic loss curves across checkpoints with equal FLOP budgets but varying model sizes. The `fit_scaling_laws` function (defined in [`lib/marin/src/marin/scaling_laws/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/scaling_laws/__init__.py)) processes these "isoflop" records to learn the relationship between parameters, tokens, and loss, enabling `predict_optimal_config` to calculate the optimal depth, width, and training duration for new budgets.

### Can I run the Delphi scaling suite on hardware other than Google TPUs?

While the reference implementation in [`experiments/references/reference_scaling_suite.py`](https://github.com/marin-community/marin/blob/main/experiments/references/reference_scaling_suite.py) targets TPUs, the **scaling laws** and **recipe logic** are hardware-agnostic. The `ScalingFit` objects and configuration prediction APIs work with any training infrastructure, though you would need to adapt the `StepRunner` execution and device placement code for GPUs or other accelerators.

### What is the Delphi chat template used for?

The **Delphi chat template** (defined in [`experiments/sft/delphi_chat_template.py`](https://github.com/marin-community/marin/blob/main/experiments/sft/delphi_chat_template.py)) implements a Llama-3 compatible token protocol including think/tool tokens. This template standardizes the conversation format for supervised fine-tuning (SFT) runs within the scaling suite, ensuring consistent data formatting across the 3×10¹⁸ to 10²³ FLOP checkpoint family.