# Delphi Scaling Law Methodology in Marin: Predicting Large Model Performance with IsoFLOP Analysis

> Discover the Delphi scaling law methodology in Marin for predicting large model performance. Learn how IsoFLOP analysis forecasts training configurations for any FLOP budget.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: deep-dive
- Published: 2026-08-28

---

**Marin predicts large model performance using a two-step IsoFLOP scaling law pipeline that fits robust quadratic curves to empirical loss data, derives optimal token counts per compute budget, and extrapolates power-law relationships to forecast training configurations for any target FLOP budget.**

The marin-community/marin repository implements a rigorous, statistically grounded approach to forecasting how large language models—specifically the Delphi family—will perform at scale. This Delphi scaling law methodology applies the Chinchilla/IsoFLOP framework to extrapolate from small-scale experiments to massive training runs. By analyzing the relationship between compute budget, token count, and loss metrics, Marin enables precise, data-driven decisions for model sizing and optimal training duration.

## The IsoFLOP Scaling Pipeline Architecture

According to the Marin source code in [`lib/marin/src/marin/scaling_laws/isoflop_analysis.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/scaling_laws/isoflop_analysis.py), the Delphi scaling law methodology decomposes into three operational phases: empirical fitting, power-law regression, and configuration prediction.

### Step 1: Fitting Empirical Optimal Token Counts

For each discrete compute bucket (FLOPs), the library fits a **robust quadratic** to the loss-versus-token-count curve using the `robust_quad_logx` function (lines 36-68). This implementation minimizes a **Huber loss** function to resist outliers from unstable training runs. The quadratic models the relationship as:

```python
loss = a·log10(tokens)^2 + b·log10(tokens) + c

```

The empirical optimal token count `D*` derives from the vertex of this parabola: `log_D_opt = -b/(2a)`, yielding `D* = 10**log_D_opt`. This represents the token count that minimizes loss for that specific compute budget.

### Step 2: Power-Law Regression Across Budgets

The set of `{(budget C, optimal tokens D*)}` pairs undergoes **log-log linear regression** within the `fit_scaling_laws` function (lines 58-75). This produces the scaling law relationship `D* ≈ A·C^α`, where the exponent α and coefficient A characterize how optimal dataset size grows with compute investment. The intermediate `ScalingFit` dataclass (lines 44-57) stores these fitted parameters indexed by model label.

### Step 3: Predicting Configurations for Target FLOPs

Given a target compute budget `C_target`, the `predict_optimal_config` function (lines 24-70) calculates the required token count as `D_target = A·C_target**α`. The system then enumerates candidate training configurations through a `ScalingHeuristic` implementation—such as `DelphiHeuristic`—and selects the most cost-effective configuration whose token budget satisfies or exceeds `D_target`, returning a `CandidateConfig` object ready for orchestration.

## Core Data Structures and Robust Statistics

### The IsoFlopRecord Dataclass

Experimental results populate `IsoFlopRecord` instances defined in [`isoflop_analysis.py`](https://github.com/marin-community/marin/blob/main/isoflop_analysis.py) (lines 81-103). These records encapsulate:

```python
@dataclass
class IsoFlopRecord:
    tokens: float    # total tokens trained

    metric: float    # e.g. bits-per-byte

    flops: float     # training FLOPs (including 3× forward/backward)

    params: float    # model parameter count

    label: str       # experiment label (e.g. "delphi")

```

### Robust Quadratic Implementation

The `robust_quad_logx` function differs from standard polynomial fitting by employing **Huber loss** rather than mean squared error. This statistical technique reduces the influence of outlier data points—such as those from failed training jobs or hardware interruptions—ensuring that the calculated optimal token counts reflect genuine model performance trends rather than experimental noise.

## Practical Code Examples

### Fitting Scaling Laws from Experiment Records

To derive scaling parameters from Delphi training runs, collect `IsoFlopRecord` instances and invoke the fitting pipeline:

```python
from marin.scaling_laws.isoflop_analysis import IsoFlopRecord, fit_scaling_laws

# Records collected from small-scale Delphi experiments

records = [
    IsoFlopRecord(tokens=1e10, metric=0.32, flops=1e18, params=2e9, label="delphi"),
    IsoFlopRecord(tokens=2e10, metric=0.30, flops=1e18, params=2e9, label="delphi"),
    # ... additional records across varying FLOP budgets

]

result = fit_scaling_laws(records)
print(result.scaling_fits)  # {'delphi': ScalingFit(alpha=0.45, A=1.2e7)}

```

This returns a `ScalingFit` object containing the power-law coefficient and exponent specific to the Delphi model family.

### Predicting Optimal Training Configurations

To forecast the optimal setup for a 5×10¹⁹ FLOP budget:

```python
from marin.scaling_laws.isoflop_analysis import predict_optimal_config
from marin.scaling_laws.tpu_utils import DelphiHeuristic

scaling_fits = result.scaling_fits
heuristic = DelphiHeuristic(vocab_size=128256)  # Concrete ScalingHeuristic

opt_cfg = predict_optimal_config(
    scaling_fits=scaling_fits,
    target_flops=5e19,
    label="delphi",
    heuristic=heuristic,
)

print(opt_cfg)  # CandidateConfig with model_config, optimizer, batch_size, etc.

```

The `predict_optimal_config` function handles the mathematical extrapolation while the `DelphiHeuristic` provides hardware-aware candidate generation.

### Visualizing Scaling Relationships

Marin includes plotting utilities to validate the fitted laws:

```python
from marin.scaling_laws.scaling_plots import create_scaling_plot

fig = create_scaling_plot(
    scaling_fits=result.scaling_fits,
    title="Delphi Scaling Law",
    output_path="delphi_scaling.html",
)
fig.show()

```

This generates interactive plots showing the IsoFLOP curves and the fitted power-law trend, aiding in validation of the scaling assumptions before committing to large-scale runs.

## Key Files and Architecture

The Delphi scaling law methodology spans several critical files in the repository:

| File | Purpose |
|------|---------|
| [`lib/marin/src/marin/scaling_laws/isoflop_analysis.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/scaling_laws/isoflop_analysis.py) | Core implementation including `robust_quad_logx`, `fit_scaling_laws`, and `predict_optimal_config` |
| [`lib/marin/src/marin/scaling_laws/scaling_plots.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/scaling_laws/scaling_plots.py) | Visualization utilities for IsoFLOP curves and scaling trends (lines 176-184) |
| [`experiments/sft/configs/delphi_1e22.py`](https://github.com/marin-community/marin/blob/main/experiments/sft/configs/delphi_1e22.py) | Concrete configuration example supplying Delphi chat templates and checkpoint references |
| [`experiments/sft/delphi_chat_template.py`](https://github.com/marin-community/marin/blob/main/experiments/sft/delphi_chat_template.py) | Token protocol definitions specific to Delphi experiments |
| [`experiments/sft/launcher.py`](https://github.com/marin-community/marin/blob/main/experiments/sft/launcher.py) | Orchestration layer integrating scaling-law predictions with training job deployment |

## Summary

The Delphi scaling law methodology in Marin provides a systematic, statistically robust approach to extrapolating model performance:

- **Robust quadratic fitting** via `robust_quad_logx` calculates empirical optimal token counts per compute budget using Huber loss for outlier resistance
- **Power-law regression** in `fit_scaling_laws` establishes the relationship `D* ≈ A·C^α` across multiple FLOP budgets
- **Configuration prediction** through `predict_optimal_config` translates target compute budgets into actionable training specifications
- **Modular heuristics** like `DelphiHeuristic` enable hardware-aware candidate selection while maintaining the statistical rigor of the IsoFLOP framework

## Frequently Asked Questions

### What is the IsoFLOP approach in Marin?

The IsoFLOP approach in Marin involves training multiple model configurations across varying token counts while holding the total compute budget (FLOPs) constant. By measuring the loss at each token count within a fixed budget, Marin identifies the optimal token count `D*` that minimizes loss for that specific compute level. Collecting these optimal points across budgets generates the data necessary for power-law fitting.

### How does Marin handle outliers when fitting scaling laws?

Marin uses **Huber loss** regression instead of standard least-squares in the `robust_quad_logx` function. This statistical technique reduces the weight of outlier data points—such as those from crashed training runs or hardware failures—ensuring that the fitted quadratic curves and subsequent optimal token calculations reflect genuine training dynamics rather than experimental artifacts.

### Can the scaling law methodology be used for models other than Delphi?

Yes. While the `delphi_1e22` configuration provides a concrete example with model-specific heuristics, the `fit_scaling_laws` and `predict_optimal_config` functions are model-agnostic. By supplying different `IsoFlopRecord` labels and implementing a custom `ScalingHeuristic` for your hardware environment, you can apply the same two-step pipeline to any transformer architecture or model family.

### What compute budgets does the Delphi scaling law support?

The methodology supports arbitrary compute budgets through mathematical extrapolation. Once fitted on empirical data from smaller runs spanning multiple orders of magnitude in FLOPs, the power law `D* ≈ A·C^α` predicts optimal token counts for target budgets significantly larger than the training data, enabling reliable forecasting for billion-parameter scale Delphi models.