# How Soup CLI Tracks Experiments and Metrics: A Deep Dive into the SQLite-Based ExperimentTracker

> Discover how Soup CLI tracks experiments and metrics with its SQLite-based ExperimentTracker. Learn about run metadata, training metrics, and cost estimates stored locally.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: deep-dive
- Published: 2026-09-06

---

**Soup CLI tracks experiments and metrics using a local SQLite database (`experiments.db`) managed by the `ExperimentTracker` class, which stores run metadata, per-step training metrics, evaluation results, and cost estimates in the user's `.soup` directory.**

The Soup CLI (MakazhanAlpamys/Soup) implements a lightweight yet comprehensive experiment tracking system that requires no external servers. At its core lies the **`ExperimentTracker`** class in [[`src/soup_cli/experiment/tracker.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/experiment/tracker.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/experiment/tracker.py), which orchestrates all experiment persistence through a local SQLite database. This article explains exactly how Soup CLI tracks experiments and metrics, from database schema to run lifecycle management.

## The SQLite Database Architecture

All experiment data lives in `$HOME/.soup/experiments.db`. When `ExperimentTracker` initializes, it executes `_ensure_schema()` to create tables if absent.

### Core Tables

| Table | Purpose |
|-------|---------|
| **runs** | Run metadata: `run_id`, `experiment_name`, timestamps, status, device info, cost estimates, JSON config |
| **metrics** | Per-step training data: `step`, `epoch`, `loss`, `lr`, `grad_norm`, `speed`, `gpu_mem`, `timestamp` |
| **eval_results** | Benchmark scores with provenance tracking for reproducible comparisons |
| **checkpoint_quality** / **forgetting_eval** | Extended analysis for model quality and catastrophic forgetting |

The schema definition resides in the `_SCHEMA_SQL` constant. For backwards compatibility, new columns like `cost_usd` and `run_kind` are added lazily via `ALTER TABLE` calls.

## Run Lifecycle Methods

The tracker implements a state machine for run execution. Here's how each method progresses a run from start to finish:

- **`generate_run_id()`** — Creates sortable unique IDs like `run_20240930_154523_a1b2c3d4`
- **`start_run()`** — Inserts a *running* row with device, GPU memory, base model, and task metadata
- **`launch_run()`** — For async launches: stores *launching* status before child process starts
- **`mark_running()`** — Updates to *running* and records the child PID
- **`finish_execution()`** — Writes terminal status (`completed`/`failed`/`terminated`) without overwriting training summaries
- **`finish_run()`** — Marks *completed*, writes summary fields, and computes cost via `soup_cli.utils.run_cost`
- **`fail_run()`** — Directly flags failures
- **`expunge_stale_launching_runs()`** — Cleans timed-out *launching* rows if PID is dead
- **`delete_run()`** — Removes a run and all associated metrics/eval data

### Orphaned Run Recovery

If a process crashes, `_reconcile_orphaned_run()` detects dead PIDs and rewrites status to *terminated* with unknown exit code. This prevents zombie entries from appearing successful.

## Recording Training Metrics

During training, any trainer (PPO, SFT, or custom) calls `log_metrics()`:

```python
tracker.log_metrics(
    run_id,
    step=global_step,
    epoch=current_epoch,
    loss=loss_value,
    lr=learning_rate,
    grad_norm=gradient_norm,
    speed=tokens_per_sec,
    gpu_mem="38GB",
)

```

Each call writes one row to the **metrics** table with an automatic timestamp. Retrieval methods include:

- **`get_metrics(run_id)`** — Returns all metric rows for a run
- **`get_metric_series(run_id, metric)`** — Extracts float list for plotting; falls back to `_eval_score_series` if no step data exists

### Example: Basic Training Script

```python
from soup_cli.experiment.tracker import ExperimentTracker

# Initialize tracker (creates DB if missing)

tracker = ExperimentTracker()

# Start run

config = {"base": "meta-llama/Llama-2-7b", "task": "sft"}
run_id = tracker.start_run(
    config,
    device="cuda",
    device_name="NVIDIA A100",
    gpu_info={"memory_total": "40GB"}
)

# Log 5 training steps

for step in range(1, 6):
    tracker.log_metrics(
        run_id,
        step=step,
        epoch=step / 5,
        loss=2.5 / step,
        lr=1e-5,
        grad_norm=0.8,
        speed=120.0,
        gpu_mem="38GB",
    )

# Finalize with summary statistics

tracker.finish_run(
    run_id,
    initial_loss=2.5,
    final_loss=0.5,
    total_steps=5,
    duration_secs=300,
    output_dir="/home/user/.soup/runs/run_20240930_154523_a1b2c3d4",
)

# Retrieve for visualization

loss_series = tracker.get_metric_series(run_id, "loss")
print("Loss per step:", loss_series)

tracker.close()

```

## Evaluation and Benchmark Tracking

After running evaluations, `save_eval_result()` persists scores with provenance:

```python
tracker.save_eval_result(
    model_path="/home/user/.soup/models/finetuned.pt",
    benchmark="hellaswag",
    score=0.842,
    details={"provenance": "baseline-v1"},
    run_id=run_id,
)

```

This stamps a `current_baseline_stamp` enabling reproducible comparisons across runs.

## Cost Estimation Integration

During `finish_run()`, the tracker computes:

1. GPU hourly rate via `lookup_gpu_rate()`
2. Estimated USD cost via `estimate_run_cost_usd()`

These populate `cost_usd` and `cost_gpu_label` columns for budget reporting.

## CLI Integration Points

The tracker serves all user-facing commands:

| Command | Usage |
|---------|-------|
| `soup train` | `start_run()` → `log_metrics()` loop → `finish_run()` |
| `soup eval` | `save_eval_result()` for benchmark scores |
| `soup runs` | Queries `get_metrics()` and run metadata for tabular display |
| `soup monitor` | Live metric forwarding via [`monitoring/callback.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/monitoring/callback.py) |
| `soup ui` | Pulls data for interactive graphs and tables |

## Key Source Files

| File | Role |
|------|------|
| [[`src/soup_cli/experiment/tracker.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/experiment/tracker.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/experiment/tracker.py) | Core `ExperimentTracker` class with all lifecycle and logging methods |
| [[`src/soup_cli/utils/constants.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/constants.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/constants.py) | Defines `EXPERIMENTS_DB` and `SOUP_DIR` paths |
| [[`src/soup_cli/trainer/ppo.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/ppo.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/ppo.py) | PPO trainer calling `log_metrics()` each step |
| [[`src/soup_cli/trainer/mlx_sft.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/mlx_sft.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/mlx_sft.py) | SFT trainer implementation with metric logging |
| [[`src/soup_cli/monitoring/callback.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/monitoring/callback.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/monitoring/callback.py) | Live monitor callback for UI metric streaming |
| [[`src/soup_cli/commands/runs.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/runs.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/runs.py) | CLI display of stored runs and metrics |
| [[`src/soup_cli/commands/ui.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/ui.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/ui.py) | UI entry point querying tracker for visualization |

## Summary

- Soup CLI tracks experiments through a **local SQLite database** at `$HOME/.soup/experiments.db`, eliminating external dependencies
- The **`ExperimentTracker`** class in [`src/soup_cli/experiment/tracker.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/experiment/tracker.py) provides complete lifecycle management: `start_run()` → `log_metrics()` → `finish_run()`
- Four main tables organize data: **runs**, **metrics**, **eval_results**, and extended quality tables
- Automatic **cost estimation** and **orphaned run reconciliation** ensure data integrity and budget visibility
- All CLI commands (`train`, `eval`, `runs`, `monitor`, `ui`) share this single tracking interface

## Frequently Asked Questions

### Where does Soup CLI store experiment data?

Soup CLI stores all data in `$HOME/.soup/experiments.db`, a local SQLite database. The path is defined in [`src/soup_cli/utils/constants.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/constants.py) via the `EXPERIMENTS_DB` constant. This local-first approach requires no cloud setup or external servers.

### How do I retrieve metrics from a completed run?

Call `tracker.get_metrics(run_id)` for full row data or `tracker.get_metric_series(run_id, "loss")` for a list of float values suitable for plotting. If no per-step metrics exist, the method automatically falls back to evaluation scores from the `eval_results` table.

### What happens if a training process crashes unexpectedly?

The `ExperimentTracker` detects orphaned runs through `_reconcile_orphaned_run()`, which checks if the recorded PID is still alive. Dead processes trigger status rewriting to *terminated* with unknown exit code, preventing stale entries from appearing successful in `soup runs` listings.

### Can I track custom metrics beyond loss and learning rate?

Yes. The `log_metrics()` method accepts arbitrary keyword arguments that map to table columns. The schema is extensible, and backwards-compatible migrations add new columns lazily via `ALTER TABLE` when newer CLI versions encounter older databases.