How Soup CLI Tracks Experiments and Metrics: A Deep Dive into the SQLite-Based ExperimentTracker

Soup CLI tracks experiments and metrics using a local SQLite database (experiments.db) managed by the ExperimentTracker class, which stores run metadata, per-step training metrics, evaluation results, and cost estimates in the user's .soup directory.

The Soup CLI (MakazhanAlpamys/Soup) implements a lightweight yet comprehensive experiment tracking system that requires no external servers. At its core lies the ExperimentTracker class in [src/soup_cli/experiment/tracker.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/experiment/tracker.py), which orchestrates all experiment persistence through a local SQLite database. This article explains exactly how Soup CLI tracks experiments and metrics, from database schema to run lifecycle management.

The SQLite Database Architecture

All experiment data lives in $HOME/.soup/experiments.db. When ExperimentTracker initializes, it executes _ensure_schema() to create tables if absent.

Core Tables

Table Purpose
runs Run metadata: run_id, experiment_name, timestamps, status, device info, cost estimates, JSON config
metrics Per-step training data: step, epoch, loss, lr, grad_norm, speed, gpu_mem, timestamp
eval_results Benchmark scores with provenance tracking for reproducible comparisons
checkpoint_quality / forgetting_eval Extended analysis for model quality and catastrophic forgetting

The schema definition resides in the _SCHEMA_SQL constant. For backwards compatibility, new columns like cost_usd and run_kind are added lazily via ALTER TABLE calls.

Run Lifecycle Methods

The tracker implements a state machine for run execution. Here's how each method progresses a run from start to finish:

  • generate_run_id() — Creates sortable unique IDs like run_20240930_154523_a1b2c3d4
  • start_run() — Inserts a running row with device, GPU memory, base model, and task metadata
  • launch_run() — For async launches: stores launching status before child process starts
  • mark_running() — Updates to running and records the child PID
  • finish_execution() — Writes terminal status (completed/failed/terminated) without overwriting training summaries
  • finish_run() — Marks completed, writes summary fields, and computes cost via soup_cli.utils.run_cost
  • fail_run() — Directly flags failures
  • expunge_stale_launching_runs() — Cleans timed-out launching rows if PID is dead
  • delete_run() — Removes a run and all associated metrics/eval data

Orphaned Run Recovery

If a process crashes, _reconcile_orphaned_run() detects dead PIDs and rewrites status to terminated with unknown exit code. This prevents zombie entries from appearing successful.

Recording Training Metrics

During training, any trainer (PPO, SFT, or custom) calls log_metrics():

tracker.log_metrics(
    run_id,
    step=global_step,
    epoch=current_epoch,
    loss=loss_value,
    lr=learning_rate,
    grad_norm=gradient_norm,
    speed=tokens_per_sec,
    gpu_mem="38GB",
)

Each call writes one row to the metrics table with an automatic timestamp. Retrieval methods include:

  • get_metrics(run_id) — Returns all metric rows for a run
  • get_metric_series(run_id, metric) — Extracts float list for plotting; falls back to _eval_score_series if no step data exists

Example: Basic Training Script

from soup_cli.experiment.tracker import ExperimentTracker

# Initialize tracker (creates DB if missing)

tracker = ExperimentTracker()

# Start run

config = {"base": "meta-llama/Llama-2-7b", "task": "sft"}
run_id = tracker.start_run(
    config,
    device="cuda",
    device_name="NVIDIA A100",
    gpu_info={"memory_total": "40GB"}
)

# Log 5 training steps

for step in range(1, 6):
    tracker.log_metrics(
        run_id,
        step=step,
        epoch=step / 5,
        loss=2.5 / step,
        lr=1e-5,
        grad_norm=0.8,
        speed=120.0,
        gpu_mem="38GB",
    )

# Finalize with summary statistics

tracker.finish_run(
    run_id,
    initial_loss=2.5,
    final_loss=0.5,
    total_steps=5,
    duration_secs=300,
    output_dir="/home/user/.soup/runs/run_20240930_154523_a1b2c3d4",
)

# Retrieve for visualization

loss_series = tracker.get_metric_series(run_id, "loss")
print("Loss per step:", loss_series)

tracker.close()

Evaluation and Benchmark Tracking

After running evaluations, save_eval_result() persists scores with provenance:

tracker.save_eval_result(
    model_path="/home/user/.soup/models/finetuned.pt",
    benchmark="hellaswag",
    score=0.842,
    details={"provenance": "baseline-v1"},
    run_id=run_id,
)

This stamps a current_baseline_stamp enabling reproducible comparisons across runs.

Cost Estimation Integration

During finish_run(), the tracker computes:

  1. GPU hourly rate via lookup_gpu_rate()
  2. Estimated USD cost via estimate_run_cost_usd()

These populate cost_usd and cost_gpu_label columns for budget reporting.

CLI Integration Points

The tracker serves all user-facing commands:

Command Usage
soup train start_run() → log_metrics() loop → finish_run()
soup eval save_eval_result() for benchmark scores
soup runs Queries get_metrics() and run metadata for tabular display
soup monitor Live metric forwarding via monitoring/callback.py
soup ui Pulls data for interactive graphs and tables

Key Source Files

File Role
[src/soup_cli/experiment/tracker.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/experiment/tracker.py) Core ExperimentTracker class with all lifecycle and logging methods
[src/soup_cli/utils/constants.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/constants.py) Defines EXPERIMENTS_DB and SOUP_DIR paths
[src/soup_cli/trainer/ppo.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/ppo.py) PPO trainer calling log_metrics() each step
[src/soup_cli/trainer/mlx_sft.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/mlx_sft.py) SFT trainer implementation with metric logging
[src/soup_cli/monitoring/callback.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/monitoring/callback.py) Live monitor callback for UI metric streaming
[src/soup_cli/commands/runs.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/runs.py) CLI display of stored runs and metrics
[src/soup_cli/commands/ui.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/ui.py) UI entry point querying tracker for visualization

Summary

  • Soup CLI tracks experiments through a local SQLite database at $HOME/.soup/experiments.db, eliminating external dependencies
  • The ExperimentTracker class in src/soup_cli/experiment/tracker.py provides complete lifecycle management: start_run() → log_metrics() → finish_run()
  • Four main tables organize data: runs, metrics, eval_results, and extended quality tables
  • Automatic cost estimation and orphaned run reconciliation ensure data integrity and budget visibility
  • All CLI commands (train, eval, runs, monitor, ui) share this single tracking interface

Frequently Asked Questions

Where does Soup CLI store experiment data?

Soup CLI stores all data in $HOME/.soup/experiments.db, a local SQLite database. The path is defined in src/soup_cli/utils/constants.py via the EXPERIMENTS_DB constant. This local-first approach requires no cloud setup or external servers.

How do I retrieve metrics from a completed run?

Call tracker.get_metrics(run_id) for full row data or tracker.get_metric_series(run_id, "loss") for a list of float values suitable for plotting. If no per-step metrics exist, the method automatically falls back to evaluation scores from the eval_results table.

What happens if a training process crashes unexpectedly?

The ExperimentTracker detects orphaned runs through _reconcile_orphaned_run(), which checks if the recorded PID is still alive. Dead processes trigger status rewriting to terminated with unknown exit code, preventing stale entries from appearing successful in soup runs listings.

Can I track custom metrics beyond loss and learning rate?

Yes. The log_metrics() method accepts arbitrary keyword arguments that map to table columns. The schema is extensible, and backwards-compatible migrations add new columns lazily via ALTER TABLE when newer CLI versions encounter older databases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →