How Soup CLI Tracks Experiments and Metrics: A Deep Dive into the SQLite-Based ExperimentTracker
Soup CLI tracks experiments and metrics using a local SQLite database (experiments.db) managed by the ExperimentTracker class, which stores run metadata, per-step training metrics, evaluation results, and cost estimates in the user's .soup directory.
The Soup CLI (MakazhanAlpamys/Soup) implements a lightweight yet comprehensive experiment tracking system that requires no external servers. At its core lies the ExperimentTracker class in [src/soup_cli/experiment/tracker.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/experiment/tracker.py), which orchestrates all experiment persistence through a local SQLite database. This article explains exactly how Soup CLI tracks experiments and metrics, from database schema to run lifecycle management.
The SQLite Database Architecture
All experiment data lives in $HOME/.soup/experiments.db. When ExperimentTracker initializes, it executes _ensure_schema() to create tables if absent.
Core Tables
| Table | Purpose |
|---|---|
| runs | Run metadata: run_id, experiment_name, timestamps, status, device info, cost estimates, JSON config |
| metrics | Per-step training data: step, epoch, loss, lr, grad_norm, speed, gpu_mem, timestamp |
| eval_results | Benchmark scores with provenance tracking for reproducible comparisons |
| checkpoint_quality / forgetting_eval | Extended analysis for model quality and catastrophic forgetting |
The schema definition resides in the _SCHEMA_SQL constant. For backwards compatibility, new columns like cost_usd and run_kind are added lazily via ALTER TABLE calls.
Run Lifecycle Methods
The tracker implements a state machine for run execution. Here's how each method progresses a run from start to finish:
generate_run_id()— Creates sortable unique IDs likerun_20240930_154523_a1b2c3d4start_run()— Inserts a running row with device, GPU memory, base model, and task metadatalaunch_run()— For async launches: stores launching status before child process startsmark_running()— Updates to running and records the child PIDfinish_execution()— Writes terminal status (completed/failed/terminated) without overwriting training summariesfinish_run()— Marks completed, writes summary fields, and computes cost viasoup_cli.utils.run_costfail_run()— Directly flags failuresexpunge_stale_launching_runs()— Cleans timed-out launching rows if PID is deaddelete_run()— Removes a run and all associated metrics/eval data
Orphaned Run Recovery
If a process crashes, _reconcile_orphaned_run() detects dead PIDs and rewrites status to terminated with unknown exit code. This prevents zombie entries from appearing successful.
Recording Training Metrics
During training, any trainer (PPO, SFT, or custom) calls log_metrics():
tracker.log_metrics(
run_id,
step=global_step,
epoch=current_epoch,
loss=loss_value,
lr=learning_rate,
grad_norm=gradient_norm,
speed=tokens_per_sec,
gpu_mem="38GB",
)
Each call writes one row to the metrics table with an automatic timestamp. Retrieval methods include:
get_metrics(run_id)— Returns all metric rows for a runget_metric_series(run_id, metric)— Extracts float list for plotting; falls back to_eval_score_seriesif no step data exists
Example: Basic Training Script
from soup_cli.experiment.tracker import ExperimentTracker
# Initialize tracker (creates DB if missing)
tracker = ExperimentTracker()
# Start run
config = {"base": "meta-llama/Llama-2-7b", "task": "sft"}
run_id = tracker.start_run(
config,
device="cuda",
device_name="NVIDIA A100",
gpu_info={"memory_total": "40GB"}
)
# Log 5 training steps
for step in range(1, 6):
tracker.log_metrics(
run_id,
step=step,
epoch=step / 5,
loss=2.5 / step,
lr=1e-5,
grad_norm=0.8,
speed=120.0,
gpu_mem="38GB",
)
# Finalize with summary statistics
tracker.finish_run(
run_id,
initial_loss=2.5,
final_loss=0.5,
total_steps=5,
duration_secs=300,
output_dir="/home/user/.soup/runs/run_20240930_154523_a1b2c3d4",
)
# Retrieve for visualization
loss_series = tracker.get_metric_series(run_id, "loss")
print("Loss per step:", loss_series)
tracker.close()
Evaluation and Benchmark Tracking
After running evaluations, save_eval_result() persists scores with provenance:
tracker.save_eval_result(
model_path="/home/user/.soup/models/finetuned.pt",
benchmark="hellaswag",
score=0.842,
details={"provenance": "baseline-v1"},
run_id=run_id,
)
This stamps a current_baseline_stamp enabling reproducible comparisons across runs.
Cost Estimation Integration
During finish_run(), the tracker computes:
- GPU hourly rate via
lookup_gpu_rate() - Estimated USD cost via
estimate_run_cost_usd()
These populate cost_usd and cost_gpu_label columns for budget reporting.
CLI Integration Points
The tracker serves all user-facing commands:
| Command | Usage |
|---|---|
soup train |
start_run() → log_metrics() loop → finish_run() |
soup eval |
save_eval_result() for benchmark scores |
soup runs |
Queries get_metrics() and run metadata for tabular display |
soup monitor |
Live metric forwarding via monitoring/callback.py |
soup ui |
Pulls data for interactive graphs and tables |
Key Source Files
Summary
- Soup CLI tracks experiments through a local SQLite database at
$HOME/.soup/experiments.db, eliminating external dependencies - The
ExperimentTrackerclass insrc/soup_cli/experiment/tracker.pyprovides complete lifecycle management:start_run()→log_metrics()→finish_run() - Four main tables organize data: runs, metrics, eval_results, and extended quality tables
- Automatic cost estimation and orphaned run reconciliation ensure data integrity and budget visibility
- All CLI commands (
train,eval,runs,monitor,ui) share this single tracking interface
Frequently Asked Questions
Where does Soup CLI store experiment data?
Soup CLI stores all data in $HOME/.soup/experiments.db, a local SQLite database. The path is defined in src/soup_cli/utils/constants.py via the EXPERIMENTS_DB constant. This local-first approach requires no cloud setup or external servers.
How do I retrieve metrics from a completed run?
Call tracker.get_metrics(run_id) for full row data or tracker.get_metric_series(run_id, "loss") for a list of float values suitable for plotting. If no per-step metrics exist, the method automatically falls back to evaluation scores from the eval_results table.
What happens if a training process crashes unexpectedly?
The ExperimentTracker detects orphaned runs through _reconcile_orphaned_run(), which checks if the recorded PID is still alive. Dead processes trigger status rewriting to terminated with unknown exit code, preventing stale entries from appearing successful in soup runs listings.
Can I track custom metrics beyond loss and learning rate?
Yes. The log_metrics() method accepts arbitrary keyword arguments that map to table columns. The schema is extensible, and backwards-compatible migrations add new columns lazily via ALTER TABLE when newer CLI versions encounter older databases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →