How Unsloth Monitors Training Progress: Loss, GPU Usage, and Real-Time Metrics
Unsloth monitors training progress through a custom TrainerCallback that populates a TrainingProgress dataclass, forwards metrics via an inter-process event queue, and queries GPU statistics through CUDA/MLX utilities and nvidia-smi.
The unslothai/unsloth repository implements a sophisticated observability layer that tracks fine-tuning metrics while isolating heavy PyTorch workloads in subprocesses. This architecture enables real-time dashboard updates without blocking the training loop, providing accurate loss curves, learning rates, and hardware utilization across CUDA and Apple Silicon backends.
The TrainingProgress Dataclass and Callback Architecture
At the core of Unsloth’s monitoring system is the TrainingProgress dataclass defined in studio/backend/core/training/trainer.py. This immutable-like structure holds step counters, epoch numbers, loss values, learning rates, elapsed time, ETA estimates, gradient norms, token counts, evaluation loss, and human-readable status messages.
The UnslothTrainer class maintains a thread-safe list of progress callbacks and exposes two primary methods for metric handling:
add_progress_callback(callback)– Registers user-supplied callables that receiveTrainingProgressinstances_update_progress(**kwargs)– Atomically updates the dataclass fields under a lock and notifies all registered listeners
# From studio/backend/core/training/trainer.py
def add_progress_callback(self, callback):
"""Register a listener that receives TrainingProgress updates."""
self.progress_callbacks.append(callback)
def _update_progress(self, **kwargs):
"""Merge new values into the dataclass and invoke all callbacks."""
with self._lock:
for key, value in kwargs.items():
if hasattr(self.training_progress, key):
setattr(self.training_progress, key, value)
for cb in self.progress_callbacks:
try:
cb(self.training_progress)
except Exception as e:
logger.error(f"Error in progress callback: {e}")
Bridging Trainer Callbacks with Transformers
Unsloth extends the Hugging Face TrainerCallback to intercept training events without modifying the underlying TRL SFTTrainer. The _create_progress_callback() method instantiates an inner class that overrides on_log, extracting metrics from the trainer's state and logs dictionaries.
When the Transformers trainer invokes on_log, the callback computes elapsed time, derives ETA from remaining steps, and pushes all values through _update_progress():
def _create_progress_callback(self):
from transformers import TrainerCallback
trainer_ref = self
class _ProgressCallback(TrainerCallback):
def on_log(self, args, state, control, logs=None, **_):
if not logs:
return
loss = logs.get("loss", logs.get("train_loss", 0.0))
step = state.global_step
elapsed = (
time.time() - trainer_ref.training_start_time
if trainer_ref.training_start_time else None
)
# ETA computation omitted for brevity
trainer_ref._update_progress(
step=step,
epoch=round(state.epoch, 2) if state.epoch else 0,
loss=loss,
learning_rate=logs.get("learning_rate", 0.0),
elapsed_seconds=elapsed,
grad_norm=logs.get("grad_norm", None),
num_tokens=logs.get("num_tokens", 0),
)
return _ProgressCallback()
Inter-Process Communication for Real-Time UI Updates
To prevent the UI from freezing during intensive GPU workloads, Unsloth runs training inside a subprocess orchestrated by studio/backend/core/training/worker.py. The worker creates an _on_progress closure that receives TrainingProgress objects and serializes them into JSON-like events for the parent process.
Each progress update triggers event_queue.put(), enabling the Studio frontend to consume metrics asynchronously:
# From studio/backend/core/training/worker.py
def _on_progress(progress: TrainingProgress):
event_queue.put({
"type": "progress",
"step": progress.step,
"epoch": progress.epoch,
"loss": progress.loss,
"learning_rate": progress.learning_rate,
"total_steps": progress.total_steps,
"elapsed_seconds": progress.elapsed_seconds,
"eta_seconds": progress.eta_seconds,
"grad_norm": progress.grad_norm,
"num_tokens": progress.num_tokens,
"eval_loss": progress.eval_loss,
"status_message": progress.status_message,
"ts": time.time(),
})
if progress.status_message:
_send_status(event_queue, progress.status_message)
trainer.add_progress_callback(_on_progress)
GPU Monitoring: Memory, Utilization, and Temperature
Unsloth tracks hardware metrics through studio/backend/utils/hardware/hardware.py, which abstracts CUDA, MLX, and CPU backends. Two primary functions expose GPU state:
get_gpu_memory_info()– Returns total, allocated, reserved, and free VRAM plus utilization percentage and device namelog_gpu_memory(context)– Formats memory statistics into structured log lines for debugging
For real-time utilization, get_gpu_utilization() shells out to nvidia-smi (when available) to capture GPU utilization percentage, temperature, power draw, and VRAM usage:
from utils.hardware import log_gpu_memory, get_gpu_utilization
# Log memory at specific lifecycle points
log_gpu_memory("After model load")
# Output: "GPU Memory [After model load] CUDA (NVIDIA L4): 2.34GB/22.17GB (10.5% used, 19.83GB free)"
# Query live statistics
stats = get_gpu_utilization()
if stats["available"]:
print(f"GPU {stats['gpu_utilization_pct']:.1f}% | Temp {stats['temperature_c']}°C "
f"| VRAM {stats['vram_used_gb']:.2f}/{stats['vram_total_gb']:.2f}GiB "
f"| Power {stats['power_draw_w']:.1f}/{stats['power_limit_w']:.1f}W")
End-to-End Training Flow
During a typical fine-tuning run, Unsloth chains these components into a cohesive pipeline:
- Worker initialization – The subprocess starts and logs initial GPU memory via
log_gpu_memory("Start training") - Model loading – Status messages update through
_update_progress(status_message="Loading...") - Dataset preparation – An optional TQDM monitor thread reports tokenization percentages
- Training loop – After each logging step, the
TrainerCallbackpopulatesTrainingProgress - Event propagation –
_on_progressforwards the dataclass to the parent viaevent_queue.put() - Dashboard rendering – The Studio backend (
studio/backend/main.py) consumes events and queriesget_gpu_utilization()to render loss curves alongside GPU temperature and VRAM charts
This design isolates fault-prone ML code while guaranteeing low-latency metric delivery to the user interface.
Summary
- Unsloth tracks training metrics through a
TrainingProgressdataclass updated atomically via_update_progress()instudio/backend/core/training/trainer.py - Callback registration allows both internal UI forwarding and user-defined hooks via
add_progress_callback() - Subprocess isolation prevents UI blocking; the worker in
studio/backend/core/training/worker.pypushes JSON events through a multiprocessing queue - GPU monitoring combines CUDA/MLX API calls in
get_gpu_memory_info()withnvidia-smipolling viaget_gpu_utilization()instudio/backend/utils/hardware/hardware.py - Real-time updates are consumed by the Studio frontend to display loss, learning rate, ETA, and hardware statistics without interrupting the training loop
Frequently Asked Questions
How does Unsloth handle training progress callbacks in multi-GPU setups?
Unsloth’s callback architecture remains process-local; each training subprocess maintains its own UnslothTrainer instance with isolated TrainingProgress state. The callback list in add_progress_callback() is not shared across processes, ensuring thread-safe updates within a single training worker. For multi-GPU configurations, each GPU’s worker reports independently to the parent process, which aggregates metrics in the Studio dashboard.
Can I register custom progress callbacks when using Unsloth programmatically?
Yes. The UnslothTrainer exposes add_progress_callback(callback) where callback is any callable accepting a TrainingProgress dataclass. Your function receives updates containing step, epoch, loss, learning rate, and GPU stats after every logging interval. This integrates with existing MLOps pipelines without modifying the trainer internals.
What GPU metrics does Unsloth expose beyond VRAM usage?
According to the hardware utilities in studio/backend/utils/hardware/hardware.py, Unsloth reports GPU utilization percentage, temperature in Celsius, power draw in watts, and power limits when running on CUDA devices. On Apple Silicon (MLX), it tracks allocated and total memory. The get_gpu_utilization() function specifically queries nvidia-smi to return gpu_utilization_pct, temperature_c, power_draw_w, and vram_used_gb alongside total capacity.
How does the event queue prevent training blocking when the UI is slow?
The training worker runs in a separate Python process spawned by studio/backend/core/training/worker.py. Progress callbacks execute _on_progress(), which performs a non-blocking event_queue.put() to a multiprocessing queue. If the parent process (UI) is busy, the queue buffers events without stalling the PyTorch training loop. The parent consumes these asynchronously in studio/backend/main.py, decoupling metric reporting from gradient computation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →