# How Unsloth Monitors Training Progress: Loss, GPU Usage, and Real-Time Metrics

> Unsloth monitors training progress with real-time metrics like loss and GPU usage. Discover how our custom TrainerCallback tracks performance and optimizes your models.

- Repository: [Unsloth AI/unsloth](https://github.com/unslothai/unsloth)
- Tags: internals
- Published: 2026-03-20

---

**Unsloth monitors training progress through a custom `TrainerCallback` that populates a `TrainingProgress` dataclass, forwards metrics via an inter-process event queue, and queries GPU statistics through CUDA/MLX utilities and nvidia-smi.**

The `unslothai/unsloth` repository implements a sophisticated observability layer that tracks fine-tuning metrics while isolating heavy PyTorch workloads in subprocesses. This architecture enables real-time dashboard updates without blocking the training loop, providing accurate loss curves, learning rates, and hardware utilization across CUDA and Apple Silicon backends.

## The TrainingProgress Dataclass and Callback Architecture

At the core of Unsloth’s monitoring system is the `TrainingProgress` dataclass defined in [`studio/backend/core/training/trainer.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/core/training/trainer.py). This immutable-like structure holds step counters, epoch numbers, loss values, learning rates, elapsed time, ETA estimates, gradient norms, token counts, evaluation loss, and human-readable status messages.

The `UnslothTrainer` class maintains a thread-safe list of progress callbacks and exposes two primary methods for metric handling:

- **`add_progress_callback(callback)`** – Registers user-supplied callables that receive `TrainingProgress` instances
- **`_update_progress(**kwargs)`** – Atomically updates the dataclass fields under a lock and notifies all registered listeners

```python

# From studio/backend/core/training/trainer.py

def add_progress_callback(self, callback):
    """Register a listener that receives TrainingProgress updates."""
    self.progress_callbacks.append(callback)

def _update_progress(self, **kwargs):
    """Merge new values into the dataclass and invoke all callbacks."""
    with self._lock:
        for key, value in kwargs.items():
            if hasattr(self.training_progress, key):
                setattr(self.training_progress, key, value)

        for cb in self.progress_callbacks:
            try:
                cb(self.training_progress)
            except Exception as e:
                logger.error(f"Error in progress callback: {e}")

```

## Bridging Trainer Callbacks with Transformers

Unsloth extends the Hugging Face `TrainerCallback` to intercept training events without modifying the underlying TRL `SFTTrainer`. The `_create_progress_callback()` method instantiates an inner class that overrides `on_log`, extracting metrics from the trainer's `state` and `logs` dictionaries.

When the Transformers trainer invokes `on_log`, the callback computes elapsed time, derives ETA from remaining steps, and pushes all values through `_update_progress()`:

```python
def _create_progress_callback(self):
    from transformers import TrainerCallback
    trainer_ref = self

    class _ProgressCallback(TrainerCallback):
        def on_log(self, args, state, control, logs=None, **_):
            if not logs:
                return
            loss = logs.get("loss", logs.get("train_loss", 0.0))
            step = state.global_step
            elapsed = (
                time.time() - trainer_ref.training_start_time
                if trainer_ref.training_start_time else None
            )
            # ETA computation omitted for brevity

            trainer_ref._update_progress(
                step=step,
                epoch=round(state.epoch, 2) if state.epoch else 0,
                loss=loss,
                learning_rate=logs.get("learning_rate", 0.0),
                elapsed_seconds=elapsed,
                grad_norm=logs.get("grad_norm", None),
                num_tokens=logs.get("num_tokens", 0),
            )
    return _ProgressCallback()

```

## Inter-Process Communication for Real-Time UI Updates

To prevent the UI from freezing during intensive GPU workloads, Unsloth runs training inside a subprocess orchestrated by [`studio/backend/core/training/worker.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/core/training/worker.py). The worker creates an `_on_progress` closure that receives `TrainingProgress` objects and serializes them into JSON-like events for the parent process.

Each progress update triggers `event_queue.put()`, enabling the Studio frontend to consume metrics asynchronously:

```python

# From studio/backend/core/training/worker.py

def _on_progress(progress: TrainingProgress):
    event_queue.put({
        "type": "progress",
        "step": progress.step,
        "epoch": progress.epoch,
        "loss": progress.loss,
        "learning_rate": progress.learning_rate,
        "total_steps": progress.total_steps,
        "elapsed_seconds": progress.elapsed_seconds,
        "eta_seconds": progress.eta_seconds,
        "grad_norm": progress.grad_norm,
        "num_tokens": progress.num_tokens,
        "eval_loss": progress.eval_loss,
        "status_message": progress.status_message,
        "ts": time.time(),
    })
    if progress.status_message:
        _send_status(event_queue, progress.status_message)

trainer.add_progress_callback(_on_progress)

```

## GPU Monitoring: Memory, Utilization, and Temperature

Unsloth tracks hardware metrics through [`studio/backend/utils/hardware/hardware.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/utils/hardware/hardware.py), which abstracts CUDA, MLX, and CPU backends. Two primary functions expose GPU state:

- **`get_gpu_memory_info()`** – Returns total, allocated, reserved, and free VRAM plus utilization percentage and device name
- **`log_gpu_memory(context)`** – Formats memory statistics into structured log lines for debugging

For real-time utilization, `get_gpu_utilization()` shells out to `nvidia-smi` (when available) to capture GPU utilization percentage, temperature, power draw, and VRAM usage:

```python
from utils.hardware import log_gpu_memory, get_gpu_utilization

# Log memory at specific lifecycle points

log_gpu_memory("After model load")

# Output: "GPU Memory [After model load] CUDA (NVIDIA L4): 2.34GB/22.17GB (10.5% used, 19.83GB free)"

# Query live statistics

stats = get_gpu_utilization()
if stats["available"]:
    print(f"GPU {stats['gpu_utilization_pct']:.1f}% | Temp {stats['temperature_c']}°C "
          f"| VRAM {stats['vram_used_gb']:.2f}/{stats['vram_total_gb']:.2f}GiB "
          f"| Power {stats['power_draw_w']:.1f}/{stats['power_limit_w']:.1f}W")

```

## End-to-End Training Flow

During a typical fine-tuning run, Unsloth chains these components into a cohesive pipeline:

1. **Worker initialization** – The subprocess starts and logs initial GPU memory via `log_gpu_memory("Start training")`
2. **Model loading** – Status messages update through `_update_progress(status_message="Loading...")`
3. **Dataset preparation** – An optional TQDM monitor thread reports tokenization percentages
4. **Training loop** – After each logging step, the `TrainerCallback` populates `TrainingProgress`
5. **Event propagation** – `_on_progress` forwards the dataclass to the parent via `event_queue.put()`
6. **Dashboard rendering** – The Studio backend ([`studio/backend/main.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/main.py)) consumes events and queries `get_gpu_utilization()` to render loss curves alongside GPU temperature and VRAM charts

This design isolates fault-prone ML code while guaranteeing low-latency metric delivery to the user interface.

## Summary

- **Unsloth tracks training metrics** through a `TrainingProgress` dataclass updated atomically via `_update_progress()` in [`studio/backend/core/training/trainer.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/core/training/trainer.py)
- **Callback registration** allows both internal UI forwarding and user-defined hooks via `add_progress_callback()`
- **Subprocess isolation** prevents UI blocking; the worker in [`studio/backend/core/training/worker.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/core/training/worker.py) pushes JSON events through a multiprocessing queue
- **GPU monitoring** combines CUDA/MLX API calls in `get_gpu_memory_info()` with `nvidia-smi` polling via `get_gpu_utilization()` in [`studio/backend/utils/hardware/hardware.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/utils/hardware/hardware.py)
- **Real-time updates** are consumed by the Studio frontend to display loss, learning rate, ETA, and hardware statistics without interrupting the training loop

## Frequently Asked Questions

### How does Unsloth handle training progress callbacks in multi-GPU setups?

Unsloth’s callback architecture remains process-local; each training subprocess maintains its own `UnslothTrainer` instance with isolated `TrainingProgress` state. The callback list in `add_progress_callback()` is not shared across processes, ensuring thread-safe updates within a single training worker. For multi-GPU configurations, each GPU’s worker reports independently to the parent process, which aggregates metrics in the Studio dashboard.

### Can I register custom progress callbacks when using Unsloth programmatically?

Yes. The `UnslothTrainer` exposes `add_progress_callback(callback)` where `callback` is any callable accepting a `TrainingProgress` dataclass. Your function receives updates containing step, epoch, loss, learning rate, and GPU stats after every logging interval. This integrates with existing MLOps pipelines without modifying the trainer internals.

### What GPU metrics does Unsloth expose beyond VRAM usage?

According to the hardware utilities in [`studio/backend/utils/hardware/hardware.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/utils/hardware/hardware.py), Unsloth reports GPU utilization percentage, temperature in Celsius, power draw in watts, and power limits when running on CUDA devices. On Apple Silicon (MLX), it tracks allocated and total memory. The `get_gpu_utilization()` function specifically queries `nvidia-smi` to return `gpu_utilization_pct`, `temperature_c`, `power_draw_w`, and `vram_used_gb` alongside total capacity.

### How does the event queue prevent training blocking when the UI is slow?

The training worker runs in a separate Python process spawned by [`studio/backend/core/training/worker.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/core/training/worker.py). Progress callbacks execute `_on_progress()`, which performs a non-blocking `event_queue.put()` to a multiprocessing queue. If the parent process (UI) is busy, the queue buffers events without stalling the PyTorch training loop. The parent consumes these asynchronously in [`studio/backend/main.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/main.py), decoupling metric reporting from gradient computation.