How Unsloth Monitors Training Progress: Loss, GPU Usage, and Real-Time Metrics

Unsloth monitors training progress through a custom TrainerCallback that populates a TrainingProgress dataclass, forwards metrics via an inter-process event queue, and queries GPU statistics through CUDA/MLX utilities and nvidia-smi.

The unslothai/unsloth repository implements a sophisticated observability layer that tracks fine-tuning metrics while isolating heavy PyTorch workloads in subprocesses. This architecture enables real-time dashboard updates without blocking the training loop, providing accurate loss curves, learning rates, and hardware utilization across CUDA and Apple Silicon backends.

The TrainingProgress Dataclass and Callback Architecture

At the core of Unsloth’s monitoring system is the TrainingProgress dataclass defined in studio/backend/core/training/trainer.py. This immutable-like structure holds step counters, epoch numbers, loss values, learning rates, elapsed time, ETA estimates, gradient norms, token counts, evaluation loss, and human-readable status messages.

The UnslothTrainer class maintains a thread-safe list of progress callbacks and exposes two primary methods for metric handling:

  • add_progress_callback(callback) – Registers user-supplied callables that receive TrainingProgress instances
  • _update_progress(**kwargs) – Atomically updates the dataclass fields under a lock and notifies all registered listeners

# From studio/backend/core/training/trainer.py

def add_progress_callback(self, callback):
    """Register a listener that receives TrainingProgress updates."""
    self.progress_callbacks.append(callback)

def _update_progress(self, **kwargs):
    """Merge new values into the dataclass and invoke all callbacks."""
    with self._lock:
        for key, value in kwargs.items():
            if hasattr(self.training_progress, key):
                setattr(self.training_progress, key, value)

        for cb in self.progress_callbacks:
            try:
                cb(self.training_progress)
            except Exception as e:
                logger.error(f"Error in progress callback: {e}")

Bridging Trainer Callbacks with Transformers

Unsloth extends the Hugging Face TrainerCallback to intercept training events without modifying the underlying TRL SFTTrainer. The _create_progress_callback() method instantiates an inner class that overrides on_log, extracting metrics from the trainer's state and logs dictionaries.

When the Transformers trainer invokes on_log, the callback computes elapsed time, derives ETA from remaining steps, and pushes all values through _update_progress():

def _create_progress_callback(self):
    from transformers import TrainerCallback
    trainer_ref = self

    class _ProgressCallback(TrainerCallback):
        def on_log(self, args, state, control, logs=None, **_):
            if not logs:
                return
            loss = logs.get("loss", logs.get("train_loss", 0.0))
            step = state.global_step
            elapsed = (
                time.time() - trainer_ref.training_start_time
                if trainer_ref.training_start_time else None
            )
            # ETA computation omitted for brevity

            trainer_ref._update_progress(
                step=step,
                epoch=round(state.epoch, 2) if state.epoch else 0,
                loss=loss,
                learning_rate=logs.get("learning_rate", 0.0),
                elapsed_seconds=elapsed,
                grad_norm=logs.get("grad_norm", None),
                num_tokens=logs.get("num_tokens", 0),
            )
    return _ProgressCallback()

Inter-Process Communication for Real-Time UI Updates

To prevent the UI from freezing during intensive GPU workloads, Unsloth runs training inside a subprocess orchestrated by studio/backend/core/training/worker.py. The worker creates an _on_progress closure that receives TrainingProgress objects and serializes them into JSON-like events for the parent process.

Each progress update triggers event_queue.put(), enabling the Studio frontend to consume metrics asynchronously:


# From studio/backend/core/training/worker.py

def _on_progress(progress: TrainingProgress):
    event_queue.put({
        "type": "progress",
        "step": progress.step,
        "epoch": progress.epoch,
        "loss": progress.loss,
        "learning_rate": progress.learning_rate,
        "total_steps": progress.total_steps,
        "elapsed_seconds": progress.elapsed_seconds,
        "eta_seconds": progress.eta_seconds,
        "grad_norm": progress.grad_norm,
        "num_tokens": progress.num_tokens,
        "eval_loss": progress.eval_loss,
        "status_message": progress.status_message,
        "ts": time.time(),
    })
    if progress.status_message:
        _send_status(event_queue, progress.status_message)

trainer.add_progress_callback(_on_progress)

GPU Monitoring: Memory, Utilization, and Temperature

Unsloth tracks hardware metrics through studio/backend/utils/hardware/hardware.py, which abstracts CUDA, MLX, and CPU backends. Two primary functions expose GPU state:

  • get_gpu_memory_info() – Returns total, allocated, reserved, and free VRAM plus utilization percentage and device name
  • log_gpu_memory(context) – Formats memory statistics into structured log lines for debugging

For real-time utilization, get_gpu_utilization() shells out to nvidia-smi (when available) to capture GPU utilization percentage, temperature, power draw, and VRAM usage:

from utils.hardware import log_gpu_memory, get_gpu_utilization

# Log memory at specific lifecycle points

log_gpu_memory("After model load")

# Output: "GPU Memory [After model load] CUDA (NVIDIA L4): 2.34GB/22.17GB (10.5% used, 19.83GB free)"

# Query live statistics

stats = get_gpu_utilization()
if stats["available"]:
    print(f"GPU {stats['gpu_utilization_pct']:.1f}% | Temp {stats['temperature_c']}°C "
          f"| VRAM {stats['vram_used_gb']:.2f}/{stats['vram_total_gb']:.2f}GiB "
          f"| Power {stats['power_draw_w']:.1f}/{stats['power_limit_w']:.1f}W")

End-to-End Training Flow

During a typical fine-tuning run, Unsloth chains these components into a cohesive pipeline:

  1. Worker initialization – The subprocess starts and logs initial GPU memory via log_gpu_memory("Start training")
  2. Model loading – Status messages update through _update_progress(status_message="Loading...")
  3. Dataset preparation – An optional TQDM monitor thread reports tokenization percentages
  4. Training loop – After each logging step, the TrainerCallback populates TrainingProgress
  5. Event propagation – _on_progress forwards the dataclass to the parent via event_queue.put()
  6. Dashboard rendering – The Studio backend (studio/backend/main.py) consumes events and queries get_gpu_utilization() to render loss curves alongside GPU temperature and VRAM charts

This design isolates fault-prone ML code while guaranteeing low-latency metric delivery to the user interface.

Summary

  • Unsloth tracks training metrics through a TrainingProgress dataclass updated atomically via _update_progress() in studio/backend/core/training/trainer.py
  • Callback registration allows both internal UI forwarding and user-defined hooks via add_progress_callback()
  • Subprocess isolation prevents UI blocking; the worker in studio/backend/core/training/worker.py pushes JSON events through a multiprocessing queue
  • GPU monitoring combines CUDA/MLX API calls in get_gpu_memory_info() with nvidia-smi polling via get_gpu_utilization() in studio/backend/utils/hardware/hardware.py
  • Real-time updates are consumed by the Studio frontend to display loss, learning rate, ETA, and hardware statistics without interrupting the training loop

Frequently Asked Questions

How does Unsloth handle training progress callbacks in multi-GPU setups?

Unsloth’s callback architecture remains process-local; each training subprocess maintains its own UnslothTrainer instance with isolated TrainingProgress state. The callback list in add_progress_callback() is not shared across processes, ensuring thread-safe updates within a single training worker. For multi-GPU configurations, each GPU’s worker reports independently to the parent process, which aggregates metrics in the Studio dashboard.

Can I register custom progress callbacks when using Unsloth programmatically?

Yes. The UnslothTrainer exposes add_progress_callback(callback) where callback is any callable accepting a TrainingProgress dataclass. Your function receives updates containing step, epoch, loss, learning rate, and GPU stats after every logging interval. This integrates with existing MLOps pipelines without modifying the trainer internals.

What GPU metrics does Unsloth expose beyond VRAM usage?

According to the hardware utilities in studio/backend/utils/hardware/hardware.py, Unsloth reports GPU utilization percentage, temperature in Celsius, power draw in watts, and power limits when running on CUDA devices. On Apple Silicon (MLX), it tracks allocated and total memory. The get_gpu_utilization() function specifically queries nvidia-smi to return gpu_utilization_pct, temperature_c, power_draw_w, and vram_used_gb alongside total capacity.

How does the event queue prevent training blocking when the UI is slow?

The training worker runs in a separate Python process spawned by studio/backend/core/training/worker.py. Progress callbacks execute _on_progress(), which performs a non-blocking event_queue.put() to a multiprocessing queue. If the parent process (UI) is busy, the queue buffers events without stalling the PyTorch training loop. The parent consumes these asynchronously in studio/backend/main.py, decoupling metric reporting from gradient computation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →