# How to Use Streamlit for Controlling LLM Training: A Complete Guide to train-llm-from-scratch

> Control LLM training from scratch with an intuitive Streamlit interface. Manage data prep, pretraining, fine-tuning, and RLHF from your browser. Perfect for train-llm-from-scratch.

- Repository: [Fareed Khan/train-llm-from-scratch](https://github.com/FareedKhan-dev/train-llm-from-scratch)
- Tags: how-to-guide
- Published: 2026-06-11

---

**The FareedKhan-dev/train-llm-from-scratch repository provides a fully-featured Streamlit control panel that enables browser-based management of the entire LLM lifecycle—from data preparation and pretraining to supervised fine-tuning (SFT), reward modeling, and RLHF algorithms like DPO, PPO, and GRPO.**

The **FareedKhan-dev/train-llm-from-scratch** repository ships with a production-ready web interface that demonstrates how to use **Streamlit for controlling LLM training** workflows without complex command-line interactions. This browser-based dashboard allows you to configure hyperparameters, launch detached training jobs, and monitor GPU utilization through an intuitive UI. The architecture cleanly separates the presentation layer from the training logic, using background subprocesses that survive page reloads while streaming real-time logs back to the browser.

## Architecture of the Streamlit Control Panel

The UI is organized into four specialized modules that handle distinct responsibilities, from defining training stages to managing background processes.

### Central Stage Registry ([`ui/stages.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/stages.py))

The file [`ui/stages.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/stages.py) serves as the single source of truth for every pipeline stage. It defines a `Stage` dataclass that bundles everything a UI page needs: a configuration dataclass, JSON file locations, the training script path, log prefixes, documentation, and multi-GPU flags.

```python
STAGES: dict[str, Stage] = {
    "pretrain": Stage(
        key="pretrain",
        title="Pretraining",
        emoji="📚",
        cfg_cls=PretrainConfig,
        config_json="configs/pretrain.json",
        smoke_json="configs/smoke/pretrain.json",
        script="scripts/pretrain_base.py",
        log_prefix="pretrain",
        doc_md="docs/02_pretraining.md",
        diagram_png="docs/diagrams/02_pretraining.png",
        multi_gpu=True,
    ),
    # ... sft, reward, dpo, ppo, grpo ...

}

```

Each entry in the `STAGES` dictionary references a **dataclass** (e.g., `PretrainConfig` from [`config/post_training_config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/post_training_config.py)) that parses the JSON config, the **script** that executes the training, and metadata for documentation rendering.

### Background Job Manager ([`ui/jobs.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/jobs.py))

The [`ui/jobs.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/jobs.py) module handles the heavy lifting of **controlling LLM training** processes outside the Streamlit thread. The `build_argv` function constructs the command line, automatically injecting `torchrun` for multi-GPU scenarios:

```python
def build_argv(script, config_json, nproc, multi_gpu, extra=None):
    base = [script, "--config", config_json] + (extra or [])
    if multi_gpu and nproc > 1:
        return [_TORCHRUN, "--standalone", f"--nproc_per_node={nproc}", *base]
    return [sys.executable, *base]

```

The `launch` function writes a JSON registry to `/ephemeral/ui_jobs/` containing the `pid`, command, start time, and status, then spawns the process using `subprocess.Popen(..., start_new_session=True)`. This ensures training continues even if you close the browser.

### UI Styling and GPU Monitoring ([`ui/theme.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/theme.py))

The [`ui/theme.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/theme.py) file defines the visual layer through a `COLORS` dictionary and CSS blocks. It provides the `gpu_status()` helper that calls `nvidia-smi` to render real-time GPU metrics, and `status_badge()` for color-coded job state indicators.

## Runtime Workflow: From Config to Training

Understanding how **Streamlit controls the LLM training** lifecycle helps you debug and extend the interface.

1. **Configuration Editing**: Sidebar pages auto-generated from `STAGES` let users edit JSON configs via Streamlit forms.
2. **Launch Sequence**: When the **Launch** button is clicked, the page calls `jobs.build_argv` → `jobs.launch`, which builds the command line and spawns a detached process.
3. **GPU Guard**: The `gpu_busy()` function scans the JSON registry for any running job with `kind == "gpu"` and blocks new launches if a GPU job is active, preventing resource conflicts.
4. **Log Streaming**: The UI polls the registry via `jobs.status()` and tails the log file at `/ephemeral/ui_jobs/<job_id>.log`, updating badges and log views in real time.

## Extending the UI with Custom Training Stages

You can add new stages to the pipeline by modifying the central registry. This is the primary extension point for customizing how **Streamlit controls your LLM training** workflow.

1. **Create a Config Dataclass**: Add a new dataclass to [`config/post_training_config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/post_training_config.py) (e.g., `MyNewConfig`).
2. **Add JSON Templates**: Create [`configs/my_new.json`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/configs/my_new.json) and an optional smoke test version at [`configs/smoke/my_new.json`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/configs/smoke/my_new.json).
3. **Write the Training Script**: Implement the logic in [`scripts/train_my_new.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_my_new.py).
4. **Register the Stage**: Insert a new entry into `STAGES` in [`ui/stages.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/stages.py):

```python
STAGES["my_new"] = Stage(
    key="my_new",
    title="My New Stage",
    emoji="🚀",
    cfg_cls=MyNewConfig,
    config_json="configs/my_new.json",
    smoke_json="configs/smoke/my_new.json",
    script="scripts/train_my_new.py",
    log_prefix="my_new",
    doc_md="docs/09_my_new.md",
    diagram_png="docs/diagrams/09_my_new.png",
    multi_gpu=False,
)

```

5. **Documentation**: Add a markdown doc and diagram, then run `streamlit run ui/app.py`—the new page appears automatically in the sidebar.

## Practical Implementation Examples

### Launching a Training Stage Programmatically

This example demonstrates how the UI internally launches an SFT job using the `jobs` module:

```python
import streamlit as st
import time
from ui import jobs, stages

# Select the SFT stage configuration

stage = stages.STAGES["sft"]
cfg_path = stage.config_json          # "configs/sft.json"

script = stage.script                 # "scripts/train_sft.py"

nproc = 2                             # Number of GPUs

# Build the command-line arguments

argv = jobs.build_argv(
    script=script,
    config_json=cfg_path,
    nproc=nproc,
    multi_gpu=stage.multi_gpu,
)

# Generate unique job ID and launch

job_id = f"sft_{int(time.time())}"
rec = jobs.launch(job_id, argv, kind="gpu")
st.success(f"Launched {stage.title} with job ID **{job_id}**")

```

### Checking GPU Availability

Use the `gpu_status` helper to programmatically check resource availability before launching:

```python
from ui.theme import gpu_status

gpus = gpu_status()
if gpus:
    for i, name, used, total, util in gpus:
        st.metric(f"GPU {i}", f"{used}/{total} GB", f"{util}% util")
else:
    st.warning("No GPUs detected")

```

### Template for a New UI Page

Create [`ui/pages/my_new.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/pages/my_new.py) to expose your custom stage:

```python
import streamlit as st
import json
import pathlib
from ui import jobs, stages, theme

stage = stages.STAGES["my_new"]
st.title(f"{stage.emoji} {stage.title}")
st.image(stage.diagram_png)

# Config form

with st.form("cfg_form"):
    epochs = st.number_input("Epochs", min_value=1, value=5)
    lr = st.text_input("Learning rate", value="5e-4")
    submitted = st.form_submit_button("Save config")
    
    if submitted:
        cfg = {"epochs": epochs, "lr": float(lr)}
        pathlib.Path(stage.config_json).write_text(
            json.dumps(cfg, indent=2)
        )

# Launch button

if st.button("Launch Training"):
    argv = jobs.build_argv(
        stage.script, 
        stage.config_json, 
        nproc=1, 
        multi_gpu=False
    )
    job_id = f"my_new_{int(time.time())}"
    jobs.launch(job_id, argv, kind="cpu")
    st.success(f"Launched {stage.title}")

```

## Summary

- **Centralized Configuration**: The [`ui/stages.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/stages.py) file acts as the single source of truth, binding together config files, training scripts, and documentation for each pipeline stage.
- **Detached Execution**: The [`ui/jobs.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/jobs.py) module launches training processes via `subprocess.Popen` with `start_new_session=True`, ensuring jobs survive browser refreshes while writing to a JSON registry for status tracking.
- **Resource Safety**: A built-in GPU guard prevents multiple GPU-intensive jobs from running simultaneously by scanning the job registry before launching new processes.
- **Extensibility**: Adding new training stages requires only registering a new `Stage` entry and creating the corresponding config dataclass and script.

## Frequently Asked Questions

### How does the UI prevent multiple GPU jobs from running simultaneously?

The `gpu_busy()` function in [`ui/jobs.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/jobs.py) scans the JSON registry located at `/ephemeral/ui_jobs/` for any active job where `kind == "gpu"`. If a running GPU job is detected, the UI blocks new launch attempts and displays a warning to the user, ensuring exclusive GPU access for training stability.

### Can I run training without using the Streamlit interface?

Yes. The training scripts in `scripts/` (e.g., [`scripts/train_sft.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_sft.py), [`scripts/pretrain_base.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/pretrain_base.py)) are standalone Python files that accept JSON configuration paths via `--config` arguments. You can execute them directly using `python` or `torchrun` for multi-GPU training, bypassing the UI entirely while using the same configuration files.

### How do I add support for multi-GPU training in a new stage?

Set `multi_gpu=True` in the `Stage` dataclass definition within [`ui/stages.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/ui/stages.py). When users launch the job with `nproc > 1`, the `build_argv` function automatically wraps the command with `torchrun --standalone --nproc_per_node=<nproc>`, enabling distributed training across multiple GPUs.

### Where are training logs stored when launched through the UI?

Logs are written to `/ephemeral/ui_jobs/<job_id>.log` as plain text. The UI tails this file in real-time using the `tail_log` function, and the job status is determined by parsing the log for "traceback" strings to detect failures.