How to Use Streamlit for Controlling LLM Training: A Complete Guide to train-llm-from-scratch
The FareedKhan-dev/train-llm-from-scratch repository provides a fully-featured Streamlit control panel that enables browser-based management of the entire LLM lifecycle—from data preparation and pretraining to supervised fine-tuning (SFT), reward modeling, and RLHF algorithms like DPO, PPO, and GRPO.
The FareedKhan-dev/train-llm-from-scratch repository ships with a production-ready web interface that demonstrates how to use Streamlit for controlling LLM training workflows without complex command-line interactions. This browser-based dashboard allows you to configure hyperparameters, launch detached training jobs, and monitor GPU utilization through an intuitive UI. The architecture cleanly separates the presentation layer from the training logic, using background subprocesses that survive page reloads while streaming real-time logs back to the browser.
Architecture of the Streamlit Control Panel
The UI is organized into four specialized modules that handle distinct responsibilities, from defining training stages to managing background processes.
Central Stage Registry (ui/stages.py)
The file ui/stages.py serves as the single source of truth for every pipeline stage. It defines a Stage dataclass that bundles everything a UI page needs: a configuration dataclass, JSON file locations, the training script path, log prefixes, documentation, and multi-GPU flags.
STAGES: dict[str, Stage] = {
"pretrain": Stage(
key="pretrain",
title="Pretraining",
emoji="📚",
cfg_cls=PretrainConfig,
config_json="configs/pretrain.json",
smoke_json="configs/smoke/pretrain.json",
script="scripts/pretrain_base.py",
log_prefix="pretrain",
doc_md="docs/02_pretraining.md",
diagram_png="docs/diagrams/02_pretraining.png",
multi_gpu=True,
),
# ... sft, reward, dpo, ppo, grpo ...
}
Each entry in the STAGES dictionary references a dataclass (e.g., PretrainConfig from config/post_training_config.py) that parses the JSON config, the script that executes the training, and metadata for documentation rendering.
Background Job Manager (ui/jobs.py)
The ui/jobs.py module handles the heavy lifting of controlling LLM training processes outside the Streamlit thread. The build_argv function constructs the command line, automatically injecting torchrun for multi-GPU scenarios:
def build_argv(script, config_json, nproc, multi_gpu, extra=None):
base = [script, "--config", config_json] + (extra or [])
if multi_gpu and nproc > 1:
return [_TORCHRUN, "--standalone", f"--nproc_per_node={nproc}", *base]
return [sys.executable, *base]
The launch function writes a JSON registry to /ephemeral/ui_jobs/ containing the pid, command, start time, and status, then spawns the process using subprocess.Popen(..., start_new_session=True). This ensures training continues even if you close the browser.
UI Styling and GPU Monitoring (ui/theme.py)
The ui/theme.py file defines the visual layer through a COLORS dictionary and CSS blocks. It provides the gpu_status() helper that calls nvidia-smi to render real-time GPU metrics, and status_badge() for color-coded job state indicators.
Runtime Workflow: From Config to Training
Understanding how Streamlit controls the LLM training lifecycle helps you debug and extend the interface.
- Configuration Editing: Sidebar pages auto-generated from
STAGESlet users edit JSON configs via Streamlit forms. - Launch Sequence: When the Launch button is clicked, the page calls
jobs.build_argv→jobs.launch, which builds the command line and spawns a detached process. - GPU Guard: The
gpu_busy()function scans the JSON registry for any running job withkind == "gpu"and blocks new launches if a GPU job is active, preventing resource conflicts. - Log Streaming: The UI polls the registry via
jobs.status()and tails the log file at/ephemeral/ui_jobs/<job_id>.log, updating badges and log views in real time.
Extending the UI with Custom Training Stages
You can add new stages to the pipeline by modifying the central registry. This is the primary extension point for customizing how Streamlit controls your LLM training workflow.
- Create a Config Dataclass: Add a new dataclass to
config/post_training_config.py(e.g.,MyNewConfig). - Add JSON Templates: Create
configs/my_new.jsonand an optional smoke test version atconfigs/smoke/my_new.json. - Write the Training Script: Implement the logic in
scripts/train_my_new.py. - Register the Stage: Insert a new entry into
STAGESinui/stages.py:
STAGES["my_new"] = Stage(
key="my_new",
title="My New Stage",
emoji="🚀",
cfg_cls=MyNewConfig,
config_json="configs/my_new.json",
smoke_json="configs/smoke/my_new.json",
script="scripts/train_my_new.py",
log_prefix="my_new",
doc_md="docs/09_my_new.md",
diagram_png="docs/diagrams/09_my_new.png",
multi_gpu=False,
)
- Documentation: Add a markdown doc and diagram, then run
streamlit run ui/app.py—the new page appears automatically in the sidebar.
Practical Implementation Examples
Launching a Training Stage Programmatically
This example demonstrates how the UI internally launches an SFT job using the jobs module:
import streamlit as st
import time
from ui import jobs, stages
# Select the SFT stage configuration
stage = stages.STAGES["sft"]
cfg_path = stage.config_json # "configs/sft.json"
script = stage.script # "scripts/train_sft.py"
nproc = 2 # Number of GPUs
# Build the command-line arguments
argv = jobs.build_argv(
script=script,
config_json=cfg_path,
nproc=nproc,
multi_gpu=stage.multi_gpu,
)
# Generate unique job ID and launch
job_id = f"sft_{int(time.time())}"
rec = jobs.launch(job_id, argv, kind="gpu")
st.success(f"Launched {stage.title} with job ID **{job_id}**")
Checking GPU Availability
Use the gpu_status helper to programmatically check resource availability before launching:
from ui.theme import gpu_status
gpus = gpu_status()
if gpus:
for i, name, used, total, util in gpus:
st.metric(f"GPU {i}", f"{used}/{total} GB", f"{util}% util")
else:
st.warning("No GPUs detected")
Template for a New UI Page
Create ui/pages/my_new.py to expose your custom stage:
import streamlit as st
import json
import pathlib
from ui import jobs, stages, theme
stage = stages.STAGES["my_new"]
st.title(f"{stage.emoji} {stage.title}")
st.image(stage.diagram_png)
# Config form
with st.form("cfg_form"):
epochs = st.number_input("Epochs", min_value=1, value=5)
lr = st.text_input("Learning rate", value="5e-4")
submitted = st.form_submit_button("Save config")
if submitted:
cfg = {"epochs": epochs, "lr": float(lr)}
pathlib.Path(stage.config_json).write_text(
json.dumps(cfg, indent=2)
)
# Launch button
if st.button("Launch Training"):
argv = jobs.build_argv(
stage.script,
stage.config_json,
nproc=1,
multi_gpu=False
)
job_id = f"my_new_{int(time.time())}"
jobs.launch(job_id, argv, kind="cpu")
st.success(f"Launched {stage.title}")
Summary
- Centralized Configuration: The
ui/stages.pyfile acts as the single source of truth, binding together config files, training scripts, and documentation for each pipeline stage. - Detached Execution: The
ui/jobs.pymodule launches training processes viasubprocess.Popenwithstart_new_session=True, ensuring jobs survive browser refreshes while writing to a JSON registry for status tracking. - Resource Safety: A built-in GPU guard prevents multiple GPU-intensive jobs from running simultaneously by scanning the job registry before launching new processes.
- Extensibility: Adding new training stages requires only registering a new
Stageentry and creating the corresponding config dataclass and script.
Frequently Asked Questions
How does the UI prevent multiple GPU jobs from running simultaneously?
The gpu_busy() function in ui/jobs.py scans the JSON registry located at /ephemeral/ui_jobs/ for any active job where kind == "gpu". If a running GPU job is detected, the UI blocks new launch attempts and displays a warning to the user, ensuring exclusive GPU access for training stability.
Can I run training without using the Streamlit interface?
Yes. The training scripts in scripts/ (e.g., scripts/train_sft.py, scripts/pretrain_base.py) are standalone Python files that accept JSON configuration paths via --config arguments. You can execute them directly using python or torchrun for multi-GPU training, bypassing the UI entirely while using the same configuration files.
How do I add support for multi-GPU training in a new stage?
Set multi_gpu=True in the Stage dataclass definition within ui/stages.py. When users launch the job with nproc > 1, the build_argv function automatically wraps the command with torchrun --standalone --nproc_per_node=<nproc>, enabling distributed training across multiple GPUs.
Where are training logs stored when launched through the UI?
Logs are written to /ephemeral/ui_jobs/<job_id>.log as plain text. The UI tails this file in real-time using the tail_log function, and the job status is determined by parsing the log for "traceback" strings to detect failures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →