Unsloth CLI Options for Model Training and Inference: A Complete Guide
The unsloth-cli tool automatically generates training and inference flags from Pydantic configuration models, exposing 40+ options for fine-tuning and 10+ parameters for text generation without requiring Python scripting.
The unslothai/unsloth repository provides a streamlined command-line interface for efficient LLM fine-tuning and inference. Built on Typer and Pydantic, the unsloth-cli dynamically constructs its argument parser from configuration schemas, ensuring CLI flags stay synchronized with the underlying Python API.
How unsloth-cli Generates Configuration Options
The CLI architecture relies on automatic reflection over Pydantic models to produce command-line arguments.
The Config Schema in config.py
In unsloth_cli/config.py, a hierarchical Config model defines all user-configurable fields. The model nests several sub-configurations:
model– Model identifier and loading parametersdata– Dataset configuration including format typetraining– Hyperparameters for optimization and sequencinglora– LoRA-specific adapter configurationlogging– Weights & Biases and TensorBoard settings
Any field added to this model automatically becomes available as a CLI flag without modifying command code.
Automatic Flag Generation in options.py
The unsloth_cli/options.py file implements the add_options_from_config decorator. This utility:
- Walks the
Configmodel recursively, including nestedBaseModelinstances - Flattens every field into a kebab-case CLI flag (e.g.,
max_seq_lengthbecomes--max-seq-length) - Generates dual boolean flags (
--enable-wandb / --no-enable-wandb) for boolean fields - Skips list-type fields (Typer cannot infer sensible representations for collections)
- Builds a
config_overridesdictionary passed to the command implementation
Commands in unsloth_cli/commands/train.py apply this decorator to receive a fully populated Config object.
Training Command Options (unsloth-cli train)
The train command, defined in unsloth_cli/commands/train.py, accepts the auto-generated flags from Config plus several explicit options:
Explicit Training Flags:
--config / -c– Path to YAML/JSON configuration file (CLI flags override file values)--hf-token– Hugging Face authentication token (falls back toHF_TOKENenvironment variable)--wandb-token– Weights & Biases API key (fallback toWANDB_API_KEY)--dry-run– Validate and print the resolved configuration without launching training
Model & Data Configuration:
--model– Hugging Face model ID or local path (e.g.,meta-llama/Meta-Llama-3-8B-Instruct)--dataset– Hugging Face dataset name to download--format-type– Dataset format selector (auto,alpaca,chatml,sharegpt)--training-type– Fine-tuning strategy (loraorfull)
Training Hyperparameters:
--max-seq-length– Maximum sequence length for the model--load-in-4bit / --no-load-in-4bit– 4-bit quantization toggle (default: enabled)--output-dir– Directory for checkpoints and logs--num-epochs– Number of training epochs--learning-rate– Optimizer learning rate--batch-size– Per-device batch size--gradient-accumulation-steps– Steps to accumulate before optimizer step--warmup-steps– Warm-up period duration--max-steps– Hard limit on training steps (0disables)--save-steps– Checkpoint frequency (0for final only)--weight-decay– Regularization coefficient--random-seed– Reproducibility seed--packing / --no-packing– Sequence packing toggle (default: disabled)--train-on-completions / --no-train-on-completions– Target the tail of samples as completion targets--gradient-checkpointing– Mode selection (unsloth,true,none)
LoRA Adapter Options:
--lora-r– LoRA rank dimension--lora-alpha– Scaling factor--lora-dropout– Dropout probability--target-modules– Comma-separated module names (e.g.,q_proj,k_proj,v_proj)--vision-all-linear / --no-vision-all-linear– Apply LoRA to all linear layers in vision models--use-rslora / --no-use-rslora– Enable RSLora regularization--use-loftq / --no-use-loftq– Enable LoFTQ quantization--finetune-vision-layers / --no-finetune-vision-layers– Tune vision-specific layers--finetune-language-layers / --no-finetune-language-layers– Tune language layers--finetune-attention-modules / --no-finetune-attention-modules– Tune attention modules--finetune-mlp-modules / --no-finetune-mlp-modules– Tune MLP modules
Logging Options:
--enable-wandb / --no-enable-wandb– Weights & Biases integration--wandb-project– Project name (default:unsloth-training)--enable-tensorboard / --no-enable-tensorboard– TensorBoard logging--tensorboard-dir– Log directory path
Note: List fields such as local_dataset are intentionally excluded from CLI exposure.
Inference Command Options (unsloth-cli inference)
The inference command in unsloth_cli/commands/inference.py defines its own Typer options rather than using the Config decorator:
Positional Arguments:
model– Hugging Face ID or local checkpoint pathprompt– Input text to send to the model
Authentication & Loading:
--hf-token– Hugging Face token--max-seq-length– Context window size (default: 2048)--load-in-4bit / --no-load-in-4bit– Quantization toggle (default: enabled)
Sampling Parameters:
--temperature– Sampling temperature (default: 0.7)--top-p– Nucleus sampling cutoff (default: 0.9)--top-k– Top-k token selection (default: 40)--max-new-tokens– Generation limit (default: 256)--repetition-penalty– Token repetition penalty (default: 1.1)--system-prompt– Optional system-level instruction prepended to the conversation
Practical Usage Examples
Launching a LoRA Fine-Tuning Run
unsloth-cli train \
--model meta-llama/Meta-Llama-3-8B-Instruct \
--dataset tatsu-lab/alpaca-gpt4-data \
--training-type lora \
--num-epochs 5 \
--learning-rate 1e-4 \
--batch-size 4 \
--lora-r 128 \
--lora-alpha 32 \
--target-modules "q_proj,k_proj,v_proj" \
--enable-wandb \
--wandb-project my-llama3-run \
--hf-token $HF_TOKEN
Validating Configuration with Dry-Run
unsloth-cli train \
--model mistralai/Mistral-7B-Instruct-v0.2 \
--dry-run \
--learning-rate 2e-4
The --dry-run flag prints the merged YAML configuration including all defaults and exits, allowing verification before allocating GPU resources.
Single-Prompt Inference
unsloth-cli inference \
meta-llama/Meta-Llama-3-8B-Instruct \
"Write a short poem about autumn." \
--temperature 0.8 \
--top-p 0.95 \
--max-new-tokens 150 \
--load-in-4bit
Inference with System Prompts
unsloth-cli inference \
"my_local/llama-8b" \
"Explain quantum entanglement in simple terms." \
--system-prompt "You are a friendly teacher." \
--max-seq-length 4096 \
--no-load-in-4bit
Summary
- The unsloth-cli dynamically generates training options from the
ConfigPydantic model inunsloth_cli/config.pyusing theadd_options_from_configdecorator inunsloth_cli/options.py. - Boolean fields automatically receive dual flags (
--flagand--no-flag), while list fields are excluded from CLI exposure. - The
traincommand supports 40+ configuration options covering quantization, LoRA parameters, optimization settings, and logging integrations. - The
inferencecommand accepts positional arguments for model and prompt, plus explicit sampling controls for temperature, top-p, and repetition penalty. - Configuration files in YAML/JSON format can seed training runs, with CLI flags taking precedence over file values.
Frequently Asked Questions
How do I load a custom configuration file in unsloth-cli?
Use the --config or -c flag followed by the path to a YAML or JSON file. Values specified via CLI flags override those in the configuration file. For example: unsloth-cli train --config base.yaml --learning-rate 5e-5.
Why are some configuration fields not available as command-line flags?
The add_options_from_config decorator in unsloth_cli/options.py skips list-type fields because Typer cannot infer a sensible string representation for collections on the command line. Use a configuration file for complex data structures like local_dataset lists.
Can I use unsloth-cli for inference without 4-bit quantization?
Yes. Pass the --no-load-in-4bit flag to disable quantization. This loads the model at full precision, useful when you have sufficient VRAM and require maximum accuracy, as shown when running inference with custom max-seq-length values.
How do I enable Weights & Biases logging from the command line?
Set --enable-wandb and optionally specify --wandb-project to define the project name. You can provide the API key via --wandb-token or the WANDB_API_KEY environment variable. The integration automatically logs training metrics, hyperparameters, and model artifacts according to the configuration in unsloth_cli/config.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →