Unsloth CLI Options for Model Training and Inference: A Complete Guide

The unsloth-cli tool automatically generates training and inference flags from Pydantic configuration models, exposing 40+ options for fine-tuning and 10+ parameters for text generation without requiring Python scripting.

The unslothai/unsloth repository provides a streamlined command-line interface for efficient LLM fine-tuning and inference. Built on Typer and Pydantic, the unsloth-cli dynamically constructs its argument parser from configuration schemas, ensuring CLI flags stay synchronized with the underlying Python API.

How unsloth-cli Generates Configuration Options

The CLI architecture relies on automatic reflection over Pydantic models to produce command-line arguments.

The Config Schema in config.py

In unsloth_cli/config.py, a hierarchical Config model defines all user-configurable fields. The model nests several sub-configurations:

  • model – Model identifier and loading parameters
  • data – Dataset configuration including format type
  • training – Hyperparameters for optimization and sequencing
  • lora – LoRA-specific adapter configuration
  • logging – Weights & Biases and TensorBoard settings

Any field added to this model automatically becomes available as a CLI flag without modifying command code.

Automatic Flag Generation in options.py

The unsloth_cli/options.py file implements the add_options_from_config decorator. This utility:

  • Walks the Config model recursively, including nested BaseModel instances
  • Flattens every field into a kebab-case CLI flag (e.g., max_seq_length becomes --max-seq-length)
  • Generates dual boolean flags (--enable-wandb / --no-enable-wandb) for boolean fields
  • Skips list-type fields (Typer cannot infer sensible representations for collections)
  • Builds a config_overrides dictionary passed to the command implementation

Commands in unsloth_cli/commands/train.py apply this decorator to receive a fully populated Config object.

Training Command Options (unsloth-cli train)

The train command, defined in unsloth_cli/commands/train.py, accepts the auto-generated flags from Config plus several explicit options:

Explicit Training Flags:

  • --config / -c – Path to YAML/JSON configuration file (CLI flags override file values)
  • --hf-token – Hugging Face authentication token (falls back to HF_TOKEN environment variable)
  • --wandb-token – Weights & Biases API key (fallback to WANDB_API_KEY)
  • --dry-run – Validate and print the resolved configuration without launching training

Model & Data Configuration:

  • --model – Hugging Face model ID or local path (e.g., meta-llama/Meta-Llama-3-8B-Instruct)
  • --dataset – Hugging Face dataset name to download
  • --format-type – Dataset format selector (auto, alpaca, chatml, sharegpt)
  • --training-type – Fine-tuning strategy (lora or full)

Training Hyperparameters:

  • --max-seq-length – Maximum sequence length for the model
  • --load-in-4bit / --no-load-in-4bit – 4-bit quantization toggle (default: enabled)
  • --output-dir – Directory for checkpoints and logs
  • --num-epochs – Number of training epochs
  • --learning-rate – Optimizer learning rate
  • --batch-size – Per-device batch size
  • --gradient-accumulation-steps – Steps to accumulate before optimizer step
  • --warmup-steps – Warm-up period duration
  • --max-steps – Hard limit on training steps (0 disables)
  • --save-steps – Checkpoint frequency (0 for final only)
  • --weight-decay – Regularization coefficient
  • --random-seed – Reproducibility seed
  • --packing / --no-packing – Sequence packing toggle (default: disabled)
  • --train-on-completions / --no-train-on-completions – Target the tail of samples as completion targets
  • --gradient-checkpointing – Mode selection (unsloth, true, none)

LoRA Adapter Options:

  • --lora-r – LoRA rank dimension
  • --lora-alpha – Scaling factor
  • --lora-dropout – Dropout probability
  • --target-modules – Comma-separated module names (e.g., q_proj,k_proj,v_proj)
  • --vision-all-linear / --no-vision-all-linear – Apply LoRA to all linear layers in vision models
  • --use-rslora / --no-use-rslora – Enable RSLora regularization
  • --use-loftq / --no-use-loftq – Enable LoFTQ quantization
  • --finetune-vision-layers / --no-finetune-vision-layers – Tune vision-specific layers
  • --finetune-language-layers / --no-finetune-language-layers – Tune language layers
  • --finetune-attention-modules / --no-finetune-attention-modules – Tune attention modules
  • --finetune-mlp-modules / --no-finetune-mlp-modules – Tune MLP modules

Logging Options:

  • --enable-wandb / --no-enable-wandb – Weights & Biases integration
  • --wandb-project – Project name (default: unsloth-training)
  • --enable-tensorboard / --no-enable-tensorboard – TensorBoard logging
  • --tensorboard-dir – Log directory path

Note: List fields such as local_dataset are intentionally excluded from CLI exposure.

Inference Command Options (unsloth-cli inference)

The inference command in unsloth_cli/commands/inference.py defines its own Typer options rather than using the Config decorator:

Positional Arguments:

  • model – Hugging Face ID or local checkpoint path
  • prompt – Input text to send to the model

Authentication & Loading:

  • --hf-token – Hugging Face token
  • --max-seq-length – Context window size (default: 2048)
  • --load-in-4bit / --no-load-in-4bit – Quantization toggle (default: enabled)

Sampling Parameters:

  • --temperature – Sampling temperature (default: 0.7)
  • --top-p – Nucleus sampling cutoff (default: 0.9)
  • --top-k – Top-k token selection (default: 40)
  • --max-new-tokens – Generation limit (default: 256)
  • --repetition-penalty – Token repetition penalty (default: 1.1)
  • --system-prompt – Optional system-level instruction prepended to the conversation

Practical Usage Examples

Launching a LoRA Fine-Tuning Run

unsloth-cli train \
    --model meta-llama/Meta-Llama-3-8B-Instruct \
    --dataset tatsu-lab/alpaca-gpt4-data \
    --training-type lora \
    --num-epochs 5 \
    --learning-rate 1e-4 \
    --batch-size 4 \
    --lora-r 128 \
    --lora-alpha 32 \
    --target-modules "q_proj,k_proj,v_proj" \
    --enable-wandb \
    --wandb-project my-llama3-run \
    --hf-token $HF_TOKEN

Validating Configuration with Dry-Run

unsloth-cli train \
    --model mistralai/Mistral-7B-Instruct-v0.2 \
    --dry-run \
    --learning-rate 2e-4

The --dry-run flag prints the merged YAML configuration including all defaults and exits, allowing verification before allocating GPU resources.

Single-Prompt Inference

unsloth-cli inference \
    meta-llama/Meta-Llama-3-8B-Instruct \
    "Write a short poem about autumn." \
    --temperature 0.8 \
    --top-p 0.95 \
    --max-new-tokens 150 \
    --load-in-4bit

Inference with System Prompts

unsloth-cli inference \
    "my_local/llama-8b" \
    "Explain quantum entanglement in simple terms." \
    --system-prompt "You are a friendly teacher." \
    --max-seq-length 4096 \
    --no-load-in-4bit

Summary

  • The unsloth-cli dynamically generates training options from the Config Pydantic model in unsloth_cli/config.py using the add_options_from_config decorator in unsloth_cli/options.py.
  • Boolean fields automatically receive dual flags (--flag and --no-flag), while list fields are excluded from CLI exposure.
  • The train command supports 40+ configuration options covering quantization, LoRA parameters, optimization settings, and logging integrations.
  • The inference command accepts positional arguments for model and prompt, plus explicit sampling controls for temperature, top-p, and repetition penalty.
  • Configuration files in YAML/JSON format can seed training runs, with CLI flags taking precedence over file values.

Frequently Asked Questions

How do I load a custom configuration file in unsloth-cli?

Use the --config or -c flag followed by the path to a YAML or JSON file. Values specified via CLI flags override those in the configuration file. For example: unsloth-cli train --config base.yaml --learning-rate 5e-5.

Why are some configuration fields not available as command-line flags?

The add_options_from_config decorator in unsloth_cli/options.py skips list-type fields because Typer cannot infer a sensible string representation for collections on the command line. Use a configuration file for complex data structures like local_dataset lists.

Can I use unsloth-cli for inference without 4-bit quantization?

Yes. Pass the --no-load-in-4bit flag to disable quantization. This loads the model at full precision, useful when you have sufficient VRAM and require maximum accuracy, as shown when running inference with custom max-seq-length values.

How do I enable Weights & Biases logging from the command line?

Set --enable-wandb and optionally specify --wandb-project to define the project name. You can provide the API key via --wandb-token or the WANDB_API_KEY environment variable. The integration automatically logs training metrics, hyperparameters, and model artifacts according to the configuration in unsloth_cli/config.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →