What Is the Autopilot Feature in Soup CLI? A Complete Guide to Zero-Configuration Fine‑Tuning
The Autopilot feature in Soup CLI (soup autopilot) is a zero-configuration fine-tuning assistant that automatically analyzes your dataset, model, and hardware to generate optimal training hyperparameters—no manual tuning required.
The Autopilot feature, introduced in Soup v0.25.0 and refined through releases like v0.53.1 (which added pre-quantized detection and live measurement), removes the guesswork from configuring transformer fine-tuning jobs. Instead of hand-tuning dozens of hyperparameters, users provide only a model identifier, dataset path, and high-level goal—Autopilot handles the rest.
How Autopilot Works: The Five-Stage Pipeline
The Autopilot command executes a deterministic, locally-run pipeline with no network dependencies. All analysis happens on your machine using only the supplied data, model name, and detected hardware.
Stage 1: Input Analysis
Autopilot begins by profiling three critical inputs in parallel:
- Dataset analysis —
src/soup_cli/autopilot/analyzer.py→analyze_dataset()examines your data file to determine format, size, and sequence characteristics - Model analysis —
analyze_model()in the same module inspects the model architecture, parameter count, and pre-quantization status - Hardware analysis —
analyze_hardware()detects available GPU/CPU resources, compute capability, and VRAM
Stage 2: Task Derivation
Based on your --goal parameter (e.g., chat, reasoning, code), Autopilot maps to a concrete training task using the GOAL_TO_TASK dictionary in src/soup_cli/autopilot/decisions.py. A chat goal typically resolves to supervised fine-tuning (sft), while reasoning may trigger specialized task configurations.
Stage 3: Hyperparameter Decisions
The decision engine in decisions.py automatically determines:
- Quantization strategy —
decide_quantization()selectsnone,8bit,4bit, or respects pre-quantized formats - PEFT configuration —
decide_peft()sets LoRA rank and alpha values - Batch sizing —
decide_batch_size()and gradient accumulation steps based on available VRAM - Learning rate —
decide_lr()scales appropriately for model size and task - Training duration —
decide_epochs()from dataset size and convergence heuristics - Sequence length —
decide_max_length()balancing memory and task requirements - Performance optimizations —
decide_performance_flags()enables FlashAttention and Liger kernels for Ampere-or-newer GPUs, plus gradient checkpointing when VRAM is constrained
Stage 4: Configuration Generation
The build_soup_config() function in src/soup_cli/autopilot/generate_config.py assembles all decisions into a validated SoupConfig Pydantic schema—ensuring type safety and schema compliance.
Stage 5: Output & Persistence
Autopilot renders a human-readable summary of all decisions. Unless --dry-run is specified, write_yaml() saves the configuration to disk (default: soup.yaml).
Running Autopilot: Code Examples
Basic Zero-Configuration Usage
The simplest invocation requires only three arguments:
soup autopilot \
--model meta-llama/Llama-3.2-1B \
--data ./my_dataset.jsonl \
--goal chat
Autopilot infers sft task, selects appropriate quantization (likely 4bit for this model size on consumer hardware), sets LoRA parameters, and writes soup.yaml.
Preview Decisions Without Writing Files
Use --dry-run to validate Autopilot's choices before committing:
soup autopilot \
--model Qwen/Qwen2.5-7B \
--data ./data/math.jsonl \
--goal reasoning \
--dry-run
Output shows decision panels only—no file creation, no side effects.
Constrain to Specific GPU Budget
Explicit VRAM limits influence quantization and batch-size decisions through decisions.parse_gpu_budget():
soup autopilot \
--model meta-llama/Llama-3.1-8B-Instruct \
--data ./data/chat.jsonl \
--goal chat \
--gpu-budget 24GB \
--output my_config.yaml
Skip Confirmation Prompts
Automation-friendly mode with --yes:
soup autopilot \
--model tiny-gpt2 \
--data ./tiny.jsonl \
--goal code \
--yes
Bypasses the overwrite confirmation when soup.yaml (or specified output) already exists.
Core Implementation Files
| File | Purpose |
|---|---|
src/soup_cli/commands/autopilot.py |
Typer command entry point; parses CLI arguments and orchestrates pipeline execution |
src/soup_cli/autopilot/analyzer.py |
Implements analyze_dataset(), analyze_model(), analyze_hardware() for input profiling |
src/soup_cli/autopilot/decisions.py |
Decision engine: task mapping, quantization, PEFT, batch sizing, learning rate, epochs, performance flags, GPU budget parsing |
src/soup_cli/autopilot/generate_config.py |
Builds final SoupConfig and handles YAML serialization via build_soup_config() and write_yaml() |
src/soup_cli/utils/gpu.py |
GPU detection, model size estimation, and batch-size estimation helpers consumed by the decision engine |
Key Design Principles
- Fully local — No API calls or network dependencies; all decisions derived from local data and hardware detection
- Deterministic — Same inputs produce identical configurations (hardware-permitting)
- Hardware-aware — Automatically adapts to detected GPU generation, VRAM, and compute capability
- Extensible — Modular decision functions allow easy addition of new optimization strategies
Summary
- Autopilot (
soup autopilot) eliminates manual hyperparameter tuning for transformer fine-tuning - Analyzes dataset, model, and hardware locally through functions in
analyzer.py - Derives optimal hyperparameters via decision functions in
decisions.py - Generates validated SoupConfig schemas and persists to YAML
- Supports dry-run preview, custom GPU budgets, and non-interactive execution modes
- Available since v0.25.0 with continuous improvements through v0.53.1+
Frequently Asked Questions
What GPU hardware does Autopilot optimize for?
Autopilot detects your GPU generation through src/soup_cli/utils/gpu.py. For Ampere or newer (SM80+) GPUs, it automatically enables FlashAttention and Liger kernels in decide_performance_flags(). For memory-constrained scenarios, it adds gradient checkpointing regardless of GPU age.
Can I override individual Autopilot decisions?
The current implementation generates complete configurations. For granular control, run with --dry-run, inspect the output, manually edit soup.yaml, or bypass Autopilot entirely with soup train --config custom.yaml. Direct per-parameter overrides in the Autopilot command are not supported—this maintains the zero-configuration design philosophy.
How does Autopilot handle pre-quantized models?
Since v0.53.1, decide_quantization() in decisions.py detects pre-quantized model formats through analyze_model(). When found, Autopilot respects the existing quantization instead of applying additional layers, preventing double-quantization degradation.
What happens if my dataset format is unsupported?
analyze_dataset() in analyzer.py validates format compatibility during Stage 1. Currently supported formats include JSONL with standard conversation or instruction-following schemas. Autopilot raises clear errors with format guidance rather than proceeding with incorrect assumptions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →