How to Use Soup's Autopilot Feature for Zero-Config Training
Soup's soup autopilot command automatically analyzes your hardware, dataset, and base model to generate a complete training configuration without manual hyperparameter tuning.
The autopilot feature in MakazhanAlpamys/Soup eliminates the complexity of finetuning large language models by inspecting your environment and making data-driven decisions about quantization strategy, LoRA settings, batch size, and performance optimizations. This guide walks through how to use Soup's autopilot for zero-config training based on the actual source implementation.
How the Autopilot Pipeline Works
The autopilot system runs three sequential stages to convert your inputs into a ready-to-train configuration.
Stage 1: Analyze (Hardware, Model, and Dataset Inspection)
The analyze stage gathers raw telemetry from your system and inputs. In src/soup_cli/autopilot/analyzer.py, three core functions perform this work:
analyze_hardware()– Usestorch.cuda(or MLX backend on Apple Silicon) to capture GPU name, available VRAM, and compute capabilityanalyze_dataset(data_path)– Extracts sample count, token length distribution, and format from your dataset directory or HuggingFace repoanalyze_model(model)– Reads model size, context window, and detects whether the checkpoint is already quantized
This analysis phase ensures all downstream decisions are grounded in actual system constraints rather than hardcoded defaults.
Stage 2: Decide (Hyperparameter Inference)
The decide stage, implemented in src/soup_cli/autopilot/decisions.py, applies rule-based heuristics to select optimal training parameters:
| Decision Function | Purpose |
|---|---|
decide_task |
Infers training objective from dataset structure and --goal flag |
decide_quantization |
Selects quantization strategy based on model format and VRAM pressure |
decide_batch_size |
Caps batch size to fit available VRAM with gradient accumulation |
decide_max_length |
Uses 95th-percentile token length from dataset analysis |
decide_peft |
Computes LoRA rank (r) and enables DoRA variant when beneficial |
decide_performance_flags |
Toggles Flash Attention, Liger kernels, gradient checkpointing |
Each function consumes the analysis profiles and returns type-safe parameters that feed into the final configuration.
Stage 3: Generate (Config Assembly and Output)
The generate stage in src/soup_cli/autopilot/generate_config.py orchestrates the pipeline through two key functions:
# build_soup_config() – lines 34-67
# write_yaml() – lines 70-82
build_soup_config() assembles all decisions into a SoupConfig Pydantic model, while write_yaml() serializes the configuration to disk. The output includes both recipe.yaml (the training config) and deploy.sh (a wrapper script).
Using the Autopilot CLI Command
The entry point is the soup autopilot subcommand, registered in the CLI at src/soup_cli/cli.py. Minimal invocation requires only three arguments:
soup autopilot \
--base meta-llama/Llama-3.2-1B \
--data /path/to/dataset \
--goal "chat"
Required Arguments
--base– HuggingFace model identifier or local checkpoint path--data– Dataset directory or HuggingFace dataset repo--goal– Training objective:"chat","completion","embedding", or other supported tasks
Optional Arguments
soup autopilot \
--base meta-llama/Llama-3.2-1B \
--data ./my_dataset \
--goal chat \
--output ./custom_output.yaml # Override default ./output/recipe.yaml
--vram-gb 24 # Manually specify VRAM instead of auto-detect
Running the Generated Training Configuration
Autopilot produces two artifacts in your working directory (or --output location):
recipe.yaml– CompleteSoupConfigvalidated againstsrc/soup_cli/config/schema.pydeploy.sh– Executable script wrappingsoup train --config recipe.yaml
Execute the training run with either method:
# Option 1: Use the generated script
bash deploy.sh
# Option 2: Invoke soup train directly
soup train --config ./output/recipe.yaml
Programmatic Autopilot Usage
Import the autopilot pipeline directly in Python for custom workflows:
from pathlib import Path
from soup_cli.autopilot.generate_config import build_soup_config, write_yaml
# Define inputs
base_model = "meta-llama/Llama-3.2-1B"
data_path = "/home/user/my_dataset"
goal = "chat"
vram_gb = None # Auto-detect from hardware
# Build complete config object
config = build_soup_config(
base=base_model,
data_path=data_path,
goal=goal,
vram_gb=vram_gb
)
# Persist to YAML
output_file = Path("./my_autopilot.yaml")
write_yaml(config, output_file)
print(f"✅ Autopilot config written to {output_file}")
Why Zero-Config Training Works
The autopilot system achieves true zero-configuration through five key mechanisms:
- Hardware awareness –
analyze_hardware()prevents out-of-memory errors by capping batch size to available VRAM - Model-aware quantization –
detect_prequantized_format()inanalyzer.pyavoids double-quantization by detecting existing quantized formats - Dataset-driven lengths –
decide_max_length()uses actual token distribution rather than arbitrary defaults - Automatic LoRA scaling –
decide_peft()adapts rank to dataset and model size, with conditional DoRA activation - Performance autotuning –
decide_performance_flags()enables optimizations only when hardware supports them
All decisions are validated through the SoupConfig schema in src/soup_cli/config/schema.py, ensuring type safety and compatibility with the training engine.
Summary
- The autopilot command runs analyze → decide → generate stages to produce complete training configs
- Three CLI arguments (
--base,--data,--goal) are sufficient to start training - Source files in
src/soup_cli/autopilot/handle hardware detection, heuristic decision-making, and YAML generation - Output includes both
recipe.yamlanddeploy.shfor immediate training execution - The system derives every hyperparameter from inputs, eliminating manual tuning for common scenarios
Frequently Asked Questions
What hardware backends does the autopilot support?
The autopilot supports CUDA GPUs via torch.cuda and Apple Silicon via MLX. The analyze_hardware() function in src/soup_cli/autopilot/analyzer.py detects the active backend automatically and adjusts VRAM calculations and performance flags accordingly.
Can I override specific hyperparameters after generation?
Yes. The generated recipe.yaml is a standard SoupConfig file that you can edit before training. The schema in src/soup_cli/config/schema.py documents all available fields. Alternatively, pass --vram-gb or other flags to influence the decision stage before generation.
Does autopilot work with pre-quantized models like GPTQ or AWQ?
Yes. The analyze_model() function detects pre-quantized formats and decide_quantization() in src/soup_cli/autopilot/decisions.py selects compatible quantization parameters to avoid re-quantization overhead or compatibility errors.
What happens if my dataset format is unrecognized?
The analyze_dataset() function attempts to infer format from file extensions and structure. If detection fails, the autopilot falls back to generic defaults and logs a warning. You may need to manually specify data configuration in edge cases, though the --goal flag helps disambiguate common formats.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →