How to Export Models to GGUF Format for Ollama Deployment with Soup
Use the soup export CLI command with --format gguf-ud and --auto-deploy ollama flags to convert a Hugging Face checkpoint into a quantized GGUF file and immediately register it with Ollama.
Soup provides a complete pipeline for converting transformers-based models into GGUF format—the binary format that powers Ollama and llama.cpp. The implementation spans multiple modules in src/soup_cli/ and handles conversion, optional importance matrix generation, quantization, and automatic Ollama registration in a single command.
The Three-Stage GGUF Export Pipeline
According to the MakazhanAlpamys/Soup source code, every GGUF export proceeds through three distinct stages implemented in src/soup_cli/utils/gguf_quant.py.
Stage 1: Convert HF Checkpoint to FP16 GGUF
The export_advanced_gguf function calls _run_convert_to_f16, which executes llama.cpp/convert_hf_to_gguf.py to produce a temporary f16.gguf file【gguf_quant.py L559-L580】.
This intermediate file preserves full precision before quantization and is automatically cleaned up after successful export (unless tests like test_issue144_gguf_export.py guard against accidental deletion).
Stage 2: Generate Importance Matrix (Optional)
For UD ladder flavours (e.g., UD-Q4_K_XL) and low-bit IQ types, an importance matrix improves quantization quality. The pipeline:
- Parses your calibration JSON-L through
_prepare_calibration_textto extract plain text【gguf_quant.py L71-L84】 - Runs
llama-imatrixvia_run_imatrixto produceimatrix.dat【gguf_quant.py L95-L111】
Flavours without UD- or IQ- prefixes skip this stage.
Stage 3: Quantize to Final GGUF
The _run_quantize_binary function builds command arguments for llama-quantize (or legacy quantize), applying your chosen flavour and imatrix when present【gguf_quant.py L122-L138】. Output writes directly to your specified path.
Choosing a GGUF Flavour
Soup organizes flavours into three categories defined in ALL_ADVANCED_GGUF_FORMATS【gguf_quant.py L70-L75】:
| Category | Examples | Needs Imatrix? |
|---|---|---|
| UD ladder (Ultra-Low Distortion) | UD-Q8_K_XL, UD-Q5_K_M, UD-Q4_K_XL |
Yes |
| IQ family (Information-Retaining Quantization) | IQ1_S, IQ2_XXS, IQ3_M |
Yes |
| Standard/ARM | Q4_0_4_4, Q5_1, Q8_0 |
No |
Complete Export Workflow
1. Prepare Calibration Data (If Required)
Create a JSON-L file for UD or IQ flavours. Each line needs a text, prompt, or content field:
{"text": "The transformer architecture revolutionized natural language processing by introducing self-attention mechanisms..."}
{"text": "Quantization reduces model size by representing weights with fewer bits while preserving task performance..."}
2. Run the Export Command
soup export \
--model /path/to/huggingface/checkpoint \
--format gguf-ud \
--gguf-flavour UD-Q4_K_XL \
--calibration-data calibration.jsonl \
--auto-deploy ollama \
--deploy-name my-custom-llama
The command flow:
- Validates the flavour in
_export_gguf_advanced(src/soup_cli/commands/export.py)【export.py L14-L34】 - Executes the three-stage pipeline with progress reporting
- Triggers
_auto_deploy_ollamawhich:- Detects Ollama binary via
detect_ollama【ollama.py L123-L136】 - Validates model name through
validate_model_name - Creates Modelfile and runs
ollama createviadeploy_to_ollama【ollama.py L208-L230】
- Detects Ollama binary via
3. Use Your Model
ollama run my-custom-llama
Programmatic Export Without CLI
Import export_advanced_gguf directly for custom workflows:
from soup_cli.utils.gguf_quant import export_advanced_gguf
# Standard ARM flavour—no calibration needed
export_advanced_gguf(
model_dir="meta-llama/Llama-2-7b-hf",
output_path="llama-2-7b.Q4_0_4_4.gguf",
flavour="Q4_0_4_4",
calibration_data=None,
llama_cpp_dir="~/.soup/llama.cpp",
)
Manual Ollama Deployment
For existing GGUF files, use deploy_to_ollama from src/soup_cli/utils/ollama.py:
from soup_cli.utils.ollama import deploy_to_ollama
modelfile = """FROM ./my_model.Q4_0_4_4.gguf
PARAMETER temperature 0.7
PARAMETER top_p 0.9
SYSTEM You are a helpful coding assistant."""
success, message = deploy_to_ollama("code-assistant", modelfile)
print(message)
Security and Path Handling
All file operations pass through enforce_under_cwd_and_no_symlink in src/soup_cli/utils/paths.py, preventing path-traversal and symlink attacks during the export process.
Key Source Files Reference
| File | Purpose |
|---|---|
src/soup_cli/utils/gguf_quant.py |
Core three-stage export, flavour validation, calibration parsing |
src/soup_cli/commands/export.py |
CLI command entry point _export_gguf_advanced |
src/soup_cli/utils/ollama.py |
Ollama detection, Modelfile generation, deployment |
src/soup_cli/utils/paths.py |
Security-hardened path utilities |
tests/test_issue144_gguf_export.py |
Guards against f16 file deletion |
tests/test_v0531_139.py |
Integration tests for full pipeline |
Summary
- Three-stage pipeline: Convert → Imatrix (optional) → Quantize, all orchestrated by
export_advanced_gguf - CLI entry:
soup export --format gguf-ud --gguf-flavour <flavour> --auto-deploy ollama - Flavour selection: UD and IQ types need calibration data; standard ARM types do not
- Security: Path traversal protection via
enforce_under_cwd_and_no_symlink - Deployment: Automatic Ollama registration with
--auto-deploy ollamaflag
Frequently Asked Questions
What calibration data format does Soup expect for GGUF export?
Soup accepts JSON-L files where each line contains a JSON object with a text, prompt, or content field. The _prepare_calibration_text function in gguf_quant.py extracts plain text from these fields for the llama-imatrix tool【gguf_quant.py L71-L84】.
Do all GGUF flavours require an importance matrix?
No. Only UD ladder (UD-Q4_K_XL, UD-Q5_K_M, etc.) and IQ family (IQ1_S, IQ2_XXS, etc.) flavours require an imatrix. Standard quantizations like Q4_0_4_4 or Q8_0 skip imatrix generation entirely, making them faster to export.
What happens to the intermediate f16.gguf file?
The temporary FP16 file created during Stage 1 is automatically removed after successful quantization. The test suite in test_issue144_gguf_export.py specifically protects against accidental deletion of user-generated files that might share similar naming patterns.
Can I deploy to Ollama without using the --auto-deploy flag?
Yes. After export, call deploy_to_ollama programmatically with a custom Modelfile, or manually create a Modelfile pointing to your GGUF and run ollama create <name> -f Modelfile. The deploy_to_ollama function in ollama.py handles the binary detection and command execution for you【ollama.py L208-L230】.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →