How to Export Models to GGUF Format for Ollama Deployment with Soup

Use the soup export CLI command with --format gguf-ud and --auto-deploy ollama flags to convert a Hugging Face checkpoint into a quantized GGUF file and immediately register it with Ollama.

Soup provides a complete pipeline for converting transformers-based models into GGUF format—the binary format that powers Ollama and llama.cpp. The implementation spans multiple modules in src/soup_cli/ and handles conversion, optional importance matrix generation, quantization, and automatic Ollama registration in a single command.

The Three-Stage GGUF Export Pipeline

According to the MakazhanAlpamys/Soup source code, every GGUF export proceeds through three distinct stages implemented in src/soup_cli/utils/gguf_quant.py.

Stage 1: Convert HF Checkpoint to FP16 GGUF

The export_advanced_gguf function calls _run_convert_to_f16, which executes llama.cpp/convert_hf_to_gguf.py to produce a temporary f16.gguf file【gguf_quant.py L559-L580】.

This intermediate file preserves full precision before quantization and is automatically cleaned up after successful export (unless tests like test_issue144_gguf_export.py guard against accidental deletion).

Stage 2: Generate Importance Matrix (Optional)

For UD ladder flavours (e.g., UD-Q4_K_XL) and low-bit IQ types, an importance matrix improves quantization quality. The pipeline:

Flavours without UD- or IQ- prefixes skip this stage.

Stage 3: Quantize to Final GGUF

The _run_quantize_binary function builds command arguments for llama-quantize (or legacy quantize), applying your chosen flavour and imatrix when present【gguf_quant.py L122-L138】. Output writes directly to your specified path.

Choosing a GGUF Flavour

Soup organizes flavours into three categories defined in ALL_ADVANCED_GGUF_FORMATS【gguf_quant.py L70-L75】:

Category Examples Needs Imatrix?
UD ladder (Ultra-Low Distortion) UD-Q8_K_XL, UD-Q5_K_M, UD-Q4_K_XL Yes
IQ family (Information-Retaining Quantization) IQ1_S, IQ2_XXS, IQ3_M Yes
Standard/ARM Q4_0_4_4, Q5_1, Q8_0 No

Complete Export Workflow

1. Prepare Calibration Data (If Required)

Create a JSON-L file for UD or IQ flavours. Each line needs a text, prompt, or content field:

{"text": "The transformer architecture revolutionized natural language processing by introducing self-attention mechanisms..."}
{"text": "Quantization reduces model size by representing weights with fewer bits while preserving task performance..."}

2. Run the Export Command

soup export \
    --model /path/to/huggingface/checkpoint \
    --format gguf-ud \
    --gguf-flavour UD-Q4_K_XL \
    --calibration-data calibration.jsonl \
    --auto-deploy ollama \
    --deploy-name my-custom-llama

The command flow:

  1. Validates the flavour in _export_gguf_advanced (src/soup_cli/commands/export.py)【export.py L14-L34】
  2. Executes the three-stage pipeline with progress reporting
  3. Triggers _auto_deploy_ollama which:
    • Detects Ollama binary via detect_ollama【ollama.py L123-L136】
    • Validates model name through validate_model_name
    • Creates Modelfile and runs ollama create via deploy_to_ollama【ollama.py L208-L230】

3. Use Your Model

ollama run my-custom-llama

Programmatic Export Without CLI

Import export_advanced_gguf directly for custom workflows:

from soup_cli.utils.gguf_quant import export_advanced_gguf

# Standard ARM flavour—no calibration needed

export_advanced_gguf(
    model_dir="meta-llama/Llama-2-7b-hf",
    output_path="llama-2-7b.Q4_0_4_4.gguf",
    flavour="Q4_0_4_4",
    calibration_data=None,
    llama_cpp_dir="~/.soup/llama.cpp",
)

Manual Ollama Deployment

For existing GGUF files, use deploy_to_ollama from src/soup_cli/utils/ollama.py:

from soup_cli.utils.ollama import deploy_to_ollama

modelfile = """FROM ./my_model.Q4_0_4_4.gguf
PARAMETER temperature 0.7
PARAMETER top_p 0.9
SYSTEM You are a helpful coding assistant."""

success, message = deploy_to_ollama("code-assistant", modelfile)
print(message)

Security and Path Handling

All file operations pass through enforce_under_cwd_and_no_symlink in src/soup_cli/utils/paths.py, preventing path-traversal and symlink attacks during the export process.

Key Source Files Reference

File Purpose
src/soup_cli/utils/gguf_quant.py Core three-stage export, flavour validation, calibration parsing
src/soup_cli/commands/export.py CLI command entry point _export_gguf_advanced
src/soup_cli/utils/ollama.py Ollama detection, Modelfile generation, deployment
src/soup_cli/utils/paths.py Security-hardened path utilities
tests/test_issue144_gguf_export.py Guards against f16 file deletion
tests/test_v0531_139.py Integration tests for full pipeline

Summary

  • Three-stage pipeline: Convert → Imatrix (optional) → Quantize, all orchestrated by export_advanced_gguf
  • CLI entry: soup export --format gguf-ud --gguf-flavour <flavour> --auto-deploy ollama
  • Flavour selection: UD and IQ types need calibration data; standard ARM types do not
  • Security: Path traversal protection via enforce_under_cwd_and_no_symlink
  • Deployment: Automatic Ollama registration with --auto-deploy ollama flag

Frequently Asked Questions

What calibration data format does Soup expect for GGUF export?

Soup accepts JSON-L files where each line contains a JSON object with a text, prompt, or content field. The _prepare_calibration_text function in gguf_quant.py extracts plain text from these fields for the llama-imatrix tool【gguf_quant.py L71-L84】.

Do all GGUF flavours require an importance matrix?

No. Only UD ladder (UD-Q4_K_XL, UD-Q5_K_M, etc.) and IQ family (IQ1_S, IQ2_XXS, etc.) flavours require an imatrix. Standard quantizations like Q4_0_4_4 or Q8_0 skip imatrix generation entirely, making them faster to export.

What happens to the intermediate f16.gguf file?

The temporary FP16 file created during Stage 1 is automatically removed after successful quantization. The test suite in test_issue144_gguf_export.py specifically protects against accidental deletion of user-generated files that might share similar naming patterns.

Can I deploy to Ollama without using the --auto-deploy flag?

Yes. After export, call deploy_to_ollama programmatically with a custom Modelfile, or manually create a Modelfile pointing to your GGUF and run ollama create <name> -f Modelfile. The deploy_to_ollama function in ollama.py handles the binary detection and command execution for you【ollama.py L208-L230】.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →