LlamaFactory Inference Backends: The Complete Guide to Engine Selection
LlamaFactory supports four inference backends—HuggingFace Transformers, vLLM, SGLang, and K-Transformers—selected via the infer_backend parameter that maps to the EngineName enumeration.
LlamaFactory provides flexible model serving through multiple inference engines optimized for different deployment scenarios. Whether you need the compatibility of standard Transformers or the throughput of vLLM, the framework abstracts backend complexity behind a unified ChatModel interface. This guide examines the four supported LlamaFactory inference backends and their implementation across the codebase.
The Four LlamaFactory Inference Backends
The framework defines supported engines in the EngineName enumeration located in src/llamafactory/extras/constants.py【/cache/repos/github.com/hiyouga/LlamaFactory/main/src/llamafactory/extras/constants.py#L22-L27】. Each backend serves distinct operational requirements:
| Backend | EngineName Value | Description | Requirements |
|---|---|---|---|
| HuggingFace | EngineName.HF ("huggingface") |
Standard 🤗 Transformers pipeline. Default option compatible with any model loadable into RAM/VRAM. | transformers |
| vLLM | EngineName.VLLM ("vllm") |
High‑performance serving engine with tensor parallelism and flash attention. Optimized for throughput. | pip install vllm |
| SGLang | EngineName.SGLANG ("sglang") |
Efficient token‑streaming backend with tool‑call support and structured generation. | pip install sglang[all] |
| K‑Transformers | EngineName.KT ("ktransformers") |
Inference engine for extremely large models using heterogeneous computing (CPU/GPU offloading). | pip install ktransformers |
How Backend Selection Works in LlamaFactory
The architecture routes backend selection through three key components:
-
Constant Definition:
src/llamafactory/extras/constants.pydefines theEngineNameenum with values"huggingface","vllm","sglang", and"ktransformers"【/cache/repos/github.com/hiyouga/LlamaFactory/main/src/llamafactory/extras/constants.py#L22-L27】. -
Argument Parsing:
src/llamafactory/hparams/model_args.pydeclaresinfer_backend: EngineName = EngineName.HFas a dataclass field, defaulting to HuggingFace【/cache/repos/github.com/hiyouga/LlamaFactory/main/src/llamafactory/hparams/model_args.py#L164-L166】. -
Engine Instantiation:
src/llamafactory/chat/chat_model.pyimplements the factory logic inChatModel.__init__. Lines 50‑78 inspectmodel_args.infer_backendand import the corresponding engine class (HuggingfaceEngine,VllmEngine,SGLangEngine, orKTransformersEngine)【/cache/repos/github.com/hiyouga/LlamaFactory/main/src/llamafactory/chat/chat_model.py#L50-L78】.
Using Different Backends in Practice
Programmatic API Usage
You can specify the backend when constructing a ChatModel instance:
from llamafactory.chat.chat_model import ChatModel
# Configuration dict mapping to ModelArguments
args = {
"model_name_or_path": "meta-llama/Meta-Llama-3-8B",
"infer_backend": "vllm", # Options: "huggingface", "vllm", "sglang", "ktransformers"
"infer_dtype": "auto",
}
# Engine instantiation occurs here based on infer_backend value
chat = ChatModel(args)
response = chat.chat([{"role": "user", "content": "Explain quantum computing"}])
print(response[0]["content"])
The ChatModel class dynamically imports the engine implementation based on the infer_backend string value【/cache/repos/github.com/hiyouga/LlamaFactory/main/src/llamafactory/chat/chat_model.py#L50-L78】.
Command-Line Interface
Pass the --infer_backend flag to llamafactory-cli:
# HuggingFace (default)
llamafactory-cli infer \
--model_name_or_path meta-llama/Meta-Llama-3-8B \
--infer_backend huggingface
# vLLM backend
llamafactory-cli infer \
--model_name_or_path meta-llama/Meta-Llama-3-8B \
--infer_backend vllm
# SGLang backend
llamafactory-cli infer \
--model_name_or_path meta-llama/Meta-Llama-3-8B \
--infer_backend sglang
# K-Transformers backend
llamafactory-cli infer \
--model_name_or_path deepseek-ai/DeepSeek-V2 \
--infer_backend ktransformers
The CLI parser validates that non-HuggingFace backends are only used in inference contexts, enforced in src/llamafactory/hparams/parser.py【/cache/repos/github.com/hiyouga/LlamaFactory/main/src/llamafactory/hparams/parser.py#L388-L389】.
Web UI Selection
The Gradio interface exposes backend selection through a dropdown in src/llamafactory/webui/components/infer.py:
import gradio as gr
# Component definition (lines 36-38)
infer_backend = gr.Dropdown(
choices=["huggingface", "vllm", "sglang"], # Note: ktransformers not yet exposed in UI
value="huggingface",
label="Inference Backend"
)
Users select the desired engine from the dropdown before launching inference【/cache/repos/github.com/hiyouga/LlamaFactory/main/src/llamafactory/webui/components/infer.py#L36-L38】.
Key Implementation Files
| File | Purpose | Key Lines |
|---|---|---|
src/llamafactory/extras/constants.py |
Defines EngineName enum with four backend values |
22‑27 |
src/llamafactory/hparams/model_args.py |
Declares infer_backend argument defaulting to EngineName.HF |
164‑166 |
src/llamafactory/chat/chat_model.py |
Factory logic that instantiates specific engine classes | 50‑78 |
src/llamafactory/webui/components/infer.py |
Gradio dropdown for backend selection in Web UI | 36‑38 |
src/llamafactory/hparams/parser.py |
Validation logic ensuring non-HF backends are inference-only | 388‑389 |
Summary
- LlamaFactory supports four inference backends: HuggingFace Transformers, vLLM, SGLang, and K-Transformers.
- Backend selection is controlled by the
infer_backendparameter, defined inModelArgumentsand defaulting to"huggingface". - The
EngineNameenum insrc/llamafactory/extras/constants.pyprovides the canonical identifiers for each backend. ChatModelacts as a factory, dynamically importing and instantiating the appropriate engine class based on the selected backend.- All four backends are available via Python API and CLI, while the Web UI currently exposes HuggingFace, vLLM, and SGLang.
Frequently Asked Questions
What is the default inference backend in LlamaFactory?
HuggingFace Transformers is the default backend. When you do not specify the infer_backend argument, ModelArguments automatically sets it to EngineName.HF (value "huggingface"), which uses the standard 🤗 Transformers pipeline for model loading and generation.
Can I use vLLM or SGLang for training in LlamaFactory?
No, vLLM, SGLang, and K-Transformers are inference-only. The argument parser in src/llamafactory/hparams/parser.py explicitly validates that these backends are only used during inference operations. For training workflows, you must use the HuggingFace backend.
How do I install dependencies for a specific backend?
Each backend requires its own optional dependency. While HuggingFace works with the base transformers installation, you must install:
- vLLM:
pip install vllm - SGLang:
pip install sglang[all] - K-Transformers:
pip install ktransformers
These packages are not included in the base LlamaFactory installation to keep the core footprint minimal.
Why is K-Transformers not available in the Web UI dropdown?
The Web UI currently exposes only HuggingFace, vLLM, and SGLang. The Gradio dropdown component in src/llamafactory/webui/components/infer.py defines choices=["huggingface", "vllm", "sglang"], excluding K-Transformers. Users requiring K-Transformers must use the Python API or CLI interface instead.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →