How ComfyUI's Checkpoint Loader Supports Different Model Formats (CKPT, Safetensors, Diffusers)
ComfyUI's checkpoint loader automatically detects and normalizes PyTorch .ckpt, .safetensors, and Diffusers formats through a unified pipeline in comfy/sd.py that maps any input to an internal state-dictionary layout.
The checkpoint loader in Comfy-Org/ComfyUI eliminates format friction by treating classic PyTorch checkpoints, SafeTensors files, and HuggingFace Diffusers repositories as interchangeable inputs. Whether you are loading a legacy Stable Diffusion 1.5 .ckpt file or a modern SDXL Diffusers folder, the loader handles format detection, key normalization, and model instantiation without requiring manual conversion. This architecture centers on the load_checkpoint_guess_config function family and a robust model detection system that identifies architecture types from weight key patterns.
File Type Detection and Raw Loading
The entry point for all checkpoint formats is comfy/utils.py, specifically the load_torch_file function. This utility performs extension-based routing to select the appropriate deserialization backend.
For SafeTensors files (.safetensors or .sft), the function uses the safetensors library's safe_open method to memory-map tensors without executing Python pickle code. For classic PyTorch checkpoints (.ckpt, .pt, .pth), it delegates to torch.load with weights_only=True and optional memory mapping for efficiency.
# comfy/utils.py – load_torch_file
def load_torch_file(ckpt, safe_load=False, device=None, return_metadata=False):
if ckpt.lower().endswith(".safetensors") or ckpt.lower().endswith(".sft"):
# SAFETENSORS path – uses safetensors.safe_open
with safetensors.safe_open(ckpt, framework="pt", device=device.type) as f:
sd = {k: f.get_tensor(k) for k in f.keys()}
metadata = f.metadata() if return_metadata else None
return (sd, metadata) if return_metadata else sd
else:
# Classic PyTorch checkpoint – torch.load (with optional mmap)
pl_sd = torch.load(ckpt, map_location=device, weights_only=True,
**({"mmap": True} if MMAP_TORCH_FILES else {}))
# Handles both “state_dict” wrappers and raw dicts
sd = pl_sd.get("state_dict", pl_sd) if isinstance(pl_sd, dict) else pl_sd
return (sd, None) if return_metadata else sd
The return_metadata=True flag ensures that SafeTensors embedded metadata (such as training tags or model hashes) propagates through the loading pipeline for downstream use.
State Dictionary Normalization
Once raw tensors are loaded, comfy/sd.py takes over via load_state_dict_guess_config. This function determines the model architecture by analyzing the state dictionary's key prefixes rather than relying on file extensions or external config files.
The normalization process follows two steps:
- Prefix Detection:
model_detection.unet_prefix_from_state_dict(sd)scans the keys to locate UNet weight groups (e.g.,model.diffusion_model.for classic checkpoints). - Config Resolution:
model_detection.model_config_from_unet(sd, diffusion_model_prefix, metadata=metadata)matches these patterns against a catalog of known architectures including SD 1.x, SD 2.x, SDXL, and Stable Cascade.
# comfy/sd.py – load_state_dict_guess_config
diffusion_model_prefix = model_detection.unet_prefix_from_state_dict(sd)
model_config = model_detection.model_config_from_unet(
sd, diffusion_model_prefix, metadata=metadata)
This approach allows the checkpoint loader to support different model formats without explicit user intervention, as the tensor key structure itself reveals the model type.
Diffusers Format Conversion
When the state dictionary does not match any built-in native config, the loader executes a Diffusers fallback path. This handles HuggingFace Diffusers checkpoints that use alternative naming conventions such as transformer.blocks.* instead of model.diffusion_model.*.
The conversion logic in comfy/model_detection.py attempts two strategies:
- MM-DiT Conversion:
convert_diffusers_mmdithandles modern Diffusers formats used by models like SD3. - UNet Conversion:
model_config_from_diffusers_unetmaps legacy Diffusers UNet structures to the internal layout.
# comfy/sd.py – part of load_state_dict_guess_config
if model_config is None:
# Try Diffusers formats
diffusion_model = load_diffusion_model_state_dict(sd, model_options={})
if diffusion_model is None:
return None
return (diffusion_model, None, VAE(sd={}), None)
The conversion functions strip Diffusers-specific prefixes and restructure the tensor dictionary to match the "original Stable Diffusion" layout expected by ComfyUI's execution engine.
Component Instantiation
After normalization, the checkpoint loader instantiates four primary components:
- UNet/Diffusion Model: Created via
model_config.get_modeland wrapped in aModelPatcherfor LoRA and other patching operations. - VAE: Extracted by stripping UNet prefixes from the state dictionary using
state_dict_prefix_replace, then passed to theVAEclass. - CLIP/Text Encoder: Constructed through
model_config.clip_targetand theCLIPclass, handling tokenizer and text model initialization. - CLIP Vision: Optionally loaded for image conditioning tasks.
The "Load Checkpoint" node in nodes.py exposes this functionality to the UI, calling load_checkpoint_guess_config with standard parameters:
# nodes.py – Load Checkpoint node
def load_checkpoint(self, ckpt_name):
ckpt_path = folder_paths.get_full_path_or_raise("checkpoints", ckpt_name)
out = comfy.sd.load_checkpoint_guess_config(
ckpt_path, output_vae=True, output_clip=True,
embedding_directory=folder_paths.get_folder_paths("embeddings"))
return out
This node returns a tuple of (model_patcher, clip, vae, clipvision), ready for connection to sampling and encoding nodes regardless of the original file format.
Practical Code Examples
Loading a Classic .ckpt File
Use load_checkpoint_guess_config directly to load legacy PyTorch checkpoints with automatic device mapping:
from comfy import sd, utils, folder_paths
ckpt_path = folder_paths.get_full_path_or_raise("checkpoints", "sd-v1-5.ckpt")
model, clip, vae, _ = sd.load_checkpoint_guess_config(
ckpt_path, output_vae=True, output_clip=True)
print("UNet device:", model.load_device) # e.g. cuda:0
print("CLIP tokenizer vocab size:", clip.tokenizer.vocab_size)
print("VAE latent channels:", vae.latent_channels)
The function automatically routes through utils.load_torch_file, which invokes torch.load for the .ckpt extension.
Loading a .safetensors Checkpoint
SafeTensors loading preserves embedded metadata and avoids Python deserialization risks:
from comfy import sd, folder_paths
ckpt_path = folder_paths.get_full_path_or_raise("checkpoints", "sdxl.safetensors")
model, clip, vae, _ = sd.load_checkpoint_guess_config(
ckpt_path, output_vae=True, output_clip=True)
# Safe-tensors are read via safetensors.safe_open (see utils.load_torch_file)
print("Loaded from safetensors – metadata keys:", model.patcher.metadata.keys())
Loading a Diffusers Format Checkpoint
For Diffusers repositories or exported single-file checkpoints, the conversion happens transparently:
from comfy import sd, folder_paths
diffusers_path = folder_paths.get_full_path_or_raise(
"checkpoints", "diffusers_sd_v1_5.safetensors")
model, clip, vae, _ = sd.load_checkpoint_guess_config(
diffusers_path, output_vae=True, output_clip=True)
print("Diffusers checkpoint converted – UNet config:", model.model.model_config)
The loader detects the Diffusers key structure and routes through convert_diffusers_mmdit or model_config_from_diffusers_unet before instantiation.
Summary
- Format Agnostic Entry Point: The
load_checkpoint_guess_configfunction incomfy/sd.pyserves as the universal interface for all checkpoint types. - Extension-Based Routing:
comfy/utils.pyselects betweensafetensors.safe_openandtorch.loadbased on file suffixes. - Key-Based Detection:
comfy/model_detection.pyidentifies model architectures by analyzing state dictionary key prefixes rather than metadata. - Automatic Diffusers Conversion: Unsupported formats fall back to Diffusers conversion utilities that remap tensor keys to the internal layout.
- Unified Output: All formats ultimately produce the same
(model, clip, vae, clipvision)tuple for consumption by ComfyUI nodes.
Frequently Asked Questions
How does ComfyUI detect whether a checkpoint is SafeTensors or PyTorch format?
ComfyUI detects the format by checking the file extension in comfy/utils.py. If the path ends with .safetensors or .sft, it uses the safetensors library's safe_open method; otherwise, it defaults to torch.load. This detection happens inside the load_torch_file function before any tensor processing begins.
Can ComfyUI load Diffusers checkpoints directly without manual conversion?
Yes, ComfyUI loads Diffusers checkpoints automatically. When model_config_from_unet fails to identify a native architecture, the loader attempts Diffusers conversion via load_diffusion_model_state_dict in comfy/sd.py, which calls conversion functions in comfy/model_detection.py to remap Diffusers-specific tensor keys to the internal format.
What components are extracted from a checkpoint during loading?
The checkpoint loader extracts four components: the UNet/diffusion model (wrapped in ModelPatcher), the VAE for latent encoding/decoding, the CLIP/text encoder for prompt processing, and optionally CLIP Vision for image conditioning. These are returned as a standardized tuple regardless of the input format.
Where is the "Load Checkpoint" node defined in the ComfyUI source code?
The "Load Checkpoint" node is defined in nodes.py (lines 605-607 in the current master branch). This node calls comfy.sd.load_checkpoint_guess_config and handles path resolution through folder_paths.get_full_path_or_raise, making the underlying format complexity invisible to end users.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →