# How ComfyUI's Checkpoint Loader Supports Different Model Formats (CKPT, Safetensors, Diffusers)

> Discover how ComfyUI's checkpoint loader seamlessly handles CKPT, Safetensors, and Diffusers model formats. Learn about its unified pipeline for efficient model loading.

- Repository: [Comfy Org/ComfyUI](https://github.com/Comfy-Org/ComfyUI)
- Tags: internals
- Published: 2026-02-26

---

**ComfyUI's checkpoint loader automatically detects and normalizes PyTorch `.ckpt`, `.safetensors`, and Diffusers formats through a unified pipeline in [`comfy/sd.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/sd.py) that maps any input to an internal state-dictionary layout.**

The checkpoint loader in Comfy-Org/ComfyUI eliminates format friction by treating classic PyTorch checkpoints, SafeTensors files, and HuggingFace Diffusers repositories as interchangeable inputs. Whether you are loading a legacy Stable Diffusion 1.5 `.ckpt` file or a modern SDXL Diffusers folder, the loader handles format detection, key normalization, and model instantiation without requiring manual conversion. This architecture centers on the `load_checkpoint_guess_config` function family and a robust model detection system that identifies architecture types from weight key patterns.

## File Type Detection and Raw Loading

The entry point for all checkpoint formats is **[`comfy/utils.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/utils.py)**, specifically the `load_torch_file` function. This utility performs extension-based routing to select the appropriate deserialization backend.

For **SafeTensors** files (`.safetensors` or `.sft`), the function uses the `safetensors` library's `safe_open` method to memory-map tensors without executing Python pickle code. For classic **PyTorch checkpoints** (`.ckpt`, `.pt`, `.pth`), it delegates to `torch.load` with `weights_only=True` and optional memory mapping for efficiency.

```python

# comfy/utils.py – load_torch_file

def load_torch_file(ckpt, safe_load=False, device=None, return_metadata=False):
    if ckpt.lower().endswith(".safetensors") or ckpt.lower().endswith(".sft"):
        # SAFETENSORS path – uses safetensors.safe_open

        with safetensors.safe_open(ckpt, framework="pt", device=device.type) as f:
            sd = {k: f.get_tensor(k) for k in f.keys()}
            metadata = f.metadata() if return_metadata else None
        return (sd, metadata) if return_metadata else sd
    else:
        # Classic PyTorch checkpoint – torch.load (with optional mmap)

        pl_sd = torch.load(ckpt, map_location=device, weights_only=True,
                           **({"mmap": True} if MMAP_TORCH_FILES else {}))
        # Handles both “state_dict” wrappers and raw dicts

        sd = pl_sd.get("state_dict", pl_sd) if isinstance(pl_sd, dict) else pl_sd
        return (sd, None) if return_metadata else sd

```

The `return_metadata=True` flag ensures that SafeTensors embedded metadata (such as training tags or model hashes) propagates through the loading pipeline for downstream use.

## State Dictionary Normalization

Once raw tensors are loaded, **[`comfy/sd.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/sd.py)** takes over via `load_state_dict_guess_config`. This function determines the model architecture by analyzing the state dictionary's key prefixes rather than relying on file extensions or external config files.

The normalization process follows two steps:

1. **Prefix Detection**: `model_detection.unet_prefix_from_state_dict(sd)` scans the keys to locate UNet weight groups (e.g., `model.diffusion_model.` for classic checkpoints).
2. **Config Resolution**: `model_detection.model_config_from_unet(sd, diffusion_model_prefix, metadata=metadata)` matches these patterns against a catalog of known architectures including SD 1.x, SD 2.x, SDXL, and Stable Cascade.

```python

# comfy/sd.py – load_state_dict_guess_config

diffusion_model_prefix = model_detection.unet_prefix_from_state_dict(sd)
model_config = model_detection.model_config_from_unet(
                 sd, diffusion_model_prefix, metadata=metadata)

```

This approach allows the checkpoint loader to support different model formats without explicit user intervention, as the tensor key structure itself reveals the model type.

## Diffusers Format Conversion

When the state dictionary does not match any built-in native config, the loader executes a **Diffusers fallback path**. This handles HuggingFace Diffusers checkpoints that use alternative naming conventions such as `transformer.blocks.*` instead of `model.diffusion_model.*`.

The conversion logic in **[`comfy/model_detection.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/model_detection.py)** attempts two strategies:

- **MM-DiT Conversion**: `convert_diffusers_mmdit` handles modern Diffusers formats used by models like SD3.
- **UNet Conversion**: `model_config_from_diffusers_unet` maps legacy Diffusers UNet structures to the internal layout.

```python

# comfy/sd.py – part of load_state_dict_guess_config

if model_config is None:
    # Try Diffusers formats

    diffusion_model = load_diffusion_model_state_dict(sd, model_options={})
    if diffusion_model is None:
        return None
    return (diffusion_model, None, VAE(sd={}), None)

```

The conversion functions strip Diffusers-specific prefixes and restructure the tensor dictionary to match the "original Stable Diffusion" layout expected by ComfyUI's execution engine.

## Component Instantiation

After normalization, the checkpoint loader instantiates four primary components:

- **UNet/Diffusion Model**: Created via `model_config.get_model` and wrapped in a `ModelPatcher` for LoRA and other patching operations.
- **VAE**: Extracted by stripping UNet prefixes from the state dictionary using `state_dict_prefix_replace`, then passed to the `VAE` class.
- **CLIP/Text Encoder**: Constructed through `model_config.clip_target` and the `CLIP` class, handling tokenizer and text model initialization.
- **CLIP Vision**: Optionally loaded for image conditioning tasks.

The "Load Checkpoint" node in **[`nodes.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/nodes.py)** exposes this functionality to the UI, calling `load_checkpoint_guess_config` with standard parameters:

```python

# nodes.py – Load Checkpoint node

def load_checkpoint(self, ckpt_name):
    ckpt_path = folder_paths.get_full_path_or_raise("checkpoints", ckpt_name)
    out = comfy.sd.load_checkpoint_guess_config(
            ckpt_path, output_vae=True, output_clip=True,
            embedding_directory=folder_paths.get_folder_paths("embeddings"))
    return out

```

This node returns a tuple of `(model_patcher, clip, vae, clipvision)`, ready for connection to sampling and encoding nodes regardless of the original file format.

## Practical Code Examples

### Loading a Classic .ckpt File

Use `load_checkpoint_guess_config` directly to load legacy PyTorch checkpoints with automatic device mapping:

```python
from comfy import sd, utils, folder_paths

ckpt_path = folder_paths.get_full_path_or_raise("checkpoints", "sd-v1-5.ckpt")
model, clip, vae, _ = sd.load_checkpoint_guess_config(
        ckpt_path, output_vae=True, output_clip=True)

print("UNet device:", model.load_device)          # e.g. cuda:0

print("CLIP tokenizer vocab size:", clip.tokenizer.vocab_size)
print("VAE latent channels:", vae.latent_channels)

```

The function automatically routes through `utils.load_torch_file`, which invokes `torch.load` for the `.ckpt` extension.

### Loading a .safetensors Checkpoint

SafeTensors loading preserves embedded metadata and avoids Python deserialization risks:

```python
from comfy import sd, folder_paths

ckpt_path = folder_paths.get_full_path_or_raise("checkpoints", "sdxl.safetensors")
model, clip, vae, _ = sd.load_checkpoint_guess_config(
        ckpt_path, output_vae=True, output_clip=True)

# Safe-tensors are read via safetensors.safe_open (see utils.load_torch_file)

print("Loaded from safetensors – metadata keys:", model.patcher.metadata.keys())

```

### Loading a Diffusers Format Checkpoint

For Diffusers repositories or exported single-file checkpoints, the conversion happens transparently:

```python
from comfy import sd, folder_paths

diffusers_path = folder_paths.get_full_path_or_raise(
        "checkpoints", "diffusers_sd_v1_5.safetensors")

model, clip, vae, _ = sd.load_checkpoint_guess_config(
        diffusers_path, output_vae=True, output_clip=True)

print("Diffusers checkpoint converted – UNet config:", model.model.model_config)

```

The loader detects the Diffusers key structure and routes through `convert_diffusers_mmdit` or `model_config_from_diffusers_unet` before instantiation.

## Summary

- **Format Agnostic Entry Point**: The `load_checkpoint_guess_config` function in [`comfy/sd.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/sd.py) serves as the universal interface for all checkpoint types.
- **Extension-Based Routing**: [`comfy/utils.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/utils.py) selects between `safetensors.safe_open` and `torch.load` based on file suffixes.
- **Key-Based Detection**: [`comfy/model_detection.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/model_detection.py) identifies model architectures by analyzing state dictionary key prefixes rather than metadata.
- **Automatic Diffusers Conversion**: Unsupported formats fall back to Diffusers conversion utilities that remap tensor keys to the internal layout.
- **Unified Output**: All formats ultimately produce the same `(model, clip, vae, clipvision)` tuple for consumption by ComfyUI nodes.

## Frequently Asked Questions

### How does ComfyUI detect whether a checkpoint is SafeTensors or PyTorch format?

ComfyUI detects the format by checking the file extension in [`comfy/utils.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/utils.py). If the path ends with `.safetensors` or `.sft`, it uses the `safetensors` library's `safe_open` method; otherwise, it defaults to `torch.load`. This detection happens inside the `load_torch_file` function before any tensor processing begins.

### Can ComfyUI load Diffusers checkpoints directly without manual conversion?

Yes, ComfyUI loads Diffusers checkpoints automatically. When `model_config_from_unet` fails to identify a native architecture, the loader attempts Diffusers conversion via `load_diffusion_model_state_dict` in [`comfy/sd.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/sd.py), which calls conversion functions in [`comfy/model_detection.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/model_detection.py) to remap Diffusers-specific tensor keys to the internal format.

### What components are extracted from a checkpoint during loading?

The checkpoint loader extracts four components: the **UNet/diffusion model** (wrapped in ModelPatcher), the **VAE** for latent encoding/decoding, the **CLIP/text encoder** for prompt processing, and optionally **CLIP Vision** for image conditioning. These are returned as a standardized tuple regardless of the input format.

### Where is the "Load Checkpoint" node defined in the ComfyUI source code?

The "Load Checkpoint" node is defined in **[`nodes.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/nodes.py)** (lines 605-607 in the current master branch). This node calls `comfy.sd.load_checkpoint_guess_config` and handles path resolution through `folder_paths.get_full_path_or_raise`, making the underlying format complexity invisible to end users.