# How to Handle Model Loading Failures with Partial Layer Distribution in DS4

> Learn how DS4 handles model loading failures by falling back layers to CPU and continuing the load. Discover resilient mixed GPU/CPU placement for uninterrupted model initialization.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-08

---

**The DS4 runtime handles GPU allocation failures by atomically marking the transport as failed, automatically falling back offending layers to the CPU tier, and continuing to load remaining tensors, enabling resilient model initialization with mixed GPU/CPU placement rather than aborting the entire load.**

When distributing large language models across heterogeneous hardware, individual GPU allocation errors must not collapse the entire inference pipeline. The **antirez/ds4** repository implements a resilient loading strategy that allows developers to handle model loading failures with partial layer distribution by spilling failed GPU layers to the CPU tier while maintaining a consistent runtime state.

## Understanding the Layer Distribution Architecture

During initialization, `ds4_engine_open` orchestrates the model loading process by first *pricing* each layer through the **layer-packer** (`ds4_compute_layer_placement`). This algorithm calculates the optimal device assignment for every tensor across available GPUs and the CPU tier. The runtime then attempts to allocate and load each tensor onto its designated device.

If a GPU cannot allocate a specific layer—due to out-of-memory conditions, driver errors, or corrupted tensors—the DS4 runtime does not abort the entire model load. Instead, it executes a five-stage recovery mechanism that preserves partial progress and guarantees a consistent placement state.

## Five-Stage Failure Recovery Mechanism

### 1. Marking the Transport as Failed

When a GPU allocation fails, the transport (TP) layer immediately records a failure flag. This atomic operation ensures that any subsequent operation on that peer can abort early without corrupting the engine state.

In [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c), the `ds4_tp_mark_failed` function sets the failure flag using explicit memory ordering:

```c
void ds4_tp_mark_failed(ds4_tp *tp) {
    if (tp) atomic_store_explicit(&tp->failed, true,
                                 memory_order_release);
}

```

This flag prevents cascading errors across the distributed layer topology.

### 2. Falling Back to the CPU Tier

The placement algorithm in [`ds4_layer_pack.c`](https://github.com/antirez/ds4/blob/main/ds4_layer_pack.c) automatically reassigns any layer that cannot fit on a GPU to the CPU tier. During the budget allocation loop in `ds4_compute_layer_placement`, the logic checks device capacity and spills to `DS4_LAYER_PACK_CPU` when exhausted:

```c
if (d < cfg->n_gpus) {
    device_for_entry[e] = d;
    budget[d] -= entry_bytes[e];
} else {
    device_for_entry[e] = DS4_LAYER_PACK_CPU;
}

```

This fallback occurs transparently during the placement computation, ensuring the layer has a valid target device before loading begins.

### 3. Continuing Layer Loading

After handling the failed allocation, the runtime continues loading subsequent tensors onto their assigned devices. Errors from the GPU side propagate to the caller through the standard error-code path (`return -1`), while the CPU-side load proceeds normally. This partial progression ensures that successful allocations are preserved rather than rolled back.

### 4. Exposing Failure Status to Callers

After the engine opens, applications can query whether any transport failure occurred during loading. The `ds4_tp_failed` function in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) provides atomic read access to the failure state:

```c
bool ds4_tp_failed(const ds4_tp *tp) {
    return tp && atomic_load_explicit(&tp->failed,
                                      memory_order_acquire);
}

```

This allows the caller to distinguish between a fully GPU-loaded model and one with CPU spillover.

### 5. Enabling Retry or Graceful Degradation

Because the CPU tier is always available (subject to system RAM), the model remains usable even with partial GPU failures. Applications may choose to:

- **Accept degraded performance**: Continue inference with slower CPU execution for spilled layers.
- **Retry with reduced budgets**: Close the engine, reduce per-GPU memory budgets, recompute placement, and reopen.
- **Adjust batch size**: Reload with smaller batch dimensions to fit within GPU constraints.

## Practical Implementation Example

The following pattern demonstrates how to open a model while handling potential GPU allocation failures gracefully:

```c
/*--- Example: Open a model with graceful handling of partial GPU failures ---*/
ds4_engine_opt opt = {0};
opt.model_path = "models/phi-2.gguf";

/* 1️⃣ Compute a placement that may spill to CPU */
int device_for_entry[DS4_MAX_LAYERS];
int rc = ds4_compute_layer_placement(entry_bytes,
                                    n_entries,
                                    &layer_cfg,
                                    device_for_entry);
if (rc) {
    fprintf(stderr, "placement computation failed: %d\n", rc);
    exit(1);
}

/* 2️⃣ Open the engine – it will try to load each layer onto its device */
ds4_engine *engine = NULL;
rc = ds4_engine_open(&engine, &opt);
if (rc) {
    /* The engine reports a failure only if the *entire* load aborts.
       Partial GPU failures are already handled by the placement logic. */
    fprintf(stderr, "engine open failed: %s\n", ds4_error());
    exit(1);
}

/* 3️⃣ After opening, check if any GPU load failed */
if (ds4_tp_failed(engine->tp)) {
    fprintf(stderr,
            "Warning: one or more layers could not be loaded on GPU; "
            "they were placed on the CPU tier.\n");
}

/* 4️⃣ Continue using the engine – inference will run with mixed GPU/CPU layers */

```

### Manual Retry After Failure

To retry loading with adjusted constraints after detecting a failure:

```c
/*--- Example: Manually retry a failed GPU load --------------------------------*/
if (ds4_tp_failed(tp)) {
    /* Reduce the per‑GPU budget and recompute placement */
    for (int d = 0; d < cfg.n_gpus; d++) cfg.gpu_budget_bytes[d] /= 2;
    ds4_compute_layer_placement(entry_bytes, n_entries, &cfg, device_for_entry);
    /* Re‑open the engine (or reload the failing layers) */
    ds4_engine_close(engine);
    rc = ds4_engine_open(&engine, &opt);
}

```

## Key Source Files and APIs

The partial layer distribution failure handling relies on coordination between these components:

- **[`ds4_layer_pack.c`](https://github.com/antirez/ds4/blob/main/ds4_layer_pack.c)**: Implements the placement algorithm (`ds4_compute_layer_placement`) that assigns layers to devices and spills to CPU when GPUs are exhausted.
- **[`ds4_layer_pack.h`](https://github.com/antirez/ds4/blob/main/ds4_layer_pack.h)**: Public API header defining `DS4_LAYER_PACK_CPU` and placement configuration structures.
- **[`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c)**: Transport layer implementing `ds4_tp_mark_failed()` and `ds4_tp_failed()` for atomic failure state management.
- **[`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)**: Engine initialization logic (`ds4_engine_open`) that orchestrates loading according to the placement plan.
- **[`tests/test_layer_pack.c`](https://github.com/antirez/ds4/blob/main/tests/test_layer_pack.c)**: Unit tests validating partial spill behavior, including `test_partial_spill_n2` for multi-GPU fallback scenarios.

## Summary

- **Automatic CPU fallback**: The layer-packer in [`ds4_layer_pack.c`](https://github.com/antirez/ds4/blob/main/ds4_layer_pack.c) assigns failed GPU layers to `DS4_LAYER_PACK_CPU` without aborting the load.
- **Atomic failure tracking**: Transport failures are recorded via `ds4_tp_mark_failed()` using `memory_order_release` semantics in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c).
- **Status querying**: Callers detect partial failures after initialization using `ds4_tp_failed()` with `memory_order_acquire`.
- **Continuous loading**: The runtime preserves successful allocations and continues loading remaining layers rather than rolling back.
- **Flexible recovery**: Applications can retry with reduced GPU budgets or accept mixed GPU/CPU inference for graceful degradation.

## Frequently Asked Questions

### What happens if a GPU runs out of memory during layer loading in DS4?

The layer-packer automatically assigns the specific layer to the CPU tier (`DS4_LAYER_PACK_CPU`) and continues loading other layers. The transport layer records the failure via `ds4_tp_mark_failed()`, allowing `ds4_engine_open` to succeed with partial GPU placement rather than returning an error.

### How can I detect if any layers failed to load on GPU in DS4?

After calling `ds4_engine_open()`, query the transport state using `ds4_tp_failed(engine->tp)`. If this returns `true`, one or more layers were placed on the CPU tier due to GPU allocation failures, and inference will run with mixed device placement.

### Can I retry loading layers on GPU after a partial failure in DS4?

Yes. Close the engine with `ds4_engine_close()`, reduce the per-GPU budget in the layer configuration (e.g., `cfg.gpu_budget_bytes[d] /= 2`), recompute placement with `ds4_compute_layer_placement()`, and reopen the engine. Alternatively, you may accept the CPU fallback for functional but slower inference.

### Does a single GPU failure corrupt the entire model state in DS4?

No. The runtime isolates failures per layer using atomic flags in the transport layer. The placement algorithm guarantees a consistent distribution across remaining GPUs and CPU without aborting the load, ensuring that a single device failure does not compromise the entire model initialization.