How to Handle Model Loading Failures with Partial Layer Distribution in DS4

The DS4 runtime handles GPU allocation failures by atomically marking the transport as failed, automatically falling back offending layers to the CPU tier, and continuing to load remaining tensors, enabling resilient model initialization with mixed GPU/CPU placement rather than aborting the entire load.

When distributing large language models across heterogeneous hardware, individual GPU allocation errors must not collapse the entire inference pipeline. The antirez/ds4 repository implements a resilient loading strategy that allows developers to handle model loading failures with partial layer distribution by spilling failed GPU layers to the CPU tier while maintaining a consistent runtime state.

Understanding the Layer Distribution Architecture

During initialization, ds4_engine_open orchestrates the model loading process by first pricing each layer through the layer-packer (ds4_compute_layer_placement). This algorithm calculates the optimal device assignment for every tensor across available GPUs and the CPU tier. The runtime then attempts to allocate and load each tensor onto its designated device.

If a GPU cannot allocate a specific layer—due to out-of-memory conditions, driver errors, or corrupted tensors—the DS4 runtime does not abort the entire model load. Instead, it executes a five-stage recovery mechanism that preserves partial progress and guarantees a consistent placement state.

Five-Stage Failure Recovery Mechanism

1. Marking the Transport as Failed

When a GPU allocation fails, the transport (TP) layer immediately records a failure flag. This atomic operation ensures that any subsequent operation on that peer can abort early without corrupting the engine state.

In ds4_tp.c, the ds4_tp_mark_failed function sets the failure flag using explicit memory ordering:

void ds4_tp_mark_failed(ds4_tp *tp) {
    if (tp) atomic_store_explicit(&tp->failed, true,
                                 memory_order_release);
}

This flag prevents cascading errors across the distributed layer topology.

2. Falling Back to the CPU Tier

The placement algorithm in ds4_layer_pack.c automatically reassigns any layer that cannot fit on a GPU to the CPU tier. During the budget allocation loop in ds4_compute_layer_placement, the logic checks device capacity and spills to DS4_LAYER_PACK_CPU when exhausted:

if (d < cfg->n_gpus) {
    device_for_entry[e] = d;
    budget[d] -= entry_bytes[e];
} else {
    device_for_entry[e] = DS4_LAYER_PACK_CPU;
}

This fallback occurs transparently during the placement computation, ensuring the layer has a valid target device before loading begins.

3. Continuing Layer Loading

After handling the failed allocation, the runtime continues loading subsequent tensors onto their assigned devices. Errors from the GPU side propagate to the caller through the standard error-code path (return -1), while the CPU-side load proceeds normally. This partial progression ensures that successful allocations are preserved rather than rolled back.

4. Exposing Failure Status to Callers

After the engine opens, applications can query whether any transport failure occurred during loading. The ds4_tp_failed function in ds4_tp.c provides atomic read access to the failure state:

bool ds4_tp_failed(const ds4_tp *tp) {
    return tp && atomic_load_explicit(&tp->failed,
                                      memory_order_acquire);
}

This allows the caller to distinguish between a fully GPU-loaded model and one with CPU spillover.

5. Enabling Retry or Graceful Degradation

Because the CPU tier is always available (subject to system RAM), the model remains usable even with partial GPU failures. Applications may choose to:

  • Accept degraded performance: Continue inference with slower CPU execution for spilled layers.
  • Retry with reduced budgets: Close the engine, reduce per-GPU memory budgets, recompute placement, and reopen.
  • Adjust batch size: Reload with smaller batch dimensions to fit within GPU constraints.

Practical Implementation Example

The following pattern demonstrates how to open a model while handling potential GPU allocation failures gracefully:

/*--- Example: Open a model with graceful handling of partial GPU failures ---*/
ds4_engine_opt opt = {0};
opt.model_path = "models/phi-2.gguf";

/* 1️⃣ Compute a placement that may spill to CPU */
int device_for_entry[DS4_MAX_LAYERS];
int rc = ds4_compute_layer_placement(entry_bytes,
                                    n_entries,
                                    &layer_cfg,
                                    device_for_entry);
if (rc) {
    fprintf(stderr, "placement computation failed: %d\n", rc);
    exit(1);
}

/* 2️⃣ Open the engine – it will try to load each layer onto its device */
ds4_engine *engine = NULL;
rc = ds4_engine_open(&engine, &opt);
if (rc) {
    /* The engine reports a failure only if the *entire* load aborts.
       Partial GPU failures are already handled by the placement logic. */
    fprintf(stderr, "engine open failed: %s\n", ds4_error());
    exit(1);
}

/* 3️⃣ After opening, check if any GPU load failed */
if (ds4_tp_failed(engine->tp)) {
    fprintf(stderr,
            "Warning: one or more layers could not be loaded on GPU; "
            "they were placed on the CPU tier.\n");
}

/* 4️⃣ Continue using the engine – inference will run with mixed GPU/CPU layers */

Manual Retry After Failure

To retry loading with adjusted constraints after detecting a failure:

/*--- Example: Manually retry a failed GPU load --------------------------------*/
if (ds4_tp_failed(tp)) {
    /* Reduce the per‑GPU budget and recompute placement */
    for (int d = 0; d < cfg.n_gpus; d++) cfg.gpu_budget_bytes[d] /= 2;
    ds4_compute_layer_placement(entry_bytes, n_entries, &cfg, device_for_entry);
    /* Re‑open the engine (or reload the failing layers) */
    ds4_engine_close(engine);
    rc = ds4_engine_open(&engine, &opt);
}

Key Source Files and APIs

The partial layer distribution failure handling relies on coordination between these components:

  • ds4_layer_pack.c: Implements the placement algorithm (ds4_compute_layer_placement) that assigns layers to devices and spills to CPU when GPUs are exhausted.
  • ds4_layer_pack.h: Public API header defining DS4_LAYER_PACK_CPU and placement configuration structures.
  • ds4_tp.c: Transport layer implementing ds4_tp_mark_failed() and ds4_tp_failed() for atomic failure state management.
  • ds4.c: Engine initialization logic (ds4_engine_open) that orchestrates loading according to the placement plan.
  • tests/test_layer_pack.c: Unit tests validating partial spill behavior, including test_partial_spill_n2 for multi-GPU fallback scenarios.

Summary

  • Automatic CPU fallback: The layer-packer in ds4_layer_pack.c assigns failed GPU layers to DS4_LAYER_PACK_CPU without aborting the load.
  • Atomic failure tracking: Transport failures are recorded via ds4_tp_mark_failed() using memory_order_release semantics in ds4_tp.c.
  • Status querying: Callers detect partial failures after initialization using ds4_tp_failed() with memory_order_acquire.
  • Continuous loading: The runtime preserves successful allocations and continues loading remaining layers rather than rolling back.
  • Flexible recovery: Applications can retry with reduced GPU budgets or accept mixed GPU/CPU inference for graceful degradation.

Frequently Asked Questions

What happens if a GPU runs out of memory during layer loading in DS4?

The layer-packer automatically assigns the specific layer to the CPU tier (DS4_LAYER_PACK_CPU) and continues loading other layers. The transport layer records the failure via ds4_tp_mark_failed(), allowing ds4_engine_open to succeed with partial GPU placement rather than returning an error.

How can I detect if any layers failed to load on GPU in DS4?

After calling ds4_engine_open(), query the transport state using ds4_tp_failed(engine->tp). If this returns true, one or more layers were placed on the CPU tier due to GPU allocation failures, and inference will run with mixed device placement.

Can I retry loading layers on GPU after a partial failure in DS4?

Yes. Close the engine with ds4_engine_close(), reduce the per-GPU budget in the layer configuration (e.g., cfg.gpu_budget_bytes[d] /= 2), recompute placement with ds4_compute_layer_placement(), and reopen the engine. Alternatively, you may accept the CPU fallback for functional but slower inference.

Does a single GPU failure corrupt the entire model state in DS4?

No. The runtime isolates failures per layer using atomic flags in the transport layer. The placement algorithm guarantees a consistent distribution across remaining GPUs and CPU without aborting the load, ensuring that a single device failure does not compromise the entire model initialization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →