# Understanding Layer Slicing with --layers for Partial Model Loading in ds4

> Master ds4 layer slicing with the --layers flag for efficient partial model loading. Optimize GPU/CPU memory by selectively loading transformer layers and skipping unselected weights.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-08

---

**The `--layers` flag in ds4 enables selective loading of specific transformer layers by parsing a comma-separated list or range, invoking `ds4_compute_layer_placement` to map selected layers across GPU/CPU memory budgets, and skipping unselected weights during the GGUF loading process.**

The `antirez/ds4` inference engine supports partial model loading through layer slicing, allowing you to load only a subset of a model’s transformer layers instead of the full network. This capability is essential for running large models on limited GPU memory or for extracting intermediate representations from early layers. By leveraging the `--layers` CLI option and the underlying `ds4_layer_pack` module, ds4 calculates optimal device placement and selectively loads only the requested weights from GGUF files.

## How --layers Enables Partial Model Loading

When you specify `--layers`, ds4 bypasses the default full-model load path and enters a selective loading mode that involves three distinct phases: CLI parsing, placement calculation, and filtered weight loading.

### CLI Argument Parsing in ds4_cli.c

The command-line interface parses the `--layers` flag in [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c), validating the input as either a comma-separated list of indices or a range specification (e.g., `0:11`). The parser stores the resulting array in `cli_config.layers` and verifies that all requested indices fall within the model’s total layer count.

In [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c), the configuration structure captures the user-defined slice:

```c
/* Parsed from --layers flag */
int *layers;
size_t n_layers;

```

This array is later passed to the engine loader to determine which GGUF tensor blocks to read.

### Layer Placement Calculation via ds4_layer_pack.h

Before loading any weights, ds4 must determine where each selected layer will reside. The header [`ds4_layer_pack.h`](https://github.com/antirez/ds4/blob/main/ds4_layer_pack.h) declares the core function `ds4_compute_layer_placement`, which calculates a monotonic, contiguous mapping of layers to devices based on memory constraints.

The placement algorithm considers:

- **entry_bytes[]**: An array containing the byte size of each layer’s weight block, extracted from the GGUF header metadata.
- **gpu_budget_bytes[]**: Per-device VRAM limits specified via `--gpu-vram` or programmatically in `ds4_engine_options`.
- **Monotonic-contiguous rule**: Once a layer cannot fit into available GPU memory, all subsequent layers are assigned to the CPU tier (`DS4_LAYER_PACK_CPU`).

The function returns a `device_for_entry[]` array that maps each requested layer (including pseudo-layers for embeddings and the output head) to a specific GPU ID or CPU.

### Selective Weight Loading in ds4_engine_load

The engine loader in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) implements the actual partial loading logic. When `ds4_engine_load` receives a non-empty `layers` list, it iterates through the GGUF file sections but only allocates and copies weights for indices present in the selection array.

As implemented in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), the loader:

1. Reads the GGUF header to populate `entry_bytes[]` for all layers.
2. Calls `ds4_compute_layer_placement` to generate the device mapping.
3. Skips tensor blocks whose indices are not in the requested set, avoiding both I/O overhead and memory allocation for unselected layers.
4. Allocates GPU memory via the backend (CUDA/Metal/ROCm) according to the placement array and copies only the selected weights.

## Using --layers from the Command Line

To load only the first 12 layers of a 32-layer model across two GPUs with 12 GB each:

```bash
ds4 --model mymodel.gguf \
    --gpu-vram 12 \
    --layers 0:11 \
    -p "Explain the concept of attention."

```

The range `0:11` instructs ds4 to load the embedding layer (pseudo-layer 0) and transformer layers 1 through 12. The output head pseudo-layer loads automatically because it remains required for token generation, regardless of the slice selection.

## Programmatic Layer Slicing with the C API

You can achieve the same partial loading behavior programmatically without using the CLI. The `ds4_engine_options` structure accepts a `layers` array and `n_layers` count:

```c
#include "ds4.h"
#include "ds4_layer_pack.h"

int main(void) {
    ds4_engine *engine = NULL;
    ds4_engine_options opts = {0};

    /* Define the layer slice: layers 0 through 11 */
    const int wanted[] = {0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11};
    opts.layers = wanted;
    opts.n_layers = sizeof(wanted) / sizeof(wanted[0]);

    /* Set per-GPU VRAM budgets (12 GB each) */
    opts.gpu_budget_bytes[0] = 12ULL << 30;
    opts.gpu_budget_bytes[1] = 12ULL << 30;
    opts.n_gpus = 2;

    /* Load only the requested slice */
    if (ds4_engine_load(&engine, "mymodel.gguf", &opts) == 0) {
        /* Run inference... */
        ds4_engine_free(engine);
    }
    return 0;
}

```

This approach passes the layer selection directly to `ds4_engine_load`, which invokes the same placement and filtering logic used by the command-line interface.

## Architecture of the Layer Packing Algorithm

The `ds4_compute_layer_placement` function in [`ds4_layer_pack.c`](https://github.com/antirez/ds4/blob/main/ds4_layer_pack.c) (declared in [`ds4_layer_pack.h`](https://github.com/antirez/ds4/blob/main/ds4_layer_pack.h)) implements a greedy bin-packing algorithm optimized for transformer architectures. It processes layers sequentially, attempting to fit each layer’s `entry_bytes` into the current GPU’s remaining `gpu_budget_bytes`.

Key characteristics of the algorithm:

- **Contiguity constraint**: The function guarantees that once layers spill to CPU (indicated by `DS4_LAYER_PACK_CPU`), all subsequent layers also reside on CPU, preventing expensive cross-device transfers during inference.
- **Pseudo-layer handling**: The algorithm treats the embedding lookup and output projection as special pseudo-layers that must always load, ensuring the model remains functional even with aggressive slicing.
- **Budget awareness**: If a single layer exceeds any individual GPU’s budget, the function immediately flags it for CPU placement, preventing allocation failures during the actual weight copy phase.

## Summary

- **Partial loading** via `--layers` allows ds4 to load specific transformer indices rather than the full model, reducing GPU memory footprint.
- **CLI parsing** in [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c) converts range syntax (e.g., `0:11`) into an integer array stored in `cli_config.layers`.
- **Placement logic** in [`ds4_layer_pack.h`](https://github.com/antirez/ds4/blob/main/ds4_layer_pack.h) uses `ds4_compute_layer_placement` to map selected layers to devices based on `entry_bytes[]` and `gpu_budget_bytes[]`, respecting a monotonic-contiguous constraint.
- **Selective loading** in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) skips unselected GGUF tensor blocks during `ds4_engine_load`, saving both I/O bandwidth and VRAM.
- The **C API** exposes this functionality through `ds4_engine_options.layers`, enabling embedded applications to implement custom slicing strategies.

## Frequently Asked Questions

### What file formats does ds4 support for partial loading?

ds4 supports partial loading exclusively with **GGUF** (GGML Universal File) format. The implementation relies on the GGUF header metadata to determine `entry_bytes[]` for each layer before loading weights, allowing the engine to skip unwanted tensor blocks efficiently.

### How does ds4 handle the output head when using --layers?

The output head and embedding layers are treated as **pseudo-layers** that load automatically regardless of the `--layers` selection. This ensures the model can still perform token generation and embedding lookup even when you load only a subset of intermediate transformer blocks.

### Can I load non-contiguous layers with --layers?

Yes, the `--layers` flag accepts comma-separated indices (e.g., `--layers 0,2,4,6`) in addition to range syntax. However, the `ds4_compute_layer_placement` algorithm still enforces a **monotonic-contiguous** memory layout on each device, meaning non-contiguous logical selections may result in fragmented memory placement or CPU fallback for gaps.

### What happens if the selected layers exceed available GPU memory?

If the cumulative size of selected layers exceeds the provided `gpu_budget_bytes`, `ds4_compute_layer_placement` assigns overflowing layers to the **CPU tier** (`DS4_LAYER_PACK_CPU`). The engine then loads those specific weights into system RAM, allowing inference to proceed with CPU offloading for the layers that do not fit on the GPU.