Understanding Layer Slicing with --layers for Partial Model Loading in ds4
The --layers flag in ds4 enables selective loading of specific transformer layers by parsing a comma-separated list or range, invoking ds4_compute_layer_placement to map selected layers across GPU/CPU memory budgets, and skipping unselected weights during the GGUF loading process.
The antirez/ds4 inference engine supports partial model loading through layer slicing, allowing you to load only a subset of a model’s transformer layers instead of the full network. This capability is essential for running large models on limited GPU memory or for extracting intermediate representations from early layers. By leveraging the --layers CLI option and the underlying ds4_layer_pack module, ds4 calculates optimal device placement and selectively loads only the requested weights from GGUF files.
How --layers Enables Partial Model Loading
When you specify --layers, ds4 bypasses the default full-model load path and enters a selective loading mode that involves three distinct phases: CLI parsing, placement calculation, and filtered weight loading.
CLI Argument Parsing in ds4_cli.c
The command-line interface parses the --layers flag in ds4_cli.c, validating the input as either a comma-separated list of indices or a range specification (e.g., 0:11). The parser stores the resulting array in cli_config.layers and verifies that all requested indices fall within the model’s total layer count.
In ds4_cli.c, the configuration structure captures the user-defined slice:
/* Parsed from --layers flag */
int *layers;
size_t n_layers;
This array is later passed to the engine loader to determine which GGUF tensor blocks to read.
Layer Placement Calculation via ds4_layer_pack.h
Before loading any weights, ds4 must determine where each selected layer will reside. The header ds4_layer_pack.h declares the core function ds4_compute_layer_placement, which calculates a monotonic, contiguous mapping of layers to devices based on memory constraints.
The placement algorithm considers:
- entry_bytes[]: An array containing the byte size of each layer’s weight block, extracted from the GGUF header metadata.
- gpu_budget_bytes[]: Per-device VRAM limits specified via
--gpu-vramor programmatically inds4_engine_options. - Monotonic-contiguous rule: Once a layer cannot fit into available GPU memory, all subsequent layers are assigned to the CPU tier (
DS4_LAYER_PACK_CPU).
The function returns a device_for_entry[] array that maps each requested layer (including pseudo-layers for embeddings and the output head) to a specific GPU ID or CPU.
Selective Weight Loading in ds4_engine_load
The engine loader in ds4.c implements the actual partial loading logic. When ds4_engine_load receives a non-empty layers list, it iterates through the GGUF file sections but only allocates and copies weights for indices present in the selection array.
As implemented in ds4.c, the loader:
- Reads the GGUF header to populate
entry_bytes[]for all layers. - Calls
ds4_compute_layer_placementto generate the device mapping. - Skips tensor blocks whose indices are not in the requested set, avoiding both I/O overhead and memory allocation for unselected layers.
- Allocates GPU memory via the backend (CUDA/Metal/ROCm) according to the placement array and copies only the selected weights.
Using --layers from the Command Line
To load only the first 12 layers of a 32-layer model across two GPUs with 12 GB each:
ds4 --model mymodel.gguf \
--gpu-vram 12 \
--layers 0:11 \
-p "Explain the concept of attention."
The range 0:11 instructs ds4 to load the embedding layer (pseudo-layer 0) and transformer layers 1 through 12. The output head pseudo-layer loads automatically because it remains required for token generation, regardless of the slice selection.
Programmatic Layer Slicing with the C API
You can achieve the same partial loading behavior programmatically without using the CLI. The ds4_engine_options structure accepts a layers array and n_layers count:
#include "ds4.h"
#include "ds4_layer_pack.h"
int main(void) {
ds4_engine *engine = NULL;
ds4_engine_options opts = {0};
/* Define the layer slice: layers 0 through 11 */
const int wanted[] = {0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11};
opts.layers = wanted;
opts.n_layers = sizeof(wanted) / sizeof(wanted[0]);
/* Set per-GPU VRAM budgets (12 GB each) */
opts.gpu_budget_bytes[0] = 12ULL << 30;
opts.gpu_budget_bytes[1] = 12ULL << 30;
opts.n_gpus = 2;
/* Load only the requested slice */
if (ds4_engine_load(&engine, "mymodel.gguf", &opts) == 0) {
/* Run inference... */
ds4_engine_free(engine);
}
return 0;
}
This approach passes the layer selection directly to ds4_engine_load, which invokes the same placement and filtering logic used by the command-line interface.
Architecture of the Layer Packing Algorithm
The ds4_compute_layer_placement function in ds4_layer_pack.c (declared in ds4_layer_pack.h) implements a greedy bin-packing algorithm optimized for transformer architectures. It processes layers sequentially, attempting to fit each layer’s entry_bytes into the current GPU’s remaining gpu_budget_bytes.
Key characteristics of the algorithm:
- Contiguity constraint: The function guarantees that once layers spill to CPU (indicated by
DS4_LAYER_PACK_CPU), all subsequent layers also reside on CPU, preventing expensive cross-device transfers during inference. - Pseudo-layer handling: The algorithm treats the embedding lookup and output projection as special pseudo-layers that must always load, ensuring the model remains functional even with aggressive slicing.
- Budget awareness: If a single layer exceeds any individual GPU’s budget, the function immediately flags it for CPU placement, preventing allocation failures during the actual weight copy phase.
Summary
- Partial loading via
--layersallows ds4 to load specific transformer indices rather than the full model, reducing GPU memory footprint. - CLI parsing in
ds4_cli.cconverts range syntax (e.g.,0:11) into an integer array stored incli_config.layers. - Placement logic in
ds4_layer_pack.husesds4_compute_layer_placementto map selected layers to devices based onentry_bytes[]andgpu_budget_bytes[], respecting a monotonic-contiguous constraint. - Selective loading in
ds4.cskips unselected GGUF tensor blocks duringds4_engine_load, saving both I/O bandwidth and VRAM. - The C API exposes this functionality through
ds4_engine_options.layers, enabling embedded applications to implement custom slicing strategies.
Frequently Asked Questions
What file formats does ds4 support for partial loading?
ds4 supports partial loading exclusively with GGUF (GGML Universal File) format. The implementation relies on the GGUF header metadata to determine entry_bytes[] for each layer before loading weights, allowing the engine to skip unwanted tensor blocks efficiently.
How does ds4 handle the output head when using --layers?
The output head and embedding layers are treated as pseudo-layers that load automatically regardless of the --layers selection. This ensures the model can still perform token generation and embedding lookup even when you load only a subset of intermediate transformer blocks.
Can I load non-contiguous layers with --layers?
Yes, the --layers flag accepts comma-separated indices (e.g., --layers 0,2,4,6) in addition to range syntax. However, the ds4_compute_layer_placement algorithm still enforces a monotonic-contiguous memory layout on each device, meaning non-contiguous logical selections may result in fragmented memory placement or CPU fallback for gaps.
What happens if the selected layers exceed available GPU memory?
If the cumulative size of selected layers exceeds the provided gpu_budget_bytes, ds4_compute_layer_placement assigns overflowing layers to the CPU tier (DS4_LAYER_PACK_CPU). The engine then loads those specific weights into system RAM, allowing inference to proceed with CPU offloading for the layers that do not fit on the GPU.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →