Supported GGUF Tensor Layouts in Dwarf Star: Why Arbitrary GGUF Files Are Rejected

Dwarf Star (DS4) only accepts GGUF files that match the exact DeepSeek V4 Flash tensor layout, with hardcoded validation for specific quantization types and shapes in ds4.c lines 11115-11284.

Dwarf Star is a specialized inference engine purpose-built for a single model architecture. Unlike general-purpose GGUF loaders, DS4 embeds the DeepSeek V4 Flash specification directly into its codebase. Understanding these supported GGUF tensor layouts explains why the engine cannot load arbitrary checkpoints.

Exact Tensor Layout Requirements

DS4 validates every tensor against fixed specifications during the loading phase. The following table lists the required GGUF tensor layouts as enforced by the validation code:

Tensor Shape Requirements GGUF Type (from gguf-tools/quants.h)
blk.N.ffn_gate_exps.weight [expert_in_dim, expert_out_dim] where expert_in_dim % QK_K == 0 DS4Q_TYPE_Q2_K (value 10)
blk.N.ffn_up_exps.weight [expert_in_dim, expert_out_dim] where expert_in_dim % QK_K == 0 DS4Q_TYPE_Q2_K
blk.N.ffn_down_exps.weight [down_in_dim, down_out_dim] where down_in_dim % QK_K == 0 DS4Q_TYPE_Q2_K
Attention Q/K/V projections (q_proj, k_proj, v_proj) [dim, dim] DS4Q_TYPE_Q4_0, DS4Q_TYPE_Q4_K, or DS4Q_TYPE_Q8_0
Shared expert FFN (ffn.shared) [dim, dim] Same as attention: Q4_0, Q4_K, or Q8_0
Token and positional embeddings [vocab_size, emb_dim] DS4Q_TYPE_F32 or DS4Q_TYPE_F16

All remaining tensors—layer normalization parameters, output projections, and bias vectors—must match exact scalar or vector dimensions with no flexibility.

When validation fails, DS4 terminates immediately with descriptive errors such as ds4_die("routed expert tensor layout is unexpected").

Why Dwarf Star Rejects Arbitrary GGUF Files

Four architectural decisions prevent DS4 from supporting general GGUF compatibility:

Hard-Wired Model Architecture

DS4 embeds 43 transformer layers, 256 MoE experts per layer, and three expert tensors per layer directly into its data structures. The ds4.h header defines these constants; the inference graph, memory pool allocation, and pipeline scheduling all assume these exact values. A GGUF with 40 layers or 128 experts would corrupt the memory layout.

Static Kernel Code

Metal and CUDA kernels compile with fixed stride and block size assumptions. For example, kernel_concat in metal/concat.metal uses threadgroup dimensions derived from the Q2_K block size. Changing the quantization scheme would require kernel recompilation, not runtime parameter adjustment.

Zero-Copy MMAP Strategy

The loader in ds4.c creates no-copy MTLBuffer references that point directly into memory-mapped GGUF files. This eliminates GPU upload overhead but requires that tensor offsets and alignment match the pre-computed layout exactly. Arbitrary GGUF files would misalign the mapping, causing GPU page faults.

Safety-First Validation

Early abort validation (lines 11115-11284) ensures that mismatched files cannot trigger undefined behavior, memory overruns, or silent numerical errors. This conservative approach prioritizes correctness over compatibility.

Code-Level Validation Example

The following C snippet demonstrates how DS4 parses CLI arguments and validates the GGUF layout during model loading:

/* Load a GGUF checkpoint with the DS4 CLI */
int main(int argc, char **argv) {
    ds4_opts opts = {0};
    ds4_parse_cli(argc, argv, &opts);          // parses --model=path/to/model.gguf
    ds4_model *model = ds4_load(&opts);        // validates tensor layout in ds4.c
    ds4_chat(model, "Explain quantum tunnelling.");
    ds4_free(model);
    return 0;
}

The ds4_load() function calls into the validation routine that checks each tensor name, dimensions, and quantization type against the tables above.

Metal Kernel Layout Dependency

Metal kernels consume GGUF type IDs through constants defined in metal/dsv4_misc.metal:

kernel void kernel_concat(
        constant ds4_metal_args_concat & args,
        device  const char * src0,
        device  const char * src1,
        device        char * dst,
        uint3   tgpig[[threadgroup_position_in_grid]],
        ushort3 tpitg[[thread_position_in_threadgroup]],
        ushort3   ntg[[threads_per_threadgroup]]) {
    // args.ne0-ne3 filled from GGUF; assumed to match Q2_K/Q4_0/Q4_K/Q8_0 block sizes
}

The constants DS4_METAL_GGUF_Q4_0, DS4_METAL_GGUF_Q4_K, and DS4_METAL_GGUF_Q8_0 map directly to the supported attention and shared-expert quantization types.

Command-Line Usage

DS4 only accepts GGUF files that pass all layout checks:


# Valid: deepseek-v4-flash.gguf matches all tensor specifications

./ds4 --model=deepseek-v4-flash.gguf --prompt="Write a short poem."

# Invalid: arbitrary GGUF fails validation with layout error

./ds4 --model=custom-model.gguf --prompt="Test"

# Output: ds4_die: routed expert tensor layout is unexpected

Key Source Files

File Purpose Relevant Content
ds4.c GGUF loader and layout validation Lines 11115-11284: tensor shape/type verification
ds4.h Model structure definitions Fixed layer count, expert count, weight structs
gguf-tools/quants.h Quantization type enumeration DS4Q_TYPE_Q2_K, DS4Q_TYPE_Q4_0, etc.
metal/dsv4_misc.metal Metal type constants DS4_METAL_GGUF_Q4_0, DS4_METAL_GGUF_Q4_K, DS4_METAL_GGUF_Q8_0

Summary

  • Fixed tensor layouts: DS4 validates GGUF files against exact shapes and quantization types for MoE experts (Q2_K), attention (Q4_0/Q4_K/Q8_0), and embeddings (F32/F16).
  • Single-model architecture: 43 layers, 256 experts, three expert tensors per layer are hardcoded in ds4.h.
  • Zero-copy constraints: Memory-mapped buffers require precise alignment; arbitrary offsets break GPU mapping.
  • Validation-first design: Early abort prevents silent errors, rejecting any GGUF that deviates from DeepSeek V4 Flash specifications.

Frequently Asked Questions

What happens if I try to load a Llama or Qwen GGUF file in DS4?

Dwarf Star immediately terminates with a layout validation error. The ds4_load() function checks tensor names against the DeepSeek V4 Flash schema; Llama's single-FFN architecture lacks the *_exps.weight tensors that DS4 requires, triggering ds4_die before any GPU allocation occurs.

Can I convert my custom model to DS4-compatible GGUF format?

No. The DS4 inference graph assumes specific attention head dimensions, expert routing logic, and layer normalization placements from DeepSeek V4 Flash. Even with identical tensor names and compatible quantization, the architectural differences would produce incorrect outputs or crashes.

Why does DS4 use Q2_K for experts but Q4_0/Q4_K/Q8_0 for attention?

The MoE expert tensors dominate model size (256 experts × 43 layers), so Q2_K provides maximum compression with acceptable quality loss for the FFN pathway. Attention tensors are smaller and more sensitive to quantization error; Q4_0/Q4_K/Q8_0 preserve accuracy for the critical Q/K/V projections.

Is DS4 open to supporting additional GGUF layouts in the future?

According to the repository structure, additional layouts would require: (1) dynamic kernel generation or JIT compilation to replace static Metal/CUDA kernels, (2) abandoning zero-copy MMAP for buffered GPU uploads, and (3) generalizing the validation logic beyond hardcoded constants. These changes would fundamentally alter DS4's design philosophy as a minimal, single-model engine.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →