GGUF Tensor Layout Requirements for Troubleshooting Model Loading Issues in ds4

The ds4 loader enforces strict GGUF tensor layout requirements including 32-byte alignment, reversed dimension ordering, specific quantization block sizes, and valid data offsets, aborting with descriptive errors when any validation check fails.

The antirez/ds4 repository implements a rigorous GGUF (GGML Unified Format) parser that validates every byte of tensor metadata before loading deep learning models. Understanding these GGUF tensor layout requirements is essential when porting models or debugging loading failures, as the codebase performs exhaustive sanity checks on file headers, alignment, rank constraints, and quantization-specific block structures.

File Header and Magic Validation

Every GGUF file must begin with the ASCII magic string "GGUF" followed by a supported version integer. In ds4.c at line 1677, the loader verifies this immediately upon opening the file:

/* Open the file and verify the magic header */
FILE *fp = fopen(path, "rb");
if (!fp) die_errno("open GGUF", path);
char magic[4];
if (fread(magic,1,4,fp)!=4 || memcmp(magic,"GGUF",4)!=0)
    die("bad GGUF template");

/* Version and tensor count */
uint32_t version = read_u32_le_fp(fp, "GGUF version");
uint64_t n_tensors = read_u64_le_fp(fp, "GGUF tensor count");

Following the magic bytes, the parser reads the tensor count and key-value count as 64-bit little-endian integers using read_u64_le_fp (lines 1681-1682). If either the magic or version check fails, the loader aborts before attempting to parse any tensor metadata.

Tensor Alignment and Offset Rules

The ds4 loader requires every tensor offset to be a multiple of the GGUF alignment value, which defaults to 32 bytes. This check appears in ds4.c lines 1683-1689:

g.alignment = DS4_GGUF_DEFAULT_ALIGNMENT;
// ... during tensor loading ...
if (offset % g.alignment) die("mis-aligned tensor");

Additionally, the declared data_offsets (start and end positions) must not exceed the actual file size. The function read_checked_len_fp in deepseek4-quantize.c (lines 71-82) enforces this boundary check, preventing out-of-bounds reads that could corrupt memory or crash the process.

Rank, Dimensions, and Ordering Constraints

Each tensor must declare a rank (n_dims) between 1 and 8 inclusive, validated at line 1748-1749 of ds4.c:

if (t->n_dims < 1 || t->n_dims > DS4Q_MAX_DIMS)
    die("bad GGUF tensor rank");

Dimension ordering follows a specific reversed convention: while stored in little-endian format, the dimensions are sequenced in reverse relative to the logical row-major order used by the runtime. The loader validates this reversal when mapping GGUF tensors to internal templates via check_reversed_shape in deepseek4-quantize.c (lines 1020-1030):

for (int i=0; i<nd; i++)
    if (tmpl->ne[i] != info->shape[nd-1-i])
        // shape mismatch error

This means a logical shape [1024, 512] appears on disk as [512, 1024] in the GGUF metadata.

Quantization Type and Block Size Validation

The loader maps GGUF type IDs to internal ds4q_type enumerations according to the table defined in deepseek4-quantize.c lines 1559-1567. Types that cannot be quantized—F32, F16, and BF16—pass through with fixed byte widths (4, 2, and 2 respectively), while quantized types must satisfy strict block-size constraints.

MXFP4 Block Requirements

For MXFP4 tensors, each block must be exactly 17 bytes. The validation occurs in deepseek4-quantize.c lines 846-847:

size_t block_bytes = ds4q_row_size(DS4Q_TYPE_MXFP4, 32);
if (block_bytes != 17) die("unexpected GGUF MXFP4 block size");

When repacking expert weights, the loader expects the packed size to match calculated expectations:

byte_buf packed = repack_fp4_weight_mxfp4(&weight, &scale);
if (packed.size != expected_per_expert) die("MXFP4 expert packed size mismatch");

FP4 and Q4_K Constraints

FP4 tensors require input dimensions divisible by 32, verified in dequant_fp4_weight (lines 86-95):

if (in_dim % 32) die("FP4 in_dim is not divisible by 32");

For Q4_K, the column count (ncols) must be a multiple of the block size returned by ds4q_block_size. The number of blocks must also match the corresponding scale tensor shape to ensure proper dequantization.

Data Payload Integrity Checks

Before reading any tensor data, the loader verifies that the on-disk payload size matches the shape-driven expectations. In deepseek4-quantize.c lines 53-56, the code compares the declared nbytes against the element count:

if (w->nbytes < weight_elems)
    die("FP8 tensor data smaller than its declared shape");

This checked_shape_product validation ensures that truncated or corrupted files are detected immediately rather than causing segmentation faults during inference.

Troubleshooting Common Layout Errors

When model loading fails, the error message indicates exactly which layout rule was violated. Use this checklist to diagnose issues, referencing the unit tests in tests/test_mxfp4_dot.c and tests/test_q4k_dot.c for reproducible layout failure scenarios:

  1. "bad GGUF template": Verify the first four bytes are ASCII "GGUF" using hexdump -C model.gguf | head -1.
  2. "mis-aligned tensor": Ensure all tensor offsets are multiples of 32. Recompute offsets in your converter or add padding.
  3. "bad GGUF tensor rank": Confirm n_dims is between 1 and 8.
  4. Shape mismatches: Check that dimensions are stored in reverse order compared to your logical layout.
  5. Quantization errors: For MXFP4, validate 17-byte blocks; for FP4, verify dimensions divisible by 32.
  6. "exceeds bytes remaining": Regenerate the GGUF file to ensure data_offsets reflect actual file size.

Summary

  • Header validation: Files must start with "GGUF" magic and valid version (ds4.c L1677).
  • Alignment: All tensor offsets must be 32-byte aligned (ds4.c L1683-1689).
  • Rank limits: Tensor dimensions must be 1-8, stored in reversed order (deepseek4-quantize.c L1020-1030).
  • Quantization rules: MXFP4 requires 17-byte blocks; FP4 requires dimensions divisible by 32 (deepseek4-quantize.c L846-847, L86-95).
  • Bounds checking: Data offsets must not exceed file size, and payload must match shape expectations (deepseek4-quantize.c L71-82).

Frequently Asked Questions

What causes "mis-aligned tensor" errors in ds4?

This error occurs when a tensor's data offset in the GGUF file is not a multiple of 32 bytes. The loader enforces strict alignment requirements at line 1683 of ds4.c to ensure SIMD-friendly memory access. To fix this, regenerate the GGUF file with proper padding or check your conversion script's offset calculations.

Why does ds4 require dimensions in reverse order?

The GGUF specification stores dimensions in little-endian format but reversed relative to the logical row-major order used by the runtime. The check_reversed_shape function in deepseek4-quantize.c (lines 1020-1030) validates this mapping. If your model expects shape [1024, 512], the GGUF metadata must list [512, 1024].

How do I validate MXFP4 block size compliance?

MXFP4 tensors must use exactly 17-byte blocks. Use the ds4q_row_size function to calculate expected sizes, as shown in deepseek4-quantize.c lines 846-847. If block_bytes != 17, your quantization parameters or converter implementation are incompatible with ds4's requirements.

What triggers the "bad GGUF template" error?

This indicates the file does not begin with the ASCII bytes "GGUF" or uses an unsupported version. Check the first four bytes of your file with a hex editor. The validation occurs immediately upon file open in ds4.c at line 1677, before any tensor metadata is read.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →