# Vertex Quantization Methods in meshoptimizer: Half-Precision, Snorm, and Unorm

> Explore vertex quantization methods in meshoptimizer: half-precision, snorm, and unorm. Optimize your 3D assets with these powerful techniques for smaller file sizes and faster loading.

- Repository: [Arseny Kapoulkine/meshoptimizer](https://github.com/zeux/meshoptimizer)
- Tags: deep-dive
- Published: 2026-07-12

---

**The `zeux/meshoptimizer` library provides four built-in vertex quantization methods: half-precision (FP16) conversion via `meshopt_quantizeHalf`, custom mantissa-bit reduction via `meshopt_quantizeFloat`, unsigned normalized (unorm) storage via `meshopt_quantizeUnorm`, and signed normalized (snorm) storage via `meshopt_quantizeSnorm`.**

Vertex data compression is critical for optimizing memory bandwidth and storage in real-time graphics applications. The `meshoptimizer` repository offers a specialized suite of quantization utilities that convert 32-bit floating-point vertex attributes into compact integer formats without requiring external dependencies.

## Half-Precision Floating-Point (FP16) Quantization

The **`meshopt_quantizeHalf`** function converts standard 32-bit IEEE-754 floats to 16-bit half-precision format. This method is implemented in [`src/quantization.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/quantization.cpp) at line 14 and handles edge cases including overflow saturation, NaN preservation, and denormal flushing to zero.

```c
unsigned short meshopt_quantizeHalf(float v);

```

When you need to preserve a wide dynamic range for position data or texture coordinates while cutting storage in half, FP16 quantization provides an efficient hardware-compatible format that most modern GPUs can decode natively.

## Custom Mantissa Quantization

For scenarios requiring bit-level precision control, **`meshopt_quantizeFloat`** rounds a 32-bit float to a configurable number of mantissa bits while preserving the sign bit and handling infinity/NaN values. This implementation resides in [`src/quantization.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/quantization.cpp) at line 37.

```c
float meshopt_quantizeFloat(float v, int N);

```

The parameter `N` specifies the number of mantissa bits to retain (valid range: 1–23). This method is ideal when you need to strip unnecessary precision from vertex weights or morph targets without converting to integer formats.

## Normalized Integer Formats (Unorm and Snorm)

For attributes that naturally exist within normalized ranges, `meshoptimizer` provides inline quantization functions in [`src/meshoptimizer.h`](https://github.com/zeux/meshoptimizer/blob/main/src/meshoptimizer.h). These map floating-point values to fixed-point integers that reconstruction shaders can decode using simple division.

### Unsigned Normalized (Unorm)

**`meshopt_quantizeUnorm`** maps values in the range `[0..1]` to an N-bit unsigned integer. The inline definition appears at line 1123 in [`src/meshoptimizer.h`](https://github.com/zeux/meshoptimizer/blob/main/src/meshoptimizer.h), with reconstruction computed as `q/(2^N-1)`.

```c
int meshopt_quantizeUnorm(float v, int N);

```

This method is optimal for texture coordinates, color channels, and ambient occlusion values that naturally clamp between zero and one.

### Signed Normalized (Snorm)

**`meshopt_quantizeSnorm`** maps values in `[-1..1]` to an N-bit signed integer. Defined inline at line 1133 in [`src/meshoptimizer.h`](https://github.com/zeux/meshoptimizer/blob/main/src/meshoptimizer.h), reconstruction uses the formula `q/(2^(N-1)-1)`.

```c
int meshopt_quantizeSnorm(float v, int N);

```

Use this for normal vectors, tangents, and bitangents where directionality must be preserved in a compact integer representation.

## Implementation Details and Source Files

The quantization utilities span two primary files in the repository:

- **[`src/quantization.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/quantization.cpp)** – Contains the implementation of `meshopt_quantizeHalf` and `meshopt_quantizeFloat`, handling IEEE-754 bit manipulation and edge case detection.
- **[`src/meshoptimizer.h`](https://github.com/zeux/meshoptimizer/blob/main/src/meshoptimizer.h)** – Provides inline definitions for `meshopt_quantizeUnorm` and `meshopt_quantizeSnorm`, enabling compile-time optimization of normalization operations.

## Practical Code Examples

Here are concrete implementations demonstrating each quantization method:

```c
/* Quantize a vertex position to half-precision */
float pos = 1.2345f;
unsigned short pos_half = meshopt_quantizeHalf(pos);  // 16-bit fp16

/* Reduce a color channel to 8-bit unorm */
float red = 0.78f;  // value in [0,1]
int red_u8 = meshopt_quantizeUnorm(red, 8);  // range: 0 to 255

/* Quantize a normal vector component to 10-bit snorm */
float nx = -0.42f;  // value in [-1,1]
int nx_s10 = meshopt_quantizeSnorm(nx, 10);  // range: -511 to 511

/* Keep only 5 mantissa bits of a skinning weight */
float weight = 0.123456f;
float weight_q = meshopt_quantizeFloat(weight, 5);  // rounded to 5 mantissa bits

```

## Summary

- **`meshopt_quantizeHalf`** converts 32-bit floats to IEEE-754 FP16 in [`src/quantization.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/quantization.cpp), handling overflow and NaN cases.
- **`meshopt_quantizeFloat`** provides configurable mantissa-bit reduction (1–23 bits) for custom precision requirements.
- **`meshopt_quantizeUnorm`** maps `[0..1]` values to N-bit unsigned integers via inline definition in [`src/meshoptimizer.h`](https://github.com/zeux/meshoptimizer/blob/main/src/meshoptimizer.h).
- **`meshopt_quantizeSnorm`** maps `[-1..1]` values to N-bit signed integers for normal vector compression.
- All four methods enable hardware-compatible vertex compression without decoding overhead on modern GPUs.

## Frequently Asked Questions

### When should I use half-precision versus unorm/snorm quantization?

**Use half-precision (`meshopt_quantizeHalf`)** when your data requires a wide dynamic range but doesn't fit cleanly into a `[0,1]` or `[-1,1]` bound, such as world-space positions or non-normalized texture coordinates. **Use unorm or snorm** when your data is already normalized, as these provide uniform quantization steps across the entire range and decode faster on the GPU via simple division constants.

### How does `meshopt_quantizeFloat` differ from `meshopt_quantizeHalf`?

`meshopt_quantizeFloat` preserves the 32-bit float container but reduces precision by zeroing out lower mantissa bits, whereas `meshopt_quantizeHalf` actually converts the value to a 16-bit representation. The former produces larger data but maintains IEEE-754 compatibility, while the latter halves storage requirements but introduces hardware-specific decoding requirements.

### How does meshoptimizer handle quantization overflow and special values?

According to the source code in [`src/quantization.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/quantization.cpp), the half-precision converter explicitly detects overflow and saturates to infinity, preserves NaN bit patterns, and flushes denormalized numbers to zero. The custom mantissa quantizer preserves sign bits and infinity/NaN values regardless of the mantissa bit count specified.

### Can I use these quantization methods for attributes other than vertex positions?

Yes. While designed for vertex data, these functions work with any floating-point values. Developers commonly use `meshopt_quantizeUnorm` for color attributes, `meshopt_quantizeSnorm` for normal maps, and `meshopt_quantizeHalf` for high-precision texture coordinates or morph target deltas.