# Embedding Quantization Perplexity Comparison: A Complete Guide to Choosing BitNet Model Formats

> Compare embedding quantization perplexity with BitNet. Learn to choose optimal BitNet model formats by measuring perplexity and analyzing results for better size speed and quality.

- Repository: [Microsoft/BitNet](https://github.com/microsoft/BitNet)
- Tags: deep-dive
- Published: 2026-03-13

---

**Use BitNet's [`utils/quantize_embeddings.py`](https://github.com/microsoft/BitNet/blob/main/utils/quantize_embeddings.py) to generate quantized embedding formats and [`utils/test_perplexity.py`](https://github.com/microsoft/BitNet/blob/main/utils/test_perplexity.py) to measure perplexity across datasets, then compare the resulting CSV files to select the format that balances size, speed, and model quality.**

Microsoft's BitNet repository provides specialized utilities for optimizing large language models through embedding quantization. By performing an embedding quantization perplexity comparison, developers can objectively measure how different quantization formats affect model performance before deployment. This workflow quantizes the token-embedding matrix into formats like Q6_K or Q8_0 and evaluates their impact on perplexity, file size, and inference throughput.

## Prerequisites: Building the Required Binaries

Before running the comparison pipeline, you must compile the underlying C++ binaries that perform quantization and perplexity calculations. The scripts expect `llama-quantize`, `llama-bench`, and `llama-perplexity` to reside in `build/bin/`.

```bash
cd BitNet
mkdir -p build && cd build
cmake .. && make -j$(nproc)

```

These binaries interface with the low-level kernels defined in [`gpu/bitnet_kernels/setup.py`](https://github.com/microsoft/BitNet/blob/main/gpu/bitnet_kernels/setup.py), which handle the actual tensor operations during benchmarking.

## Step 1: Quantize the Embedding Matrix

The [`utils/quantize_embeddings.py`](https://github.com/microsoft/BitNet/blob/main/utils/quantize_embeddings.py) script automates the creation of quantized models for every supported embedding format. It generates files like `ggml-model-i2_s-embed-q6_k.gguf` and records throughput metrics using `llama-bench`.

```bash
python3 utils/quantize_embeddings.py \
  --input models/BitNet-b1.58-2B-4T/ggml-model-f32.gguf \
  --output-dir models/BitNet-b1.58-2B-4T \
  --stats-dir stats \
  --skip-existing

```

This script produces `stats/embedding_benchmark.csv`, which contains **tokens-per-second** measurements for 1, 2, 4, and 8 CPU threads, alongside quantization time and file size information.

## Step 2: Run Perplexity Tests Across All Formats

With the `--test-embeddings` flag, [`utils/test_perplexity.py`](https://github.com/microsoft/BitNet/blob/main/utils/test_perplexity.py) automatically quantizes the model to each embedding type and executes `llama-perplexity` on every dataset in your data directory.

```bash
python3 utils/test_perplexity.py \
  --model models/BitNet-b1.58-2B-4T/ggml-model-f32.gguf \
  --data-dir data \
  --threads 16 \
  --ctx-size 512 \
  --test-embeddings \
  --csv-output results/embedding_perplexity.csv

```

The resulting CSV contains perplexity scores, execution times, and status indicators for each combination of embedding format and test corpus. Lower perplexity indicates better language modeling capability, while higher values signal quality degradation from aggressive quantization.

## Step 3: Analyze the Embedding Quantization Perplexity Comparison

To make an informed decision, merge the two CSV files on the `embedding_type` column and evaluate the trade-offs between:

- **Model size**: Check the file system for the size of each `ggml-model-i2_s-embed-*.gguf` file or reference the output logs from the quantizer.
- **Inference speed**: Compare the `threads_X` columns in `embedding_benchmark.csv` to determine throughput at various concurrency levels.
- **Prediction quality**: Verify that perplexity in `embedding_perplexity.csv` remains within an acceptable delta (typically less than 5%) of the baseline F32 model.

Select the most aggressive quantization format that satisfies your latency budget while keeping perplexity close to the original float32 baseline.

## Accelerated Testing with Quick Mode

For rapid iteration during development, add the `--quick` flag to [`test_perplexity.py`](https://github.com/microsoft/BitNet/blob/main/test_perplexity.py). This mode limits datasets to the first 4096 characters and reduces the context window to 128 tokens, dramatically reducing execution time while still exposing significant perplexity regressions.

```bash
python3 utils/test_perplexity.py \
  --model models/BitNet-b1.58-2B-4T/ggml-model-f32.gguf \
  --data-dir data \
  --test-embeddings \
  --quick \
  --csv-output results/quick_perplexity.csv

```

## Key Implementation Files

Understanding the source architecture helps interpret the results:

- **[`utils/quantize_embeddings.py`](https://github.com/microsoft/BitNet/blob/main/utils/quantize_embeddings.py)**: Orchestrates the quantization pipeline and invokes `llama-bench` for throughput measurement.
- **[`utils/test_perplexity.py`](https://github.com/microsoft/BitNet/blob/main/utils/test_perplexity.py)**: Executes `llama-perplexity` and aggregates results into structured CSV reports.
- **[`gpu/convert_checkpoint.py`](https://github.com/microsoft/BitNet/blob/main/gpu/convert_checkpoint.py)**: Handles the initial conversion of model checkpoints to GGUF format, which serves as input for the quantization scripts.
- **[`gpu/bitnet_kernels/setup.py`](https://github.com/microsoft/BitNet/blob/main/gpu/bitnet_kernels/setup.py)**: Defines the optimized kernels that execute during benchmarking and perplexity evaluation.

## Summary

- **Embedding quantization perplexity comparison** requires running [`utils/quantize_embeddings.py`](https://github.com/microsoft/BitNet/blob/main/utils/quantize_embeddings.py) followed by [`utils/test_perplexity.py`](https://github.com/microsoft/BitNet/blob/main/utils/test_perplexity.py) with the `--test-embeddings` flag.
- The workflow generates two CSV files: `embedding_benchmark.csv` for speed metrics and `embedding_perplexity.csv` for quality metrics.
- Compare **file size**, **tokens-per-second** (throughput), and **perplexity** to identify the format offering the best compression-to-quality ratio.
- Use `--quick` mode for preliminary testing, then run full tests before final deployment.
- All scripts require compiled binaries from `build/bin/` (produced via `cmake` and `make`).

## Frequently Asked Questions

### What is embedding quantization perplexity?

Embedding quantization perplexity measures how well a language model predicts text after its token-embedding matrix has been compressed to a lower-precision format like Q6_K or Q8_0. Lower perplexity scores indicate the quantized model retains the original's predictive capabilities, while higher scores suggest significant information loss from aggressive compression.

### Which embedding format offers the best trade-off in BitNet?

The optimal format depends on your deployment constraints, but Q6_K often provides a practical balance. According to the benchmark output from [`utils/quantize_embeddings.py`](https://github.com/microsoft/BitNet/blob/main/utils/quantize_embeddings.py), Q6_K typically reduces file size substantially while maintaining perplexity within 1-2% of the F32 baseline, though you should verify this against your specific dataset using the comparison CSVs.

### How long does a full embedding quantization perplexity comparison take?

A complete comparison across all supported formats (F32, F16, Q8_0, Q6_K, etc.) and multiple datasets can take several hours on CPU-bound systems. Use the `--quick` flag in [`utils/test_perplexity.py`](https://github.com/microsoft/BitNet/blob/main/utils/test_perplexity.py) to reduce this to minutes by limiting evaluation to 4096 characters per dataset and a 128-token context window.

### Can I compare formats without building the C++ binaries?

No. The Python scripts in `utils/` are wrappers that shell out to `llama-quantize`, `llama-bench`, and `llama-perplexity`. You must compile these binaries using `cmake .. && make -j$(nproc)` in the `build/` directory before running any embedding quantization perplexity comparison workflow.