Embedding Quantization Perplexity Comparison: A Complete Guide to Choosing BitNet Model Formats

Use BitNet's utils/quantize_embeddings.py to generate quantized embedding formats and utils/test_perplexity.py to measure perplexity across datasets, then compare the resulting CSV files to select the format that balances size, speed, and model quality.

Microsoft's BitNet repository provides specialized utilities for optimizing large language models through embedding quantization. By performing an embedding quantization perplexity comparison, developers can objectively measure how different quantization formats affect model performance before deployment. This workflow quantizes the token-embedding matrix into formats like Q6_K or Q8_0 and evaluates their impact on perplexity, file size, and inference throughput.

Prerequisites: Building the Required Binaries

Before running the comparison pipeline, you must compile the underlying C++ binaries that perform quantization and perplexity calculations. The scripts expect llama-quantize, llama-bench, and llama-perplexity to reside in build/bin/.

cd BitNet
mkdir -p build && cd build
cmake .. && make -j$(nproc)

These binaries interface with the low-level kernels defined in gpu/bitnet_kernels/setup.py, which handle the actual tensor operations during benchmarking.

Step 1: Quantize the Embedding Matrix

The utils/quantize_embeddings.py script automates the creation of quantized models for every supported embedding format. It generates files like ggml-model-i2_s-embed-q6_k.gguf and records throughput metrics using llama-bench.

python3 utils/quantize_embeddings.py \
  --input models/BitNet-b1.58-2B-4T/ggml-model-f32.gguf \
  --output-dir models/BitNet-b1.58-2B-4T \
  --stats-dir stats \
  --skip-existing

This script produces stats/embedding_benchmark.csv, which contains tokens-per-second measurements for 1, 2, 4, and 8 CPU threads, alongside quantization time and file size information.

Step 2: Run Perplexity Tests Across All Formats

With the --test-embeddings flag, utils/test_perplexity.py automatically quantizes the model to each embedding type and executes llama-perplexity on every dataset in your data directory.

python3 utils/test_perplexity.py \
  --model models/BitNet-b1.58-2B-4T/ggml-model-f32.gguf \
  --data-dir data \
  --threads 16 \
  --ctx-size 512 \
  --test-embeddings \
  --csv-output results/embedding_perplexity.csv

The resulting CSV contains perplexity scores, execution times, and status indicators for each combination of embedding format and test corpus. Lower perplexity indicates better language modeling capability, while higher values signal quality degradation from aggressive quantization.

Step 3: Analyze the Embedding Quantization Perplexity Comparison

To make an informed decision, merge the two CSV files on the embedding_type column and evaluate the trade-offs between:

  • Model size: Check the file system for the size of each ggml-model-i2_s-embed-*.gguf file or reference the output logs from the quantizer.
  • Inference speed: Compare the threads_X columns in embedding_benchmark.csv to determine throughput at various concurrency levels.
  • Prediction quality: Verify that perplexity in embedding_perplexity.csv remains within an acceptable delta (typically less than 5%) of the baseline F32 model.

Select the most aggressive quantization format that satisfies your latency budget while keeping perplexity close to the original float32 baseline.

Accelerated Testing with Quick Mode

For rapid iteration during development, add the --quick flag to test_perplexity.py. This mode limits datasets to the first 4096 characters and reduces the context window to 128 tokens, dramatically reducing execution time while still exposing significant perplexity regressions.

python3 utils/test_perplexity.py \
  --model models/BitNet-b1.58-2B-4T/ggml-model-f32.gguf \
  --data-dir data \
  --test-embeddings \
  --quick \
  --csv-output results/quick_perplexity.csv

Key Implementation Files

Understanding the source architecture helps interpret the results:

Summary

  • Embedding quantization perplexity comparison requires running utils/quantize_embeddings.py followed by utils/test_perplexity.py with the --test-embeddings flag.
  • The workflow generates two CSV files: embedding_benchmark.csv for speed metrics and embedding_perplexity.csv for quality metrics.
  • Compare file size, tokens-per-second (throughput), and perplexity to identify the format offering the best compression-to-quality ratio.
  • Use --quick mode for preliminary testing, then run full tests before final deployment.
  • All scripts require compiled binaries from build/bin/ (produced via cmake and make).

Frequently Asked Questions

What is embedding quantization perplexity?

Embedding quantization perplexity measures how well a language model predicts text after its token-embedding matrix has been compressed to a lower-precision format like Q6_K or Q8_0. Lower perplexity scores indicate the quantized model retains the original's predictive capabilities, while higher scores suggest significant information loss from aggressive compression.

Which embedding format offers the best trade-off in BitNet?

The optimal format depends on your deployment constraints, but Q6_K often provides a practical balance. According to the benchmark output from utils/quantize_embeddings.py, Q6_K typically reduces file size substantially while maintaining perplexity within 1-2% of the F32 baseline, though you should verify this against your specific dataset using the comparison CSVs.

How long does a full embedding quantization perplexity comparison take?

A complete comparison across all supported formats (F32, F16, Q8_0, Q6_K, etc.) and multiple datasets can take several hours on CPU-bound systems. Use the --quick flag in utils/test_perplexity.py to reduce this to minutes by limiting evaluation to 4096 characters per dataset and a 128-token context window.

Can I compare formats without building the C++ binaries?

No. The Python scripts in utils/ are wrappers that shell out to llama-quantize, llama-bench, and llama-perplexity. You must compile these binaries using cmake .. && make -j$(nproc) in the build/ directory before running any embedding quantization perplexity comparison workflow.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →