How to Convert HuggingFace FP8 Weights to DeepSeek-V3 Format Using convert.py
The convert.py script reshapes and shards FP8 SafeTensor checkpoints from HuggingFace into a model-parallel format compatible with the DeepSeek-V3 inference engine.
The convert.py utility in the deepseek-ai/DeepSeek-V3 repository bridges the gap between HuggingFace-released FP8 weights and the custom tensor layout required by the native inference stack. This tool performs name remapping, expert filtering for Mixture-of-Experts (MoE) layers, and model-parallel sharding to produce checkpoint files ready for distributed loading.
How convert.py Processes FP8 Checkpoints
The conversion pipeline implemented in inference/convert.py executes three distinct operations on the source weights.
Loading and Name Normalization
First, the script loads each SafeTensor file using safetensors.torch.safe_open【L50-L55】. It then rewrites tensor names to match DeepSeek-V3’s internal naming scheme:
- The leading
model.prefix is stripped【L56-L58】 - Legacy substrings are replaced (
self_attn→attn,mlp→ffn,weight_scale_inv→scale)【L58-L61】 - Layer-specific keys (e.g.,
q_proj) are mapped to new identifiers via themappingdictionary defined at the top of the file【L11-L30】
Model-Parallel Sharding Logic
After normalization, the script applies model-parallel (MP) sharding based on the --model-parallel argument:
- Expert filtering: When a tensor contains
"experts"in its name, the script retains only the expert indices that fall within the current MP rank’s slice【L68-L72】 - Tensor splitting: For tensors mapped with a split dimension (
dimis notNone), the code asserts divisibility by the MP factor, then extracts the local slice usingtensor.narrow【L73-L76】
Output Generation
Processed tensors are accumulated into per-rank dictionaries (state_dicts) and serialized as model{i}-mp{mp}.safetensors【L80-L82】. Finally, auxiliary token files (e.g., tokenizer.model, tokenizer_config.json) are copied unchanged to the destination folder【L83-L85】, yielding a complete inference-ready checkpoint.
Running the Conversion from the Command Line
Install the minimal dependencies and execute the script with your specific model configuration:
# Install dependencies
pip install tqdm safetensors torch
# Run the converter
python inference/convert.py \
--hf-ckpt-path path/to/hf_fp8_checkpoint/ \
--save-path path/to/deepseek_v3_format/ \
--n-experts 64 \ # Total experts in the MoE model
--model-parallel 8 # MP factor (must divide n-experts evenly)
Verify the output contains sharded weights and tokenizer files:
ls path/to/deepseek_v3_format/
# → model0-mp8.safetensors model1-mp8.safetensors … model7-mp8.safetensors
# tokenizer.model tokenizer_config.json …
Programmatic Usage
You can also embed the conversion logic directly in Python scripts:
from inference.convert import main as convert
hf_dir = "path/to/hf_fp8_checkpoint"
out_dir = "path/to/deepseek_v3_format"
n_exp = 64
mp = 8
convert(hf_dir, out_dir, n_exp, mp)
Summary
- Input format: HuggingFace FP8 SafeTensor checkpoints with standard naming conventions
- Name remapping: Strips
model.prefixes and translates layer names via a predefinedmappingdictionary inconvert.py - Sharding strategy: Filters MoE experts by rank and splits tensors using
tensor.narrowaccording to the model-parallel factor - Output format: Rank-specific
model{i}-mp{mp}.safetensorsfiles plus copied tokenizer assets - Integration: Output is consumed directly by
inference/model.pyandinference/generate.pyin the DeepSeek-V3 repository
Frequently Asked Questions
What dependencies are required to run convert.py?
You need torch, safetensors, and tqdm. These handle tensor I/O, SafeTensor serialization, and progress bars respectively. No GPU is required for the conversion itself.
Why does the conversion require a model-parallel (MP) factor?
The MP factor determines how many shards to create and how to split expert weights and attention tensors. It must evenly divide the total expert count (e.g., 64 experts with MP=8 creates 8 shards of 8 experts each). This matches the parallelism strategy used during inference.
What is the difference between convert.py and fp8_cast_bf16.py?
convert.py restructures and shards FP8 weights into the DeepSeek-V3 native format while keeping the FP8 quantization. fp8_cast_bf16.py (found in the same inference/ directory) is a separate utility that decompresses FP8 weights to BF16 precision for hardware that requires higher-precision inference.
How do I verify the conversion succeeded?
Check that the destination directory contains model{i}-mp{mp}.safetensors files numbering equal to your MP factor (e.g., 8 files for MP=8) and that each file loads without errors in inference/model.py. The conversion logs progress via tqdm, reporting any tensor shape mismatches or mapping failures immediately.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →