Weight Update Validation Options in Miles: `--check-weight-update-equal` and `--check-weight-update-selector` Explained
Miles provides five command-line flags for verifying model weight synchronization after training steps, with --check-weight-update-equal enabling strict equality checks and --check-weight-update-selector controlling which models are validated.
The Miles training framework includes built-in verification mechanisms to detect weight synchronization failures between the training backend and inference engine. These checks are essential when debugging distributed training setups, mixed-precision quantization pipelines, or Mixture-of-Experts (MoE) draft models. The flags are implemented in miles/utils/arguments.py and propagate through the launcher scripts to the inference controller.
Core Weight Update Validation Flags
--check-weight-update-equal: Enable Strict Equality Verification
This store_true flag activates a post-update comparison between the target model's canonical checkpoint weights and the weights actually used by the inference engine.
When enabled, Miles aborts training immediately if any weight tensor mismatch is detected. This catches subtle synchronization bugs that might otherwise manifest as silent correctness errors.
python -m miles.main.train \
--model qwen3-5-35B-A3B \
--checkpoint /path/to/checkpoint \
--check-weight-update-equal
--check-weight-update-selector: Choose Which Models to Validate
This string selector (default: "all", choices: ["all", "target", "draft"]) fine-tunes the scope of the equality check:
- all — Validates both the target model and draft/MTP worker weights
- target — Validates only the target model (useful when MTP training is disabled)
- draft — Validates only the draft/MTP worker weights
# Validate only the target model, skip draft verification
python -m miles.main.train \
--model qwen3-5-35B-A3B \
--checkpoint /path/to/checkpoint \
--check-weight-update-equal \
--check-weight-update-selector target
These arguments are defined in miles/utils/arguments.py at lines 2217-2241 according to the source analysis.
Advanced Tuning Options
--check-weight-update-skip-list: Exclude Specific Layers
Supply a list of substrings to exclude matching weight names from validation. Mismatches for these layers are downgraded to informational messages rather than fatal errors.
# Exclude visual and audio layers (common for MTP setups)
python -m miles.main.train \
--model qwen3-5-35B-A3B \
--check-weight-update-equal \
--check-weight-update-skip-list visual audio
This is particularly useful when the draft model contains auxiliary heads (vision, audio) not present in the target checkpoint.
--check-weight-update-allow-quant-error: Tolerate Quantization Rounding
This store_true flag permits quantized tensors to differ by up to one unit in the last place (1 ULP) when comparison occurs in de-quantized space.
# Allow 1-ULP quantization difference
python -m miles.main.train \
--model qwen3-5-35B-A3B \
--check-weight-update-equal \
--check-weight-update-allow-quant-error
This option is required when running weight-update validation on models employing quantization, as exact bit-wise equality is impossible across quantization boundaries.
--check-lora-weight-equal: LoRA-Specific Validation
A LoRA-specific analogue that verifies adapter weights are correctly transferred from Megatron to SGLang. Enable alongside standard weight checks when training with Low-Rank Adaptation.
Integration in Training Pipelines
The validation flags are typically added to launcher scripts such as scripts/run_nemotron_3_ultra_550b_a55b.py. The arguments propagate from the command-line parser through to the inference controller, where the actual weight-sync logic executes the checks.
Test coverage exists in:
tests/fast/test_train.py— Fast CI validation of flag propagationtests/e2e/megatron/— End-to-end integration tests
Summary
--check-weight-update-equal— Master switch for post-update weight equality verification--check-weight-update-selector— Scope control:all,target, ordraftmodels--check-weight-update-skip-list— Exclude specific layer name patterns from checks--check-weight-update-allow-quant-error— Tolerate 1-ULP quantization rounding errors--check-lora-weight-equal— Dedicated validation for LoRA adapter synchronization
Frequently Asked Questions
What triggers a validation failure with --check-weight-update-equal?
Any bit-wise mismatch between the canonical checkpoint weights and inference engine weights causes immediate training abort, unless the mismatching parameter name matches a substring in --check-weight-update-skip-list or quantization tolerance is enabled via --check-weight-update-allow-quant-error.
When should I use --check-weight-update-selector target instead of all?
Use target when training without an MTP draft model, or when you want to isolate whether synchronization failures originate in the main model or auxiliary workers. This reduces validation overhead and narrows debugging scope.
Does --check-weight-update-allow-quant-error make validation less reliable?
The 1-ULP tolerance is mathematically sound for IEEE-754 floating-point comparisons and represents the minimum possible rounding error from quantization/de-quantization. It catches real synchronization bugs while avoiding false positives from benign rounding differences.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →