megadlms
GPU-optimized framework for training diffusion language models at any scale. The backend of Quokka, Super Data Learners, and OpenMoE 2 training.
Discover the core components of the MegaDLMs framework architecture. Learn how MegaDLMs enhances Megatron-LM for scalable diffusion language model training with specialized layers and utilities.
How to Convert HuggingFace Checkpoints to Megatron Format Using MegaDLMsConvert HuggingFace checkpoints to Megatron format with MegaDLMs. Learn how to reshape tensors and permute QKV for seamless integration using our dedicated conversion utility.
Supported Tokenizer Types for MegaDLMs Data Preprocessing: Complete GuideExplore ten supported tokenizer types for MegaDLMs data preprocessing, including BERT, GPT-2, SentencePiece, Llama2 & multimodal options. Select your ideal tokenizer type with --tokenizer-type.
How MegaDLMs Process and Tokenize Input Data for Training: A Complete Technical GuideDiscover how MegaDLMs processes and tokenizes input data with a three-stage pipeline. Learn about padded tokenizers, multiprocessed JSON to token IDs, and indexed binary files for efficient GPU training.
How Expert Parallelism in MegaDLMs Optimizes MoE Model TrainingDiscover how Expert Parallelism in MegaDLMs optimizes MoE model training. Reduce GPU memory and speed up training with efficient token routing across ranks. Learn more now.
Context Parallelism in MegaDLMs: How It Handles Long SequencesExplore Context Parallelism in MegaDLMs. Discover how it processes extra long sequences by partitioning data across GPUs, enabling full attention computation through efficient key-value communication.
Data Parallelism in MegaDLMs: DDP vs FSDP2 Implementation GuideExplore Data Parallelism in MegaDLMs with our DDP vs FSDP2 implementation guide. Learn to optimize distributed training for standard or memory-constrained extreme-scale models.
How Activation Checkpointing in MegaDLMs Reduces Memory UsageDiscover how activation checkpointing in MegaDLMs slashes GPU memory usage by recomputing activations during backward passes, enabling larger models and faster training.
Benefits of Using FP8 Precision with MegaDLMs and Required HardwareUnlock up to 50% faster MegaDLM training with FP8 precision. Reduce memory usage and accelerate inference on NVIDIA Hopper Ada or Blackwell GPUs. Discover hardware requirements and benefits.
How MegaDLMs Implements Mixture of Experts (MoE) ArchitecturesMegaDLMs implements Mixture of Experts (MoE) with expert parallelism, configurable routing, and optimized token dispatchers. Learn how MegaDLMs supports MoE architectures.
The Three Stages of the DLM Training Pipeline in MegaDLMs: A Complete GuideExplore the three stages of the DLM training pipeline in MegaDLMs: pre-training, SFT, and RL. This guide details each phase for effective Diffusion Language Model development.
DLM vs Autoregressive LM Training in MegaDLMs: Key Architectural DifferencesExplore DLM vs Autoregressive LM training within MegaDLMs. Understand the key architectural differences in bidirectional masking versus causal conditioning and how MegaDLMs switch modes for optimal performance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →