megadlms

GPU-optimized framework for training diffusion language models at any scale. The backend of Quokka, Super Data Learners, and OpenMoE 2 training.

22 articles 327 View on GitHub ↗
22 articles
MegaDLMs Framework Architecture: Core Components Explained

Discover the core components of the MegaDLMs framework architecture. Learn how MegaDLMs enhances Megatron-LM for scalable diffusion language model training with specialized layers and utilities.

architecture
Mar 4, 2026
How to Convert HuggingFace Checkpoints to Megatron Format Using MegaDLMs

Convert HuggingFace checkpoints to Megatron format with MegaDLMs. Learn how to reshape tensors and permute QKV for seamless integration using our dedicated conversion utility.

how-to-guide
Mar 4, 2026
Supported Tokenizer Types for MegaDLMs Data Preprocessing: Complete Guide

Explore ten supported tokenizer types for MegaDLMs data preprocessing, including BERT, GPT-2, SentencePiece, Llama2 & multimodal options. Select your ideal tokenizer type with --tokenizer-type.

tutorial
Mar 4, 2026
How MegaDLMs Process and Tokenize Input Data for Training: A Complete Technical Guide

Discover how MegaDLMs processes and tokenizes input data with a three-stage pipeline. Learn about padded tokenizers, multiprocessed JSON to token IDs, and indexed binary files for efficient GPU training.

deep-dive
Mar 4, 2026
How Expert Parallelism in MegaDLMs Optimizes MoE Model Training

Discover how Expert Parallelism in MegaDLMs optimizes MoE model training. Reduce GPU memory and speed up training with efficient token routing across ranks. Learn more now.

deep-dive
Mar 4, 2026
Context Parallelism in MegaDLMs: How It Handles Long Sequences

Explore Context Parallelism in MegaDLMs. Discover how it processes extra long sequences by partitioning data across GPUs, enabling full attention computation through efficient key-value communication.

deep-dive
Mar 4, 2026
Data Parallelism in MegaDLMs: DDP vs FSDP2 Implementation Guide

Explore Data Parallelism in MegaDLMs with our DDP vs FSDP2 implementation guide. Learn to optimize distributed training for standard or memory-constrained extreme-scale models.

how-to-guide
Mar 4, 2026
How Activation Checkpointing in MegaDLMs Reduces Memory Usage

Discover how activation checkpointing in MegaDLMs slashes GPU memory usage by recomputing activations during backward passes, enabling larger models and faster training.

deep-dive
Mar 4, 2026
Benefits of Using FP8 Precision with MegaDLMs and Required Hardware

Unlock up to 50% faster MegaDLM training with FP8 precision. Reduce memory usage and accelerate inference on NVIDIA Hopper Ada or Blackwell GPUs. Discover hardware requirements and benefits.

deep-dive
Mar 4, 2026
How MegaDLMs Implements Mixture of Experts (MoE) Architectures

MegaDLMs implements Mixture of Experts (MoE) with expert parallelism, configurable routing, and optimized token dispatchers. Learn how MegaDLMs supports MoE architectures.

deep-dive
Mar 4, 2026
The Three Stages of the DLM Training Pipeline in MegaDLMs: A Complete Guide

Explore the three stages of the DLM training pipeline in MegaDLMs: pre-training, SFT, and RL. This guide details each phase for effective Diffusion Language Model development.

how-to-guide
Mar 4, 2026
DLM vs Autoregressive LM Training in MegaDLMs: Key Architectural Differences

Explore DLM vs Autoregressive LM training within MegaDLMs. Understand the key architectural differences in bidirectional masking versus causal conditioning and how MegaDLMs switch modes for optimal performance.

deep-dive
Mar 4, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →