Understanding the Post-Training Pipeline for LLMs: From SFT to GRPO

TLDR: The post-training pipeline for LLMs in the FareedKhan-dev/train-llm-from-scratch repository implements a complete workflow from supervised fine-tuning through reward modeling and preference alignment (DPO/ORPO/KTO) to reinforcement learning (PPO/GRPO), with modular configurations and distributed training support.

The FareedKhan-dev/train-llm-from-scratch repository provides a comprehensive, production-ready implementation of the post-training pipeline for LLMs. This pipeline transforms a pretrained language model into an instruction-following, reward-aligned system through a series of modular stages. Each stage is backed by dedicated source files that handle everything from configuration management to distributed training.

Configuration and CLI Infrastructure

The foundation of the pipeline rests on centralized configuration management and a generic command-line interface.

In config/post_training_config.py, the repository defines dataclass configurations for every stage: SFTConfig, RewardConfig, DPOConfig, PPOConfig, and GRPOConfig. These dataclasses centralize hyperparameters such as learning rates, batch sizes, and KL coefficients.

The src/post_training/cli.py module provides parse_config and parse_config_with_json functions. These utilities automatically convert dataclass fields into --field value command-line arguments, supporting JSON overrides and ensuring reproducible experiments.

Data Loading and Preparation

The data layer handles tokenized datasets for different training objectives.

These loaders provide batch iterators that are DDP-aware, ensuring proper data distribution across multiple GPUs.

Stage 1: Supervised Fine-Tuning (SFT)

The SFT stage trains the model on instruction data while masking out prompt tokens.

The training script scripts/train_sft.py orchestrates the loop, while src/post_training/sft.py implements the prompt-masked cross-entropy loss. This loss function uses a boolean loss_mask to silence gradients on user prompts, forcing the model to learn only the assistant response generation.

PYTHONPATH=. python scripts/train_sft.py \
    --config configs/sft.json \
    --lr 2e-5 \
    --batch_size 16 \
    --epochs 2 \
    --device cuda

This stage consumes a pretrained_ckpt and produces sft.pt, which serves as the base for subsequent alignment stages.

Stage 2: Reward Model Training

Before reinforcement learning, the pipeline trains a scalar reward model on preference pairs.

The scripts/train_reward.py script loads the SFT checkpoint (sft_ckpt) and preference data (preferences.jsonl). In src/post_training/reward_train.py, the bradley_terry_loss computes -log(sigmoid(r_chosen - r_rejected)), providing a principled pairwise preference objective.

PYTHONPATH=. python scripts/train_reward.py \
    --config configs/reward.json \
    --lr 1e-5 \
    --batch_size 8 \
    --epochs 1

The output reward.pt provides the reward signal for PPO training.

Stage 3: Preference-Based Alignment (DPO, ORPO, KTO)

This stage aligns the policy directly to preferences without requiring a learned value function.

The scripts/train_dpo.py script supports three algorithms:

  • Direct Preference Optimization (DPO)
  • Odds Ratio Preference Optimization (ORPO)
  • Kahneman-Tversky Optimization (KTO)

These objectives operate on summed response log-probs computed by sequence_logprobs in src/post_training/rollout.py. The core losses—dpo_loss, orpo_loss, and kto_loss—reside in src/post_training/dpo.py, using the SFT checkpoint as both the policy and frozen reference model (ref_*_logps).

PYTHONPATH=. python scripts/train_dpo.py \
    --config configs/dpo.json \
    --beta 0.1 \
    --loss_type dpo

Stage 4: Reinforcement Learning with PPO

Classic PPO implements RLHF with a learned value head and Generalized Advantage Estimation (GAE).

The scripts/train_ppo.py script generates rollouts via src/post_training/rollout.py, then computes advantages using compute_gae in src/post_training/ppo.py. This combines rewards, bootstrapped values, and a λ-trace to reduce variance. The update applies a clipped surrogate loss (ppo_policy_loss) and a clipped value loss (ppo_value_loss), with an optional KL penalty against the reference policy.

PYTHONPATH=. python scripts/train_ppo.py \
    --config configs/ppo.json \
    --iterations 1000 \
    --ppo_epochs 4 \
    --clip 0.2 \
    --kl_coef 0.05

Key utilities like masked_mean in src/post_training/utils.py ensure proper normalization across variable-length sequences.

Stage 5: Group Relative Policy Optimization (GRPO)

GRPO removes the need for a learned value network by using group-relative baselines.

The scripts/train_grpo.py script samples a group of completions per prompt (controlled by --group_size). Instead of a critic, the group_advantages function in src/post_training/grpo.py normalizes each sample's reward by its group mean and standard deviation. The loss applies a per-token clipped surrogate plus a KL penalty using the k3_kl estimator, which provides an unbiased, non-negative estimate of KL(policy‖ref).

PYTHONPATH=. python scripts/train_grpo.py \
    --config configs/grpo.json \
    --iterations 1000 \
    --group_size 8 \
    --clip 0.2 \
    --kl_coef 0.04

Shared Utilities: Rollout, Distributed Training, and Evaluation

All stages share common utilities to ensure consistency.

Generation and Log-Prob Utilities: The src/post_training/rollout.py module provides generate_with_logprobs and compute_logprobs for generation with per-token log-probability recording. This ensures that log-prob computations during training match the sampling distribution, which is essential for unbiased RL updates.

Distributed Training: The src/post_training/distributed.py module contains ddp_setup, ddp_wrap, and reduce_scalar to handle torch.distributed boilerplate, while src/post_training/logging_utils.py manages metrics tracking with optional W&B integration. These enable seamless scaling from single GPU to multi-node clusters.

Evaluation: The scripts/eval_post_training.py script reloads final checkpoints and runs inference on held-out prompts, computing metrics such as preference accuracy (implicit_accuracy in dpo.py) or task-specific scores like GSM-8K.

Summary

The post-training pipeline for LLMs in this repository provides a complete, modular pathway from base model to aligned assistant:

  • Configuration management via dataclasses in config/post_training_config.py ensures reproducibility across all stages.
  • Supervised Fine-Tuning uses prompt-masked loss to focus learning on assistant responses.
  • Reward modeling applies the Bradley-Terry loss to learn from human preferences.
  • Preference alignment supports DPO, ORPO, and KTO objectives without requiring a separate value model.
  • Reinforcement learning implements both value-based PPO with GAE and value-free GRPO with group-relative advantages.
  • Shared utilities for rollouts, distributed training, and evaluation ensure consistency and scalability.

Frequently Asked Questions

What is the difference between PPO and GRPO in the post-training pipeline?

PPO (Proximal Policy Optimization) relies on a learned value network (critic) to estimate advantages using Generalized Advantage Estimation (GAE), as implemented in src/post_training/ppo.py. GRPO (Group Relative Policy Optimization), found in src/post_training/grpo.py, eliminates the value network entirely by sampling a group of completions per prompt and computing advantages relative to the group mean. This reduces memory overhead and removes the need to train a separate critic, making GRPO more efficient for certain applications.

How does the prompt-masked loss in SFT prevent the model from learning to copy prompts?

The sft_loss function in src/post_training/sft.py accepts a boolean loss_mask that identifies which tokens belong to the user prompt versus the assistant response. During the forward pass, gradients are only computed for tokens where loss_mask is True (the assistant portion). This prevents the model from wasting capacity learning to copy input prompts and focuses training entirely on generating appropriate responses.

Why is the k3 KL estimator used in GRPO instead of standard KL divergence?

The k3_kl function in src/post_training/grpo.py provides an unbiased, non-negative estimator of the KL divergence between the policy and reference model. Unlike standard estimators that can be negative or biased, k3 ensures stable KL penalty calculations during training. This stability is crucial when applying the KL coefficient to prevent the policy from diverging too far from the reference while maintaining learning signal.

Can I run the post-training pipeline on a single GPU or does it require multi-GPU setup?

The pipeline supports both single-GPU and multi-GPU training. The src/post_training/distributed.py utilities automatically handle DDP initialization via ddp_setup and ddp_wrap, but the scripts check for distributed availability before wrapping models. You can run any stage on a single GPU by omitting the torchrun wrapper and using --device cuda, as shown in the SFT example above.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →