# Understanding the Post-Training Pipeline for LLMs: From SFT to GRPO

> Explore the LLM post-training pipeline from SFT to GRPO. Learn about reward modeling, preference alignment, and reinforcement learning in the FareedKhan-dev/train-llm-from-scratch repository.

- Repository: [Fareed Khan/train-llm-from-scratch](https://github.com/FareedKhan-dev/train-llm-from-scratch)
- Tags: deep-dive
- Published: 2026-06-11

---

**TLDR:** The post-training pipeline for LLMs in the FareedKhan-dev/train-llm-from-scratch repository implements a complete workflow from supervised fine-tuning through reward modeling and preference alignment (DPO/ORPO/KTO) to reinforcement learning (PPO/GRPO), with modular configurations and distributed training support.

The FareedKhan-dev/train-llm-from-scratch repository provides a comprehensive, production-ready implementation of the post-training pipeline for LLMs. This pipeline transforms a pretrained language model into an instruction-following, reward-aligned system through a series of modular stages. Each stage is backed by dedicated source files that handle everything from configuration management to distributed training.

## Configuration and CLI Infrastructure

The foundation of the pipeline rests on centralized configuration management and a generic command-line interface.

In [`config/post_training_config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/post_training_config.py), the repository defines dataclass configurations for every stage: `SFTConfig`, `RewardConfig`, `DPOConfig`, `PPOConfig`, and `GRPOConfig`. These dataclasses centralize hyperparameters such as learning rates, batch sizes, and KL coefficients.

The [`src/post_training/cli.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/cli.py) module provides `parse_config` and `parse_config_with_json` functions. These utilities automatically convert dataclass fields into `--field value` command-line arguments, supporting JSON overrides and ensuring reproducible experiments.

## Data Loading and Preparation

The data layer handles tokenized datasets for different training objectives.

- [`data_loader/sft_dataset.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/sft_dataset.py) loads instruction-following data for supervised fine-tuning.
- [`data_loader/preference_dataset.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/preference_dataset.py) manages preference pairs (chosen vs. rejected) for reward modeling and DPO.
- [`data_loader/prompt_dataset.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/prompt_dataset.py) provides prompts for reinforcement learning stages.

These loaders provide batch iterators that are DDP-aware, ensuring proper data distribution across multiple GPUs.

## Stage 1: Supervised Fine-Tuning (SFT)

The SFT stage trains the model on instruction data while masking out prompt tokens.

The training script [`scripts/train_sft.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_sft.py) orchestrates the loop, while [`src/post_training/sft.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/sft.py) implements the **prompt-masked cross-entropy loss**. This loss function uses a boolean `loss_mask` to silence gradients on user prompts, forcing the model to learn only the assistant response generation.

```bash
PYTHONPATH=. python scripts/train_sft.py \
    --config configs/sft.json \
    --lr 2e-5 \
    --batch_size 16 \
    --epochs 2 \
    --device cuda

```

This stage consumes a `pretrained_ckpt` and produces `sft.pt`, which serves as the base for subsequent alignment stages.

## Stage 2: Reward Model Training

Before reinforcement learning, the pipeline trains a scalar reward model on preference pairs.

The [`scripts/train_reward.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_reward.py) script loads the SFT checkpoint (`sft_ckpt`) and preference data (`preferences.jsonl`). In [`src/post_training/reward_train.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/reward_train.py), the `bradley_terry_loss` computes `-log(sigmoid(r_chosen - r_rejected))`, providing a principled pairwise preference objective.

```bash
PYTHONPATH=. python scripts/train_reward.py \
    --config configs/reward.json \
    --lr 1e-5 \
    --batch_size 8 \
    --epochs 1

```

The output `reward.pt` provides the reward signal for PPO training.

## Stage 3: Preference-Based Alignment (DPO, ORPO, KTO)

This stage aligns the policy directly to preferences without requiring a learned value function.

The [`scripts/train_dpo.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_dpo.py) script supports three algorithms:
- **Direct Preference Optimization (DPO)**
- **Odds Ratio Preference Optimization (ORPO)**
- **Kahneman-Tversky Optimization (KTO)**

These objectives operate on summed response log-probs computed by `sequence_logprobs` in [`src/post_training/rollout.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/rollout.py). The core losses—`dpo_loss`, `orpo_loss`, and `kto_loss`—reside in [`src/post_training/dpo.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/dpo.py), using the SFT checkpoint as both the policy and frozen reference model (`ref_*_logps`).

```bash
PYTHONPATH=. python scripts/train_dpo.py \
    --config configs/dpo.json \
    --beta 0.1 \
    --loss_type dpo

```

## Stage 4: Reinforcement Learning with PPO

Classic PPO implements RLHF with a learned value head and Generalized Advantage Estimation (GAE).

The [`scripts/train_ppo.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_ppo.py) script generates rollouts via [`src/post_training/rollout.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/rollout.py), then computes advantages using `compute_gae` in [`src/post_training/ppo.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/ppo.py). This combines rewards, bootstrapped values, and a λ-trace to reduce variance. The update applies a clipped surrogate loss (`ppo_policy_loss`) and a clipped value loss (`ppo_value_loss`), with an optional KL penalty against the reference policy.

```bash
PYTHONPATH=. python scripts/train_ppo.py \
    --config configs/ppo.json \
    --iterations 1000 \
    --ppo_epochs 4 \
    --clip 0.2 \
    --kl_coef 0.05

```

Key utilities like `masked_mean` in [`src/post_training/utils.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/utils.py) ensure proper normalization across variable-length sequences.

## Stage 5: Group Relative Policy Optimization (GRPO)

GRPO removes the need for a learned value network by using group-relative baselines.

The [`scripts/train_grpo.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_grpo.py) script samples a *group* of completions per prompt (controlled by `--group_size`). Instead of a critic, the `group_advantages` function in [`src/post_training/grpo.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/grpo.py) normalizes each sample's reward by its group mean and standard deviation. The loss applies a per-token clipped surrogate plus a KL penalty using the `k3_kl` estimator, which provides an unbiased, non-negative estimate of KL(policy‖ref).

```bash
PYTHONPATH=. python scripts/train_grpo.py \
    --config configs/grpo.json \
    --iterations 1000 \
    --group_size 8 \
    --clip 0.2 \
    --kl_coef 0.04

```

## Shared Utilities: Rollout, Distributed Training, and Evaluation

All stages share common utilities to ensure consistency.

**Generation and Log-Prob Utilities:** The [`src/post_training/rollout.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/rollout.py) module provides `generate_with_logprobs` and `compute_logprobs` for generation with per-token log-probability recording. This ensures that log-prob computations during training match the sampling distribution, which is essential for unbiased RL updates.

**Distributed Training:** The [`src/post_training/distributed.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/distributed.py) module contains `ddp_setup`, `ddp_wrap`, and `reduce_scalar` to handle `torch.distributed` boilerplate, while [`src/post_training/logging_utils.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/logging_utils.py) manages metrics tracking with optional W&B integration. These enable seamless scaling from single GPU to multi-node clusters.

**Evaluation:** The [`scripts/eval_post_training.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/eval_post_training.py) script reloads final checkpoints and runs inference on held-out prompts, computing metrics such as preference accuracy (`implicit_accuracy` in [`dpo.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/dpo.py)) or task-specific scores like GSM-8K.

## Summary

The post-training pipeline for LLMs in this repository provides a complete, modular pathway from base model to aligned assistant:

- **Configuration management** via dataclasses in [`config/post_training_config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/post_training_config.py) ensures reproducibility across all stages.
- **Supervised Fine-Tuning** uses prompt-masked loss to focus learning on assistant responses.
- **Reward modeling** applies the Bradley-Terry loss to learn from human preferences.
- **Preference alignment** supports DPO, ORPO, and KTO objectives without requiring a separate value model.
- **Reinforcement learning** implements both value-based PPO with GAE and value-free GRPO with group-relative advantages.
- **Shared utilities** for rollouts, distributed training, and evaluation ensure consistency and scalability.

## Frequently Asked Questions

### What is the difference between PPO and GRPO in the post-training pipeline?

**PPO (Proximal Policy Optimization)** relies on a learned value network (critic) to estimate advantages using Generalized Advantage Estimation (GAE), as implemented in [`src/post_training/ppo.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/ppo.py). **GRPO (Group Relative Policy Optimization)**, found in [`src/post_training/grpo.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/grpo.py), eliminates the value network entirely by sampling a group of completions per prompt and computing advantages relative to the group mean. This reduces memory overhead and removes the need to train a separate critic, making GRPO more efficient for certain applications.

### How does the prompt-masked loss in SFT prevent the model from learning to copy prompts?

The `sft_loss` function in [`src/post_training/sft.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/sft.py) accepts a boolean `loss_mask` that identifies which tokens belong to the user prompt versus the assistant response. During the forward pass, gradients are only computed for tokens where `loss_mask` is True (the assistant portion). This prevents the model from wasting capacity learning to copy input prompts and focuses training entirely on generating appropriate responses.

### Why is the k3 KL estimator used in GRPO instead of standard KL divergence?

The `k3_kl` function in [`src/post_training/grpo.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/grpo.py) provides an unbiased, non-negative estimator of the KL divergence between the policy and reference model. Unlike standard estimators that can be negative or biased, k3 ensures stable KL penalty calculations during training. This stability is crucial when applying the KL coefficient to prevent the policy from diverging too far from the reference while maintaining learning signal.

### Can I run the post-training pipeline on a single GPU or does it require multi-GPU setup?

The pipeline supports both single-GPU and multi-GPU training. The [`src/post_training/distributed.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/distributed.py) utilities automatically handle DDP initialization via `ddp_setup` and `ddp_wrap`, but the scripts check for distributed availability before wrapping models. You can run any stage on a single GPU by omitting the `torchrun` wrapper and using `--device cuda`, as shown in the SFT example above.