# How to Train a Reward Model for LLM Alignment: A Complete Implementation Guide

> Learn to train a reward model for LLM alignment. This guide details implementation using a scalar reward head and Bradley-Terry loss on human preference data.

- Repository: [Fareed Khan/train-llm-from-scratch](https://github.com/FareedKhan-dev/train-llm-from-scratch)
- Tags: how-to-guide
- Published: 2026-06-11

---

**Training a reward model for LLM alignment involves attaching a scalar reward head to a supervised fine-tuned (SFT) backbone and optimizing it with Bradley-Terry pairwise loss on human preference data.**

In the `FareedKhan-dev/train-llm-from-scratch` repository, reward model training represents the third stage of the RLHF pipeline. This component learns to assign scalar scores to generated responses based on human preferences, providing the critical signal needed for PPO or DPO fine-tuning. The implementation pairs a lightweight reward head with a pairwise ranking loss to efficiently capture human judgment patterns.

## Architecture

### The RewardModel Class

The `RewardModel` class in [`src/post_training/reward_model.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/reward_model.py) wraps your SFT transformer and projects hidden states into scalar reward values. Unlike the base model, it discards the language modeling head in favor of a single linear layer.

```python
class RewardModel(nn.Module):
    """Wrap a Transformer and add a scalar reward head (no lm_head used)."""
    def __init__(self, transformer: Transformer) -> None:
        super().__init__()
        self.transformer = transformer
        n_embed = transformer.lm_head.in_features          # ← hidden size

        self.reward_head = nn.Linear(n_embed, 1, bias=False)
        nn.init.zeros_(self.reward_head.weight)            # start near‑zero rewards

    def token_rewards(self, idx: torch.Tensor) -> torch.Tensor:
        # (B, T) → per‑token scalar reward

        hidden = self.transformer.forward_hidden(idx)
        return self.reward_head(hidden).squeeze(-1)

    def forward(self, idx: torch.Tensor,
                seq_lengths: torch.Tensor | None = None) -> torch.Tensor:
        # (B,) scalar reward per sequence

        rewards = self.token_rewards(idx)                # (B, T)

        if seq_lengths is None:
            return rewards[:, -1]                        # last token of full‑length seq

        return gather_last(rewards, seq_lengths)        # last *real* token

```

The **scalar head** is a `nn.Linear` layer mapping from the transformer's hidden dimension to a single value, initialized with zeros to start training from neutral rewards. The `forward` method extracts the hidden state of the **last real token** using `gather_last` (defined in [`src/post_training/utils.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/utils.py)), which indexes `rewards[i, seq_lengths[i]‑1]` to ignore padding tokens.

### Bradley-Terry Loss Function

The training objective uses the Bradley-Terry pairwise comparison loss implemented in [`src/post_training/reward_train.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/reward_train.py):

```python
def bradley_tyrrey_loss(chosen, rejected):
    # L = -log σ(chosen - rejected)

    return -F.logsigmoid(chosen - rejected).mean()

```

This loss pushes the reward of the **chosen** response above that of the **rejected** response by maximizing the log-probability that the preferred response scores higher. The repository monitors **preference accuracy** (`chosen > rejected`) as the primary training metric alongside the reward margin.

### Preference Data Flow

The [`data_loader/preference_dataset.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/preference_dataset.py) loader prepares tuples of prompt-response pairs with specific tensor formatting:

```python
batch = {
    "chosen_ids":   Tensor(B, L),   # token ids of prompt+chosen+EOT

    "chosen_len":   Tensor(B),      # true length of each chosen sequence

    "rejected_ids": Tensor(B, L),   # token ids of prompt+rejected+EOT

    "rejected_len": Tensor(B),      # true length of each rejected sequence

}

```

Sequences are **right-padded** to uniform length, which remains safe because the underlying Transformer uses causal attention masks. The last real token never attends to padding tokens that follow it, ensuring valid reward computation via `gather_last`.

## Training Procedure

### Running the Training Loop

The entry point [`scripts/train_reward.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_reward.py) orchestrates the training process. A typical execution looks like:

```bash
PYTHONPATH=. python scripts/train_reward.py \
    --cfg configs/post_training_config.yaml \
    --data_path data/preference/train.jsonl \
    --ckpt_dir ckpts/reward \
    --lr 1e-5 --max_len 768

```

Inside the training loop, the implementation concatenates chosen and rejected sequences for computational efficiency:

```python

# Build the SFT backbone from the config

backbone = build_model_from_config(cfg)
rm = RewardModel(backbone).to(device)

# Preference iterator yields collated tensors

iterator = get_preference_iterator(
    path=args.data_path,
    batch_size=args.batch_size,
    max_len=args.max_len,
    device=device,
)

optim = torch.optim.AdamW(rm.parameters(), lr=args.lr)

for batch in iterator:
    # Concatenate chosen + rejected so we get a single forward pass (2B rows)

    ids = torch.cat([batch["chosen_ids"], batch["rejected_ids"]], dim=0)
    lens = torch.cat([batch["chosen_len"], batch["rejected_len"]], dim=0)

    # Forward → (2B,) rewards

    rewards = rm(ids, seq_lengths=lens).float()
    chosen_r, rejected_r = rewards[:B], rewards[B:]

    loss = bradley_tyrrey_loss(chosen_r, rejected_r)
    loss.backward()
    optim.step()
    optim.zero_grad()

```

By processing both sides of the preference pair in a **single forward pass** (2B rows), the implementation maximizes GPU utilization. After the forward pass, the reward tensor splits back into `chosen_r` and `rejected_r` for loss computation. By default, the optimizer updates only the reward head parameters while the SFT backbone remains frozen.

### Saving and Loading Checkpoints

Save checkpoints using standard PyTorch serialization:

```python
torch.save({"model_state_dict": rm.state_dict()}, "ckpts/reward/reward.pt")

```

For inference, use the `load_reward_model` helper from [`src/post_training/reward_model.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/reward_model.py), which handles Distributed Data Parallel (DDP) prefix stripping (`module.`) and automatically sets evaluation mode:

```python
from src.post_training.reward_model import load_reward_model
from src.config.loader import load_config

cfg = load_config("configs/post_training_config.yaml")
rm = load_reward_model(cfg, "ckpts/reward/reward.pt", device="cpu")

```

## Practical Code Examples

### Instantiating from a Checkpoint

Load a trained reward model for evaluation or downstream RL training:

```python
from src.post_training.reward_model import load_reward_model
from src.config.loader import load_config

cfg = load_config("configs/post_training_config.yaml")
rm = load_reward_model(cfg, "ckpts/reward/reward.pt", device="cpu")

```

### Scoring Individual Responses

Generate scalar rewards for specific model outputs:

```python
prompt = "Explain the greenhouse effect in one sentence."
response = "Greenhouse gases trap heat, warming the planet."
ids, _ = encode_chat([{"role":"user","content":prompt},
                      {"role":"assistant","content":response}])
ids = torch.tensor([ids], dtype=torch.long)          # (1, T)

reward = rm(ids)                                     # → scalar tensor

print("Reward:", reward.item())

```

### Integrating with PPO

Use the trained reward model to provide feedback signals during PPO fine-tuning:

```python
from src.post_training.ppo import PPOTrainer   # simplified view

trainer = PPOTrainer(
    policy=model,                     # the SFT model

    reward_model=rm,                  # the trained RM

    cfg=ppo_cfg,
)
trainer.train()

```

The PPO trainer internally calls `rm(ids, seq_lengths)` for each rollout to compute advantage estimates and policy gradients.

## Summary

Training a reward model for LLM alignment in this repository follows a structured pipeline:

* **Architecture**: Attach a zero-initialized linear reward head to the SFT backbone, extracting rewards from the last real token using `gather_last`
* **Data**: Load preference pairs via `get_preference_iterator`, which returns right-padded tensors with explicit sequence lengths
* **Training**: Concatenate chosen and rejected batches for efficiency, compute Bradley-Terry loss, and optimize with AdamW while monitoring preference accuracy
* **Inference**: Load trained models using `load_reward_model` to remove DDP prefixes and enable evaluation mode
* **Integration**: Feed the reward model into PPO or DPO trainers to provide human preference signals for policy optimization

## Frequently Asked Questions

### What is the difference between a reward model and the base LLM?

The base LLM predicts next-token probabilities using a language modeling head, while the reward model replaces this with a **scalar reward head** that outputs a single value representing human preference. According to the `FareedKhan-dev/train-llm-from-scratch` implementation, the reward model uses `forward_hidden` to extract representations and applies a linear layer mapping to a scalar score rather than vocabulary logits.

### Why use Bradley-Terry loss instead of regression loss?

**Bradley-Terry loss** directly optimizes for ranking consistency by maximizing the margin between chosen and rejected responses, whereas regression loss (MSE) requires absolute score labels that are difficult to calibrate across different annotators. The implementation in [`src/post_training/reward_train.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/reward_train.py) uses `-logsigmoid(chosen - rejected)`, which naturally handles the relative nature of human preferences without requiring normalized absolute scores.

### Should I freeze the transformer backbone during reward model training?

Yes, the default configuration in [`scripts/train_reward.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_reward.py) freezes the SFT backbone and updates only the **reward head parameters**. This preserves the language capabilities learned during supervised fine-tuning while adapting only the preference-scoring component. You can unfreeze backbone layers by explicitly passing those parameters to the optimizer if your dataset requires deeper capacity.

### How do I handle variable sequence lengths in the reward model?

Pass the `seq_lengths` tensor to the `forward` method, which uses `gather_last` from [`src/post_training/utils.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/utils.py) to index the reward at position `seq_lengths[i] - 1` for each batch element. This ensures the model extracts rewards from the **last real token** before any padding, maintaining valid gradients even with right-padded variable-length sequences.