How to Train a Reward Model for LLM Alignment: A Complete Implementation Guide
Training a reward model for LLM alignment involves attaching a scalar reward head to a supervised fine-tuned (SFT) backbone and optimizing it with Bradley-Terry pairwise loss on human preference data.
In the FareedKhan-dev/train-llm-from-scratch repository, reward model training represents the third stage of the RLHF pipeline. This component learns to assign scalar scores to generated responses based on human preferences, providing the critical signal needed for PPO or DPO fine-tuning. The implementation pairs a lightweight reward head with a pairwise ranking loss to efficiently capture human judgment patterns.
Architecture
The RewardModel Class
The RewardModel class in src/post_training/reward_model.py wraps your SFT transformer and projects hidden states into scalar reward values. Unlike the base model, it discards the language modeling head in favor of a single linear layer.
class RewardModel(nn.Module):
"""Wrap a Transformer and add a scalar reward head (no lm_head used)."""
def __init__(self, transformer: Transformer) -> None:
super().__init__()
self.transformer = transformer
n_embed = transformer.lm_head.in_features # ← hidden size
self.reward_head = nn.Linear(n_embed, 1, bias=False)
nn.init.zeros_(self.reward_head.weight) # start near‑zero rewards
def token_rewards(self, idx: torch.Tensor) -> torch.Tensor:
# (B, T) → per‑token scalar reward
hidden = self.transformer.forward_hidden(idx)
return self.reward_head(hidden).squeeze(-1)
def forward(self, idx: torch.Tensor,
seq_lengths: torch.Tensor | None = None) -> torch.Tensor:
# (B,) scalar reward per sequence
rewards = self.token_rewards(idx) # (B, T)
if seq_lengths is None:
return rewards[:, -1] # last token of full‑length seq
return gather_last(rewards, seq_lengths) # last *real* token
The scalar head is a nn.Linear layer mapping from the transformer's hidden dimension to a single value, initialized with zeros to start training from neutral rewards. The forward method extracts the hidden state of the last real token using gather_last (defined in src/post_training/utils.py), which indexes rewards[i, seq_lengths[i]‑1] to ignore padding tokens.
Bradley-Terry Loss Function
The training objective uses the Bradley-Terry pairwise comparison loss implemented in src/post_training/reward_train.py:
def bradley_tyrrey_loss(chosen, rejected):
# L = -log σ(chosen - rejected)
return -F.logsigmoid(chosen - rejected).mean()
This loss pushes the reward of the chosen response above that of the rejected response by maximizing the log-probability that the preferred response scores higher. The repository monitors preference accuracy (chosen > rejected) as the primary training metric alongside the reward margin.
Preference Data Flow
The data_loader/preference_dataset.py loader prepares tuples of prompt-response pairs with specific tensor formatting:
batch = {
"chosen_ids": Tensor(B, L), # token ids of prompt+chosen+EOT
"chosen_len": Tensor(B), # true length of each chosen sequence
"rejected_ids": Tensor(B, L), # token ids of prompt+rejected+EOT
"rejected_len": Tensor(B), # true length of each rejected sequence
}
Sequences are right-padded to uniform length, which remains safe because the underlying Transformer uses causal attention masks. The last real token never attends to padding tokens that follow it, ensuring valid reward computation via gather_last.
Training Procedure
Running the Training Loop
The entry point scripts/train_reward.py orchestrates the training process. A typical execution looks like:
PYTHONPATH=. python scripts/train_reward.py \
--cfg configs/post_training_config.yaml \
--data_path data/preference/train.jsonl \
--ckpt_dir ckpts/reward \
--lr 1e-5 --max_len 768
Inside the training loop, the implementation concatenates chosen and rejected sequences for computational efficiency:
# Build the SFT backbone from the config
backbone = build_model_from_config(cfg)
rm = RewardModel(backbone).to(device)
# Preference iterator yields collated tensors
iterator = get_preference_iterator(
path=args.data_path,
batch_size=args.batch_size,
max_len=args.max_len,
device=device,
)
optim = torch.optim.AdamW(rm.parameters(), lr=args.lr)
for batch in iterator:
# Concatenate chosen + rejected so we get a single forward pass (2B rows)
ids = torch.cat([batch["chosen_ids"], batch["rejected_ids"]], dim=0)
lens = torch.cat([batch["chosen_len"], batch["rejected_len"]], dim=0)
# Forward → (2B,) rewards
rewards = rm(ids, seq_lengths=lens).float()
chosen_r, rejected_r = rewards[:B], rewards[B:]
loss = bradley_tyrrey_loss(chosen_r, rejected_r)
loss.backward()
optim.step()
optim.zero_grad()
By processing both sides of the preference pair in a single forward pass (2B rows), the implementation maximizes GPU utilization. After the forward pass, the reward tensor splits back into chosen_r and rejected_r for loss computation. By default, the optimizer updates only the reward head parameters while the SFT backbone remains frozen.
Saving and Loading Checkpoints
Save checkpoints using standard PyTorch serialization:
torch.save({"model_state_dict": rm.state_dict()}, "ckpts/reward/reward.pt")
For inference, use the load_reward_model helper from src/post_training/reward_model.py, which handles Distributed Data Parallel (DDP) prefix stripping (module.) and automatically sets evaluation mode:
from src.post_training.reward_model import load_reward_model
from src.config.loader import load_config
cfg = load_config("configs/post_training_config.yaml")
rm = load_reward_model(cfg, "ckpts/reward/reward.pt", device="cpu")
Practical Code Examples
Instantiating from a Checkpoint
Load a trained reward model for evaluation or downstream RL training:
from src.post_training.reward_model import load_reward_model
from src.config.loader import load_config
cfg = load_config("configs/post_training_config.yaml")
rm = load_reward_model(cfg, "ckpts/reward/reward.pt", device="cpu")
Scoring Individual Responses
Generate scalar rewards for specific model outputs:
prompt = "Explain the greenhouse effect in one sentence."
response = "Greenhouse gases trap heat, warming the planet."
ids, _ = encode_chat([{"role":"user","content":prompt},
{"role":"assistant","content":response}])
ids = torch.tensor([ids], dtype=torch.long) # (1, T)
reward = rm(ids) # → scalar tensor
print("Reward:", reward.item())
Integrating with PPO
Use the trained reward model to provide feedback signals during PPO fine-tuning:
from src.post_training.ppo import PPOTrainer # simplified view
trainer = PPOTrainer(
policy=model, # the SFT model
reward_model=rm, # the trained RM
cfg=ppo_cfg,
)
trainer.train()
The PPO trainer internally calls rm(ids, seq_lengths) for each rollout to compute advantage estimates and policy gradients.
Summary
Training a reward model for LLM alignment in this repository follows a structured pipeline:
- Architecture: Attach a zero-initialized linear reward head to the SFT backbone, extracting rewards from the last real token using
gather_last - Data: Load preference pairs via
get_preference_iterator, which returns right-padded tensors with explicit sequence lengths - Training: Concatenate chosen and rejected batches for efficiency, compute Bradley-Terry loss, and optimize with AdamW while monitoring preference accuracy
- Inference: Load trained models using
load_reward_modelto remove DDP prefixes and enable evaluation mode - Integration: Feed the reward model into PPO or DPO trainers to provide human preference signals for policy optimization
Frequently Asked Questions
What is the difference between a reward model and the base LLM?
The base LLM predicts next-token probabilities using a language modeling head, while the reward model replaces this with a scalar reward head that outputs a single value representing human preference. According to the FareedKhan-dev/train-llm-from-scratch implementation, the reward model uses forward_hidden to extract representations and applies a linear layer mapping to a scalar score rather than vocabulary logits.
Why use Bradley-Terry loss instead of regression loss?
Bradley-Terry loss directly optimizes for ranking consistency by maximizing the margin between chosen and rejected responses, whereas regression loss (MSE) requires absolute score labels that are difficult to calibrate across different annotators. The implementation in src/post_training/reward_train.py uses -logsigmoid(chosen - rejected), which naturally handles the relative nature of human preferences without requiring normalized absolute scores.
Should I freeze the transformer backbone during reward model training?
Yes, the default configuration in scripts/train_reward.py freezes the SFT backbone and updates only the reward head parameters. This preserves the language capabilities learned during supervised fine-tuning while adapting only the preference-scoring component. You can unfreeze backbone layers by explicitly passing those parameters to the optimizer if your dataset requires deeper capacity.
How do I handle variable sequence lengths in the reward model?
Pass the seq_lengths tensor to the forward method, which uses gather_last from src/post_training/utils.py to index the reward at position seq_lengths[i] - 1 for each batch element. This ensures the model extracts rewards from the last real token before any padding, maintaining valid gradients even with right-padded variable-length sequences.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →