How GRPO, GSPO, PPO, and REINFORCE++ Differ in Implementation: A Deep Dive into the Miles Codebase
GRPO and GSPO use group-wise advantage normalization without a value function, PPO relies on clipped surrogate objectives with a value head, and REINFORCE++ applies pure policy gradients with optional regularizers.
The Miles repository implements four on-policy reinforcement learning algorithms for LLM fine-tuning: GRPO (Group-Regularized Policy Optimization), GSPO (Group-Structured Policy Optimization), PPO (Proximal Policy Optimization), and REINFORCE++. While they share a common training pipeline—data loading, rollout generation, loss aggregation, and optimizer steps—each algorithm diverges in three critical areas: advantage estimation, KL penalty computation, and policy update rules. This analysis breaks down the implementation differences using actual source files from radixark/miles.
Advantage Estimation: Group vs. Value-Function Baselines
GRPO: Group-Wise Centered and Scaled Advantages
GRPO calculates a group-wise advantage where all sampled completions for a single prompt form a group. The per-sample advantage is first centered by subtracting the group mean, then scaled by the group-wise standard deviation (or clipped variance). This creates an implicit baseline that eliminates the need for a learned value function.
The implementation lives in miles/rollout/on_policy_distillation.py within the compute_group_advantage function:
# Conceptual flow from compute_group_advantage
advantages = (rewards - group_mean) / (group_std + eps)
The --use-grpo flag in miles/utils/arguments.py activates this path.
GSPO: Same Grouping, Different KL Treatment
GSPO shares the identical grouping and advantage normalization logic with GRPO—both use compute_group_advantage and the RolloutBatch data structure with group_id fields from miles/rollout/batch.py. The divergence appears in how KL penalties are computed (see next section).
PPO: Value-Function Baseline with Reward-to-Go
PPO uses a value-function baseline (or a moving-average baseline when no value head exists). The advantage equals reward-to-go minus baseline. No group-level operations occur. This branch resides in miles/backends/training_utils/loss_hub/losses.py:
# PPO advantage: standard TD(λ) or GAE
advantages = returns - value_preds
REINFORCE++: Pure Policy Gradient with Variance Reduction
REINFORCE++ implements the classic gradient ∇θ log πθ · R augmented with variance-reduction tricks: a baseline, entropy bonus, and optional KL penalty. The "++" denotes these regularizers plus a smoothed reward that bypasses clipping. The path is handled in miles/rollout/on_policy_distillation.py under the REINFORCE++ branch, with support for custom reducer functions via --custom-pg-loss-reducer-function-path.
KL Penalty Computation: Per-Token vs. Per-Sequence
The KL divergence treatment creates the sharpest distinction between algorithms.
| Algorithm | KL Scope | Implementation Details |
|---|---|---|
| GRPO | Per-token | Computed between current policy and reference checkpoint; added as penalty λ · KL in loss |
| GSPO | Per-sequence | KL summed over all tokens, then averaged across group; matches GSPO paper formulation |
| PPO | Per-token (implicit) | Not a hard penalty; clipping mechanism clip(ratio, 1-ε, 1+ε) bounds KL implicitly |
| REINFORCE++ | Optional per-token | Treated as regularizer when supplied; no clipping enforced |
The compute_kl function in miles/backends/training_utils/loss_hub/losses.py accepts a per_sequence flag—True for GSPO, False for GRPO and PPO.
Policy Update Rules: Clipping and Gradient Formulation
GRPO: Unclipped Group-Normalized Gradient
# From grpo_policy_loss in losses.py
loss = -advantage * log_prob + kl_coef * kl_penalty
The advantage arrives pre-normalized from compute_group_advantage, eliminating the need for ratio clipping.
GSPO: Identical Structure, Sequence-Level KL
# From gspo_policy_loss in losses.py
loss = -advantage * log_prob + kl_coef * sequence_kl_penalty
Same update rule as GRPO, but kl_penalty aggregates per-sequence.
PPO: Clipped Surrogate Objective
# From ppo_policy_loss in losses.py
ratio = torch.exp(log_prob - old_log_prob)
clipped_ratio = torch.clamp(ratio, 1 - eps, 1 + eps)
loss = -torch.min(ratio * advantage, clipped_ratio * advantage)
The only algorithm with torch.clamp on policy ratios. This is PPO's signature trust-region mechanism.
REINFORCE++: Unclipped with Entropy Regularization
# From on_policy_distillation.py REINFORCE++ path
loss = -advantage * log_prob + kl_coef * kl_penalty - entropy_coef * entropy
No clipping. Entropy bonus is explicit and configurable.
Group Formation and Data Structures
Both GRPO and GSPO depend on the RolloutBatch class in miles/rollout/batch.py. This structure tracks group_id for each sample, enabling the advantage function to aggregate statistics across completions belonging to the same prompt. The group size is controlled via CLI:
# GRPO run with group size 8
python -m miles.main.scripts.run_qwen3_dense \
--model qwen3-4b \
--algo grpo \
--group-size 8 \
--kl-coef 0.1
# GSPO run with larger groups
python -m miles.main.scripts.run_qwen3_next_80b_a3b \
--model qwen3-next-80b \
--algo gspo \
--group-size 16 \
--kl-coef 0.05
Practical Launch Examples
PPO Configuration
python -m miles.main.examples.ppo.run_qwen3_4b_ppo \
--model qwen3-4b \
--algo ppo \
--clip-eps 0.2 \
--kl-coef 0.01
The clip-eps parameter controls the trust region—unique to PPO.
REINFORCE++ Configuration
python -m miles.main.examples.on_policy_distillation.run \
--model qwen3-5b \
--algo reinforce_pp \
--entropy-coef 0.01 \
--kl-coef 0.0
Setting kl-coef to zero demonstrates the pure policy-gradient baseline.
Key Implementation Files
miles/backends/training_utils/loss_hub/losses.py— Core loss implementations:grpo_policy_loss,gspo_policy_loss,ppo_policy_loss, andcompute_klmiles/rollout/on_policy_distillation.py— Advantage computation and REINFORCE++ logicmiles/utils/arguments.py— CLI flags:--algo,--use-grpo,--use-gspo,--group-sizemiles/rollout/batch.py—RolloutBatchwithgroup_idtracking
Summary
- GRPO eliminates value functions via group-wise advantage normalization and uses per-token KL penalties
- GSPO matches GRPO's grouping but switches to per-sequence KL for metric alignment
- PPO is the only algorithm with clipped policy ratios and explicit value-function baselines
- REINFORCE++ provides the simplest gradient form with optional regularizers and no clipping
Frequently Asked Questions
When should I choose GRPO over PPO in Miles?
Select GRPO when you can sample multiple completions per prompt (large group sizes) and want to avoid training a value head. The group baseline provides strong variance reduction without additional parameters. PPO remains preferable when you have a reliable value head and need the stability of clipped updates.
What is the practical difference between GRPO and GSPO?
The algorithms differ only in KL penalty granularity. GRPO penalizes per-token KL divergence, which can be noisier but more fine-grained. GSPO's per-sequence KL better matches downstream evaluation metrics that score entire outputs. Use GSPO when token-level KL correlates poorly with your target metric.
Does REINFORCE++ support the same features as GRPO?
REINFORCE++ supports entropy regularization and optional KL penalties like GRPO, but lacks group-wise advantage normalization. It serves simpler research prototypes or scenarios requiring pure policy gradients. The --custom-pg-loss-reducer-function-path flag enables custom variance-reduction strategies not available in other algorithms.
Where is the clipping operation implemented in Miles PPO?
The clipping occurs in miles/backends/training_utils/loss_hub/losses.py within ppo_policy_loss. Specifically, torch.clamp(ratio, 1 - eps, 1 + eps) bounds the policy ratio. GRPO, GSPO, and REINFORCE++ contain no equivalent clamping operation—their trust regions come from implicit constraints (group normalization) or none at all.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →