# How GRPO, GSPO, PPO, and REINFORCE++ Differ in Implementation: A Deep Dive into the Miles Codebase

> Explore GRPO, GSPO, PPO, and REINFORCE++ implementation differences in Miles. Understand their core mechanics for advanced reinforcement learning.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: deep-dive
- Published: 2026-09-06

---

**GRPO and GSPO use group-wise advantage normalization without a value function, PPO relies on clipped surrogate objectives with a value head, and REINFORCE++ applies pure policy gradients with optional regularizers.**

The Miles repository implements four on-policy reinforcement learning algorithms for LLM fine-tuning: **GRPO** (Group-Regularized Policy Optimization), **GSPO** (Group-Structured Policy Optimization), **PPO** (Proximal Policy Optimization), and **REINFORCE++**. While they share a common training pipeline—data loading, rollout generation, loss aggregation, and optimizer steps—each algorithm diverges in three critical areas: **advantage estimation**, **KL penalty computation**, and **policy update rules**. This analysis breaks down the implementation differences using actual source files from `radixark/miles`.

## Advantage Estimation: Group vs. Value-Function Baselines

### GRPO: Group-Wise Centered and Scaled Advantages

GRPO calculates a **group-wise advantage** where all sampled completions for a single prompt form a group. The per-sample advantage is first **centered by subtracting the group mean**, then **scaled by the group-wise standard deviation** (or clipped variance). This creates an **implicit baseline** that eliminates the need for a learned value function.

The implementation lives in [`miles/rollout/on_policy_distillation.py`](https://github.com/radixark/miles/blob/main/miles/rollout/on_policy_distillation.py) within the `compute_group_advantage` function:

```python

# Conceptual flow from compute_group_advantage

advantages = (rewards - group_mean) / (group_std + eps)

```

The `--use-grpo` flag in [`miles/utils/arguments.py`](https://github.com/radixark/miles/blob/main/miles/utils/arguments.py) activates this path.

### GSPO: Same Grouping, Different KL Treatment

GSPO shares the identical grouping and advantage normalization logic with GRPO—both use `compute_group_advantage` and the `RolloutBatch` data structure with `group_id` fields from [`miles/rollout/batch.py`](https://github.com/radixark/miles/blob/main/miles/rollout/batch.py). The divergence appears in how KL penalties are computed (see next section).

### PPO: Value-Function Baseline with Reward-to-Go

PPO uses a **value-function baseline** (or a moving-average baseline when no value head exists). The advantage equals **reward-to-go minus baseline**. No group-level operations occur. This branch resides in [`miles/backends/training_utils/loss_hub/losses.py`](https://github.com/radixark/miles/blob/main/miles/backends/training_utils/loss_hub/losses.py):

```python

# PPO advantage: standard TD(λ) or GAE

advantages = returns - value_preds

```

### REINFORCE++: Pure Policy Gradient with Variance Reduction

REINFORCE++ implements the classic gradient `∇θ log πθ · R` augmented with **variance-reduction tricks**: a baseline, entropy bonus, and optional KL penalty. The "++" denotes these regularizers plus a **smoothed reward** that bypasses clipping. The path is handled in [`miles/rollout/on_policy_distillation.py`](https://github.com/radixark/miles/blob/main/miles/rollout/on_policy_distillation.py) under the REINFORCE++ branch, with support for custom reducer functions via `--custom-pg-loss-reducer-function-path`.

## KL Penalty Computation: Per-Token vs. Per-Sequence

The KL divergence treatment creates the sharpest distinction between algorithms.

| Algorithm | KL Scope | Implementation Details |
|-----------|----------|------------------------|
| **GRPO** | Per-token | Computed between current policy and reference checkpoint; added as penalty `λ · KL` in loss |
| **GSPO** | Per-sequence | KL summed over all tokens, then averaged across group; matches GSPO paper formulation |
| **PPO** | Per-token (implicit) | Not a hard penalty; clipping mechanism `clip(ratio, 1-ε, 1+ε)` bounds KL implicitly |
| **REINFORCE++** | Optional per-token | Treated as regularizer when supplied; no clipping enforced |

The `compute_kl` function in [`miles/backends/training_utils/loss_hub/losses.py`](https://github.com/radixark/miles/blob/main/miles/backends/training_utils/loss_hub/losses.py) accepts a `per_sequence` flag—`True` for GSPO, `False` for GRPO and PPO.

## Policy Update Rules: Clipping and Gradient Formulation

### GRPO: Unclipped Group-Normalized Gradient

```python

# From grpo_policy_loss in losses.py

loss = -advantage * log_prob + kl_coef * kl_penalty

```

The advantage arrives pre-normalized from `compute_group_advantage`, eliminating the need for ratio clipping.

### GSPO: Identical Structure, Sequence-Level KL

```python

# From gspo_policy_loss in losses.py

loss = -advantage * log_prob + kl_coef * sequence_kl_penalty

```

Same update rule as GRPO, but `kl_penalty` aggregates per-sequence.

### PPO: Clipped Surrogate Objective

```python

# From ppo_policy_loss in losses.py

ratio = torch.exp(log_prob - old_log_prob)
clipped_ratio = torch.clamp(ratio, 1 - eps, 1 + eps)
loss = -torch.min(ratio * advantage, clipped_ratio * advantage)

```

The **only algorithm with `torch.clamp`** on policy ratios. This is PPO's signature trust-region mechanism.

### REINFORCE++: Unclipped with Entropy Regularization

```python

# From on_policy_distillation.py REINFORCE++ path

loss = -advantage * log_prob + kl_coef * kl_penalty - entropy_coef * entropy

```

No clipping. Entropy bonus is explicit and configurable.

## Group Formation and Data Structures

Both GRPO and GSPO depend on the `RolloutBatch` class in [`miles/rollout/batch.py`](https://github.com/radixark/miles/blob/main/miles/rollout/batch.py). This structure tracks `group_id` for each sample, enabling the advantage function to aggregate statistics across completions belonging to the same prompt. The group size is controlled via CLI:

```python

# GRPO run with group size 8

python -m miles.main.scripts.run_qwen3_dense \
    --model qwen3-4b \
    --algo grpo \
    --group-size 8 \
    --kl-coef 0.1

```

```python

# GSPO run with larger groups

python -m miles.main.scripts.run_qwen3_next_80b_a3b \
    --model qwen3-next-80b \
    --algo gspo \
    --group-size 16 \
    --kl-coef 0.05

```

## Practical Launch Examples

### PPO Configuration

```python
python -m miles.main.examples.ppo.run_qwen3_4b_ppo \
    --model qwen3-4b \
    --algo ppo \
    --clip-eps 0.2 \
    --kl-coef 0.01

```

The `clip-eps` parameter controls the trust region—unique to PPO.

### REINFORCE++ Configuration

```python
python -m miles.main.examples.on_policy_distillation.run \
    --model qwen3-5b \
    --algo reinforce_pp \
    --entropy-coef 0.01 \
    --kl-coef 0.0

```

Setting `kl-coef` to zero demonstrates the pure policy-gradient baseline.

## Key Implementation Files

- **[`miles/backends/training_utils/loss_hub/losses.py`](https://github.com/radixark/miles/blob/main/miles/backends/training_utils/loss_hub/losses.py)** — Core loss implementations: `grpo_policy_loss`, `gspo_policy_loss`, `ppo_policy_loss`, and `compute_kl`
- **[`miles/rollout/on_policy_distillation.py`](https://github.com/radixark/miles/blob/main/miles/rollout/on_policy_distillation.py)** — Advantage computation and REINFORCE++ logic
- **[`miles/utils/arguments.py`](https://github.com/radixark/miles/blob/main/miles/utils/arguments.py)** — CLI flags: `--algo`, `--use-grpo`, `--use-gspo`, `--group-size`
- **[`miles/rollout/batch.py`](https://github.com/radixark/miles/blob/main/miles/rollout/batch.py)** — `RolloutBatch` with `group_id` tracking

## Summary

- **GRPO** eliminates value functions via group-wise advantage normalization and uses per-token KL penalties
- **GSPO** matches GRPO's grouping but switches to per-sequence KL for metric alignment
- **PPO** is the only algorithm with clipped policy ratios and explicit value-function baselines
- **REINFORCE++** provides the simplest gradient form with optional regularizers and no clipping

## Frequently Asked Questions

### When should I choose GRPO over PPO in Miles?

Select **GRPO** when you can sample multiple completions per prompt (large group sizes) and want to avoid training a value head. The group baseline provides strong variance reduction without additional parameters. PPO remains preferable when you have a reliable value head and need the stability of clipped updates.

### What is the practical difference between GRPO and GSPO?

The algorithms differ only in **KL penalty granularity**. GRPO penalizes per-token KL divergence, which can be noisier but more fine-grained. GSPO's per-sequence KL better matches downstream evaluation metrics that score entire outputs. Use GSPO when token-level KL correlates poorly with your target metric.

### Does REINFORCE++ support the same features as GRPO?

**REINFORCE++** supports entropy regularization and optional KL penalties like GRPO, but lacks group-wise advantage normalization. It serves simpler research prototypes or scenarios requiring pure policy gradients. The `--custom-pg-loss-reducer-function-path` flag enables custom variance-reduction strategies not available in other algorithms.

### Where is the clipping operation implemented in Miles PPO?

The clipping occurs in [`miles/backends/training_utils/loss_hub/losses.py`](https://github.com/radixark/miles/blob/main/miles/backends/training_utils/loss_hub/losses.py) within `ppo_policy_loss`. Specifically, `torch.clamp(ratio, 1 - eps, 1 + eps)` bounds the policy ratio. GRPO, GSPO, and REINFORCE++ contain no equivalent clamping operation—their trust regions come from implicit constraints (group normalization) or none at all.