# What Reinforcement Learning Algorithms Are Implemented in the CartPole Lab?

> Explore the CartPole lab in AI-For-Beginners and discover its actor-critic reinforcement learning algorithm implemented in TensorFlow and PyTorch.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: deep-dive
- Published: 2026-08-23

---

**The CartPole lab in Microsoft's AI-For-Beginners repository implements a single reinforcement learning algorithm: the actor-critic (policy-gradient) method, available in both TensorFlow/Keras and PyTorch implementations.**

The CartPole balancing task serves as the introductory reinforcement learning exercise in the `microsoft/AI-For-Beginners` curriculum. Located in the deep reinforcement learning section at `/lessons/6-Other/22-DeepRL/`, this lab focuses specifically on policy-gradient methods rather than value-based alternatives, making it ideal for beginners to understand how neural networks directly learn control policies through gradient ascent.

## Actor-Critic (Policy-Gradient) Algorithm Overview

The lab implements an **actor-critic** architecture, a hybrid reinforcement learning algorithm that combines policy optimization with value estimation. According to the source code in both `/lessons/6-Other/22-DeepRL/CartPole-RL-TF.ipynb` and `/lessons/6-Other/22-DeepRL/CartPole-RL-PyTorch.ipynb`, the implementation uses two separate neural networks:

- **Actor network**: Learns the policy directly by mapping environment states to action logits (probabilities)
- **Critic network**: Estimates state values to provide a baseline, reducing variance in the policy gradient updates

Both notebooks utilize the **REINFORCE gradient** with a baseline (the critic's value estimate) to stabilize training. This approach maximizes expected return by adjusting policy parameters in the direction of greater cumulative reward, while the critic's value estimates help distinguish whether actions were genuinely better than average or merely lucky.

## TensorFlow/Keras Implementation

The TensorFlow version located at `/lessons/6-Other/22-DeepRL/CartPole-RL-TF.ipynb` constructs separate Keras models for the actor and critic using the functional API. The architecture features two hidden layers with 24 units and ReLU activations, mapping the 4-dimensional CartPole observation space to action probabilities and state values.

```python
import tensorflow as tf
from tensorflow.keras import layers

# Simple policy network (actor) – maps state → action logits

def build_policy_network(state_shape, action_dim):
    inputs = tf.keras.Input(shape=state_shape)
    x = layers.Dense(24, activation="relu")(inputs)
    x = layers.Dense(24, activation="relu")(x)
    action_logits = layers.Dense(action_dim)(x)
    return tf.keras.Model(inputs, action_logits)

# Value network (critic) – maps state → state‑value estimate

def build_value_network(state_shape):
    inputs = tf.keras.Input(shape=state_shape)
    x = layers.Dense(24, activation="relu")(inputs)
    x = layers.Dense(24, activation="relu")(x)
    value = layers.Dense(1)(x)
    return tf.keras.Model(inputs, value)

```

## PyTorch Implementation

The PyTorch counterpart at `/lessons/6-Other/22-DeepRL/CartPole-RL-PyTorch.ipynb` provides equivalent functionality using `nn.Module` subclasses. Both networks maintain identical architecture—two hidden layers of 24 units with ReLU activations—ensuring consistency across deep learning frameworks while demonstrating framework-agnostic RL concepts.

```python
import torch
import torch.nn as nn
import torch.nn.functional as F

class PolicyNetwork(nn.Module):
    def __init__(self, obs_dim, n_actions):
        super().__init__()
        self.fc1 = nn.Linear(obs_dim, 24)
        self.fc2 = nn.Linear(24, 24)
        self.logits = nn.Linear(24, n_actions)

    def forward(self, x):
        x = F.relu(self.fc1(x))
        x = F.relu(self.fc2(x))
        return self.logits(x)

class ValueNetwork(nn.Module):
    def __init__(self, obs_dim):
        super().__init__()
        self.fc1 = nn.Linear(obs_dim, 24)
        self.fc2 = nn.Linear(24, 24)
        self.value = nn.Linear(24, 1)

    def forward(self, x):
        x = F.relu(self.fc1(x))
        x = F.relu(self.fc2(x))
        return self.value(x)

```

## Why Policy-Gradient for This Lab?

Unlike comprehensive RL suites that implement DQN, SARSA, or Monte-Carlo methods, the microsoft/AI-For-Beginners repository intentionally focuses solely on the actor-critic approach for the CartPole lab. This pedagogical choice emphasizes **educational clarity over algorithmic breadth**, allowing beginners to grasp policy-gradient fundamentals—how neural networks parametrize stochastic policies and how advantage estimates improve gradient stability—before encountering value-based or model-based alternatives.

## Summary

- The CartPole lab implements exclusively the **actor-critic (policy-gradient)** reinforcement learning algorithm.
- Two complete implementations exist: **TensorFlow/Keras** (`CartPole-RL-TF.ipynb`) and **PyTorch** (`CartPole-RL-PyTorch.ipynb`).
- The architecture separates **actor** (policy network) and **critic** (value network) into distinct neural networks with 24-unit hidden layers.
- Training utilizes **REINFORCE with baseline**, where the critic's value estimates reduce variance in policy gradient updates.
- No other RL algorithms (DQN, SARSA, Q-learning) are included in this specific lab, maintaining focus on policy-gradient fundamentals.

## Frequently Asked Questions

### Does the CartPole lab implement DQN or other value-based methods?

No, the lab does not implement Deep Q-Networks (DQN), SARSA, or other value-based algorithms. Both notebook files exclusively demonstrate the actor-critic policy-gradient method, concentrating on direct policy optimization rather than learning action-value functions.

### What is the difference between the actor and critic networks in this implementation?

The **actor network** outputs action logits representing the probability distribution over possible actions (left or right), effectively defining the agent's behavior policy. The **critic network** estimates the expected cumulative reward (value) from a given state, providing a baseline that helps determine whether selected actions performed better or worse than average, thus reducing variance during gradient updates.

### Which deep learning frameworks are supported in the CartPole RL lab?

The lab provides complete implementations in both **TensorFlow/Keras** and **PyTorch**. The TensorFlow version uses the Keras functional API to define `build_policy_network()` and `build_value_network()` functions, while the PyTorch version implements `PolicyNetwork` and `ValueNetwork` classes inheriting from `nn.Module`.

### Where can I find the training loop and gradient updates in the source files?

The training logic resides within the respective Jupyter notebooks at `/lessons/6-Other/22-DeepRL/CartPole-RL-TF.ipynb` and `/lessons/6-Other/22-DeepRL/CartPole-RL-PyTorch.ipynb`. Both implementations calculate policy gradients using the log-probability of actions weighted by advantages (returns minus critic values), then apply standard backpropagation to update the network weights.