What Reinforcement Learning Algorithms Are Implemented in the CartPole Lab?
The CartPole lab in Microsoft's AI-For-Beginners repository implements a single reinforcement learning algorithm: the actor-critic (policy-gradient) method, available in both TensorFlow/Keras and PyTorch implementations.
The CartPole balancing task serves as the introductory reinforcement learning exercise in the microsoft/AI-For-Beginners curriculum. Located in the deep reinforcement learning section at /lessons/6-Other/22-DeepRL/, this lab focuses specifically on policy-gradient methods rather than value-based alternatives, making it ideal for beginners to understand how neural networks directly learn control policies through gradient ascent.
Actor-Critic (Policy-Gradient) Algorithm Overview
The lab implements an actor-critic architecture, a hybrid reinforcement learning algorithm that combines policy optimization with value estimation. According to the source code in both /lessons/6-Other/22-DeepRL/CartPole-RL-TF.ipynb and /lessons/6-Other/22-DeepRL/CartPole-RL-PyTorch.ipynb, the implementation uses two separate neural networks:
- Actor network: Learns the policy directly by mapping environment states to action logits (probabilities)
- Critic network: Estimates state values to provide a baseline, reducing variance in the policy gradient updates
Both notebooks utilize the REINFORCE gradient with a baseline (the critic's value estimate) to stabilize training. This approach maximizes expected return by adjusting policy parameters in the direction of greater cumulative reward, while the critic's value estimates help distinguish whether actions were genuinely better than average or merely lucky.
TensorFlow/Keras Implementation
The TensorFlow version located at /lessons/6-Other/22-DeepRL/CartPole-RL-TF.ipynb constructs separate Keras models for the actor and critic using the functional API. The architecture features two hidden layers with 24 units and ReLU activations, mapping the 4-dimensional CartPole observation space to action probabilities and state values.
import tensorflow as tf
from tensorflow.keras import layers
# Simple policy network (actor) – maps state → action logits
def build_policy_network(state_shape, action_dim):
inputs = tf.keras.Input(shape=state_shape)
x = layers.Dense(24, activation="relu")(inputs)
x = layers.Dense(24, activation="relu")(x)
action_logits = layers.Dense(action_dim)(x)
return tf.keras.Model(inputs, action_logits)
# Value network (critic) – maps state → state‑value estimate
def build_value_network(state_shape):
inputs = tf.keras.Input(shape=state_shape)
x = layers.Dense(24, activation="relu")(inputs)
x = layers.Dense(24, activation="relu")(x)
value = layers.Dense(1)(x)
return tf.keras.Model(inputs, value)
PyTorch Implementation
The PyTorch counterpart at /lessons/6-Other/22-DeepRL/CartPole-RL-PyTorch.ipynb provides equivalent functionality using nn.Module subclasses. Both networks maintain identical architecture—two hidden layers of 24 units with ReLU activations—ensuring consistency across deep learning frameworks while demonstrating framework-agnostic RL concepts.
import torch
import torch.nn as nn
import torch.nn.functional as F
class PolicyNetwork(nn.Module):
def __init__(self, obs_dim, n_actions):
super().__init__()
self.fc1 = nn.Linear(obs_dim, 24)
self.fc2 = nn.Linear(24, 24)
self.logits = nn.Linear(24, n_actions)
def forward(self, x):
x = F.relu(self.fc1(x))
x = F.relu(self.fc2(x))
return self.logits(x)
class ValueNetwork(nn.Module):
def __init__(self, obs_dim):
super().__init__()
self.fc1 = nn.Linear(obs_dim, 24)
self.fc2 = nn.Linear(24, 24)
self.value = nn.Linear(24, 1)
def forward(self, x):
x = F.relu(self.fc1(x))
x = F.relu(self.fc2(x))
return self.value(x)
Why Policy-Gradient for This Lab?
Unlike comprehensive RL suites that implement DQN, SARSA, or Monte-Carlo methods, the microsoft/AI-For-Beginners repository intentionally focuses solely on the actor-critic approach for the CartPole lab. This pedagogical choice emphasizes educational clarity over algorithmic breadth, allowing beginners to grasp policy-gradient fundamentals—how neural networks parametrize stochastic policies and how advantage estimates improve gradient stability—before encountering value-based or model-based alternatives.
Summary
- The CartPole lab implements exclusively the actor-critic (policy-gradient) reinforcement learning algorithm.
- Two complete implementations exist: TensorFlow/Keras (
CartPole-RL-TF.ipynb) and PyTorch (CartPole-RL-PyTorch.ipynb). - The architecture separates actor (policy network) and critic (value network) into distinct neural networks with 24-unit hidden layers.
- Training utilizes REINFORCE with baseline, where the critic's value estimates reduce variance in policy gradient updates.
- No other RL algorithms (DQN, SARSA, Q-learning) are included in this specific lab, maintaining focus on policy-gradient fundamentals.
Frequently Asked Questions
Does the CartPole lab implement DQN or other value-based methods?
No, the lab does not implement Deep Q-Networks (DQN), SARSA, or other value-based algorithms. Both notebook files exclusively demonstrate the actor-critic policy-gradient method, concentrating on direct policy optimization rather than learning action-value functions.
What is the difference between the actor and critic networks in this implementation?
The actor network outputs action logits representing the probability distribution over possible actions (left or right), effectively defining the agent's behavior policy. The critic network estimates the expected cumulative reward (value) from a given state, providing a baseline that helps determine whether selected actions performed better or worse than average, thus reducing variance during gradient updates.
Which deep learning frameworks are supported in the CartPole RL lab?
The lab provides complete implementations in both TensorFlow/Keras and PyTorch. The TensorFlow version uses the Keras functional API to define build_policy_network() and build_value_network() functions, while the PyTorch version implements PolicyNetwork and ValueNetwork classes inheriting from nn.Module.
Where can I find the training loop and gradient updates in the source files?
The training logic resides within the respective Jupyter notebooks at /lessons/6-Other/22-DeepRL/CartPole-RL-TF.ipynb and /lessons/6-Other/22-DeepRL/CartPole-RL-PyTorch.ipynb. Both implementations calculate policy gradients using the log-probability of actions weighted by advantages (returns minus critic values), then apply standard backpropagation to update the network weights.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →