How to Generate Text Using RNNs in AI for Beginners: A Complete Implementation Guide
Recurrent Neural Networks generate text by maintaining a hidden state across sequential tokens, using architectures like SimpleRNN and LSTM to predict the next character or word based on previous context.
The Microsoft AI-For-Beginners curriculum provides hands-on implementations of text generation using RNNs in its Natural Language Processing section. Located in lessons/5-NLP/, these notebooks demonstrate complete pipelines from basic recurrent layers to advanced character-level language models using both TensorFlow and PyTorch.
RNN Fundamentals in the Curriculum
The AI-For-Beginners repository organizes its RNN material across two core lessons. The RNNTF.ipynb notebook in lessons/5-NLP/16-RNN/ introduces sequential processing concepts, while the generative notebooks in lessons/5-NLP/17-GenerativeNetworks/ implement end-to-end text generation systems.
SimpleRNN and Hidden States
A SimpleRNN layer processes text by updating a single hidden vector at each time step. In lessons/5-NLP/16-RNN/RNNTF.ipynb, the implementation combines a TextVectorization layer, an Embedding layer, and a SimpleRNN layer for basic classification:
import tensorflow as tf
from tensorflow import keras
vocab_size = 20000
embed_dim = 64
vectorizer = keras.layers.TextVectorization(max_tokens=vocab_size)
model = keras.Sequential([
vectorizer,
keras.layers.Embedding(vocab_size, embed_dim),
keras.layers.SimpleRNN(16),
keras.layers.Dense(4, activation='softmax')
])
model.compile(loss='sparse_categorical_crossentropy',
optimizer='adam',
metrics=['accuracy'])
This architecture maintains a hidden state dimension of 16, passing the recurrent output through a final dense classifier.
Long Short-Term Memory (LSTM)
LSTM layers mitigate the vanishing gradient problem by maintaining both a cell state (c) and a hidden state (h). The same notebook demonstrates replacing SimpleRNN with keras.layers.LSTM() for improved long-range dependency modeling, which is essential for coherent text generation across longer sequences.
Character-Level Text Generation Pipeline
The complete workflow for generating text using RNNs in AI for Beginners follows four distinct stages: corpus preparation, model architecture design, next-token training, and probabilistic sampling.
Step 1: Corpus Tokenization
The GenerativeTF.ipynb notebook implements character-level modeling by mapping each unique character to an integer index. This approach treats text as a sequence of discrete symbols rather than words, reducing vocabulary complexity:
import string
import numpy as np
text = open('sample.txt').read().lower()
vocab = sorted(set(text))
char2idx = {c: i+1 for i, c in enumerate(vocab)} # Reserve 0 for padding
idx2char = np.array([''] + list(vocab))
seq_len = 100
step = 1
sequences = [text[i:i+seq_len] for i in range(0, len(text)-seq_len, step)]
X = np.array([[char2idx[c] for c in seq] for seq in sequences])
y = np.array([char2idx[seq[-1]] for seq in sequences])
This creates input sequences of 100 characters (seq_len) and targets the subsequent character for prediction.
Step 2: LSTM Architecture Design
The generative model stacks an Embedding layer with masking enabled, followed by an LSTM layer and a dense output layer with softmax activation:
model = keras.Sequential([
keras.layers.Embedding(len(vocab)+1, 128, mask_zero=True),
keras.layers.LSTM(256),
keras.layers.Dense(len(vocab)+1, activation='softmax')
])
model.compile(loss='sparse_categorical_crossentropy', optimizer='adam')
The mask_zero=True parameter instructs Keras to ignore padding tokens (index 0) during backpropagation, improving training efficiency when handling variable-length sequences.
Step 3: Training Configuration
The model trains on categorical cross-entropy loss using the Adam optimizer across 20 epochs with a batch size of 64:
model.fit(X, y, batch_size=64, epochs=20)
This configuration learns the probability distribution of character sequences, enabling the model to predict likely next characters given any input prefix.
Advanced RNN Techniques for Better Generation
Bidirectional Processing
Wrapping recurrent layers in keras.layers.Bidirectional() allows the network to process sequences from both directions simultaneously. By setting return_sequences=True when stacking multiple RNN layers, the model captures patterns from future and past context, though this approach is typically reserved for classification tasks rather than autoregressive generation.
Sequence Masking
Setting mask_zero=True on the embedding layer (or explicitly adding a Masking layer) tells Keras to skip padding tokens during training. This technique accelerates convergence by ensuring the model only learns from actual content rather than padded zeros, as demonstrated in the classification examples within RNNTF.ipynb.
Sampling Text from Trained Models
The final stage involves iteratively predicting characters and feeding outputs back as inputs. The GenerativeTF.ipynb notebook implements temperature scaling to control randomness during generation:
def sample(model, seed, length=200, temperature=1.0):
input_seq = np.array([[char2idx[c] for c in seed]])
result = seed
for _ in range(length):
preds = model.predict(input_seq, verbose=0)[0]
preds = np.log(preds + 1e-8) / temperature
probs = tf.nn.softmax(preds).numpy()
next_idx = np.random.choice(len(probs), p=probs)
next_char = idx2char[next_idx]
result += next_char
input_seq = np.append(input_seq[:,1:], [[next_idx]], axis=1)
return result
print(sample(model, seed='once upon a time ', length=300, temperature=0.8))
Temperature scaling divides the log-probabilities by the temperature parameter: values below 1.0 produce conservative, repetitive text, while values above 1.0 increase creativity and diversity in the generated output.
Summary
- The AI-For-Beginners repository implements text generation using RNNs in
lessons/5-NLP/16-RNN/andlessons/5-NLP/17-GenerativeNetworks/ - SimpleRNN provides basic sequential processing, while LSTM architectures handle long-range dependencies through cell and hidden states
- Character-level generation requires mapping text to integer indices, training an
Embedding → LSTM → Densestack on next-character prediction - Setting
mask_zero=Trueon embedding layers optimizes training by ignoring padding tokens - Temperature scaling during sampling controls the creativity of generated text by adjusting the probability distribution sharpness
- The curriculum provides both TensorFlow (
GenerativeTF.ipynb) and PyTorch (GenerativePyTorch.ipynb) implementations
Frequently Asked Questions
What is the difference between SimpleRNN and LSTM for text generation?
SimpleRNN maintains a single hidden state vector that updates at each time step, making it prone to vanishing gradients during backpropagation through long sequences. LSTM (Long Short-Term Memory) introduces a cell state and gating mechanisms (input, forget, and output gates) that preserve information across many time steps, enabling the model to learn longer dependencies essential for coherent text generation.
How does temperature scaling affect text generation quality?
Temperature scaling modifies the probability distribution of predicted characters by dividing log-probabilities by a temperature parameter before applying softmax. A temperature of 1.0 produces unbiased sampling from the learned distribution. Lower temperatures (e.g., 0.5) sharpen the distribution, favoring high-probability characters and producing conservative, repetitive but coherent text. Higher temperatures (e.g., 1.5) flatten the distribution, increasing randomness and creativity but potentially reducing grammatical coherence.
Why use character-level instead of word-level tokenization for RNN text generation?
Character-level tokenization treats each letter as a discrete token, resulting in a small vocabulary size (typically 50-100 unique characters) that reduces memory requirements and embedding matrix size. This approach allows the model to learn spelling patterns, punctuation, and morphological rules implicitly. While word-level models capture semantic meaning more directly, they require massive vocabularies and struggle with out-of-vocabulary words, whereas character models can generate any string composed of known characters.
Where can I find the PyTorch implementation of RNN text generation in the repository?
The PyTorch implementation resides in lessons/5-NLP/17-GenerativeNetworks/GenerativePyTorch.ipynb, which adapts the same LSTM character-level architecture to PyTorch's nn.LSTM module. This notebook demonstrates equivalent functionality to the TensorFlow version, including data preparation, model definition using torch.nn layers, and custom sampling loops with temperature control, providing a framework-agnostic understanding of RNN text generation principles.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →