# How to Generate Text Using RNNs in AI for Beginners: A Complete Implementation Guide

> Learn to generate text with RNNs in AI for beginners. This guide covers SimpleRNN and LSTM for predicting text based on context. Start your AI journey now.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: how-to-guide
- Published: 2026-08-29

---

**Recurrent Neural Networks generate text by maintaining a hidden state across sequential tokens, using architectures like SimpleRNN and LSTM to predict the next character or word based on previous context.**

The Microsoft AI-For-Beginners curriculum provides hands-on implementations of text generation using RNNs in its Natural Language Processing section. Located in `lessons/5-NLP/`, these notebooks demonstrate complete pipelines from basic recurrent layers to advanced character-level language models using both TensorFlow and PyTorch.

## RNN Fundamentals in the Curriculum

The AI-For-Beginners repository organizes its RNN material across two core lessons. The `RNNTF.ipynb` notebook in `lessons/5-NLP/16-RNN/` introduces sequential processing concepts, while the generative notebooks in `lessons/5-NLP/17-GenerativeNetworks/` implement end-to-end text generation systems.

### SimpleRNN and Hidden States

A **SimpleRNN** layer processes text by updating a single hidden vector at each time step. In `lessons/5-NLP/16-RNN/RNNTF.ipynb`, the implementation combines a `TextVectorization` layer, an `Embedding` layer, and a `SimpleRNN` layer for basic classification:

```python
import tensorflow as tf
from tensorflow import keras

vocab_size = 20000
embed_dim = 64

vectorizer = keras.layers.TextVectorization(max_tokens=vocab_size)
model = keras.Sequential([
    vectorizer,
    keras.layers.Embedding(vocab_size, embed_dim),
    keras.layers.SimpleRNN(16),
    keras.layers.Dense(4, activation='softmax')
])
model.compile(loss='sparse_categorical_crossentropy',
              optimizer='adam',
              metrics=['accuracy'])

```

This architecture maintains a hidden state dimension of 16, passing the recurrent output through a final dense classifier.

### Long Short-Term Memory (LSTM)

**LSTM** layers mitigate the vanishing gradient problem by maintaining both a cell state (`c`) and a hidden state (`h`). The same notebook demonstrates replacing `SimpleRNN` with `keras.layers.LSTM()` for improved long-range dependency modeling, which is essential for coherent text generation across longer sequences.

## Character-Level Text Generation Pipeline

The complete workflow for generating text using RNNs in AI for Beginners follows four distinct stages: corpus preparation, model architecture design, next-token training, and probabilistic sampling.

### Step 1: Corpus Tokenization

The `GenerativeTF.ipynb` notebook implements character-level modeling by mapping each unique character to an integer index. This approach treats text as a sequence of discrete symbols rather than words, reducing vocabulary complexity:

```python
import string
import numpy as np

text = open('sample.txt').read().lower()
vocab = sorted(set(text))
char2idx = {c: i+1 for i, c in enumerate(vocab)}  # Reserve 0 for padding

idx2char = np.array([''] + list(vocab))

seq_len = 100
step = 1
sequences = [text[i:i+seq_len] for i in range(0, len(text)-seq_len, step)]
X = np.array([[char2idx[c] for c in seq] for seq in sequences])
y = np.array([char2idx[seq[-1]] for seq in sequences])

```

This creates input sequences of 100 characters (`seq_len`) and targets the subsequent character for prediction.

### Step 2: LSTM Architecture Design

The generative model stacks an `Embedding` layer with masking enabled, followed by an LSTM layer and a dense output layer with softmax activation:

```python
model = keras.Sequential([
    keras.layers.Embedding(len(vocab)+1, 128, mask_zero=True),
    keras.layers.LSTM(256),
    keras.layers.Dense(len(vocab)+1, activation='softmax')
])
model.compile(loss='sparse_categorical_crossentropy', optimizer='adam')

```

The `mask_zero=True` parameter instructs Keras to ignore padding tokens (index 0) during backpropagation, improving training efficiency when handling variable-length sequences.

### Step 3: Training Configuration

The model trains on categorical cross-entropy loss using the Adam optimizer across 20 epochs with a batch size of 64:

```python
model.fit(X, y, batch_size=64, epochs=20)

```

This configuration learns the probability distribution of character sequences, enabling the model to predict likely next characters given any input prefix.

## Advanced RNN Techniques for Better Generation

### Bidirectional Processing

Wrapping recurrent layers in `keras.layers.Bidirectional()` allows the network to process sequences from both directions simultaneously. By setting `return_sequences=True` when stacking multiple RNN layers, the model captures patterns from future and past context, though this approach is typically reserved for classification tasks rather than autoregressive generation.

### Sequence Masking

Setting `mask_zero=True` on the embedding layer (or explicitly adding a `Masking` layer) tells Keras to skip padding tokens during training. This technique accelerates convergence by ensuring the model only learns from actual content rather than padded zeros, as demonstrated in the classification examples within `RNNTF.ipynb`.

## Sampling Text from Trained Models

The final stage involves iteratively predicting characters and feeding outputs back as inputs. The `GenerativeTF.ipynb` notebook implements temperature scaling to control randomness during generation:

```python
def sample(model, seed, length=200, temperature=1.0):
    input_seq = np.array([[char2idx[c] for c in seed]])
    result = seed
    for _ in range(length):
        preds = model.predict(input_seq, verbose=0)[0]
        preds = np.log(preds + 1e-8) / temperature
        probs = tf.nn.softmax(preds).numpy()
        next_idx = np.random.choice(len(probs), p=probs)
        next_char = idx2char[next_idx]
        result += next_char
        input_seq = np.append(input_seq[:,1:], [[next_idx]], axis=1)
    return result

print(sample(model, seed='once upon a time ', length=300, temperature=0.8))

```

**Temperature scaling** divides the log-probabilities by the temperature parameter: values below 1.0 produce conservative, repetitive text, while values above 1.0 increase creativity and diversity in the generated output.

## Summary

- The AI-For-Beginners repository implements text generation using RNNs in `lessons/5-NLP/16-RNN/` and `lessons/5-NLP/17-GenerativeNetworks/`
- **SimpleRNN** provides basic sequential processing, while **LSTM** architectures handle long-range dependencies through cell and hidden states
- Character-level generation requires mapping text to integer indices, training an `Embedding → LSTM → Dense` stack on next-character prediction
- Setting `mask_zero=True` on embedding layers optimizes training by ignoring padding tokens
- Temperature scaling during sampling controls the creativity of generated text by adjusting the probability distribution sharpness
- The curriculum provides both TensorFlow (`GenerativeTF.ipynb`) and PyTorch (`GenerativePyTorch.ipynb`) implementations

## Frequently Asked Questions

### What is the difference between SimpleRNN and LSTM for text generation?

**SimpleRNN** maintains a single hidden state vector that updates at each time step, making it prone to vanishing gradients during backpropagation through long sequences. **LSTM** (Long Short-Term Memory) introduces a cell state and gating mechanisms (input, forget, and output gates) that preserve information across many time steps, enabling the model to learn longer dependencies essential for coherent text generation.

### How does temperature scaling affect text generation quality?

Temperature scaling modifies the probability distribution of predicted characters by dividing log-probabilities by a temperature parameter before applying softmax. A **temperature of 1.0** produces unbiased sampling from the learned distribution. **Lower temperatures** (e.g., 0.5) sharpen the distribution, favoring high-probability characters and producing conservative, repetitive but coherent text. **Higher temperatures** (e.g., 1.5) flatten the distribution, increasing randomness and creativity but potentially reducing grammatical coherence.

### Why use character-level instead of word-level tokenization for RNN text generation?

Character-level tokenization treats each letter as a discrete token, resulting in a small vocabulary size (typically 50-100 unique characters) that reduces memory requirements and embedding matrix size. This approach allows the model to learn spelling patterns, punctuation, and morphological rules implicitly. While word-level models capture semantic meaning more directly, they require massive vocabularies and struggle with out-of-vocabulary words, whereas character models can generate any string composed of known characters.

### Where can I find the PyTorch implementation of RNN text generation in the repository?

The PyTorch implementation resides in `lessons/5-NLP/17-GenerativeNetworks/GenerativePyTorch.ipynb`, which adapts the same LSTM character-level architecture to PyTorch's `nn.LSTM` module. This notebook demonstrates equivalent functionality to the TensorFlow version, including data preparation, model definition using `torch.nn` layers, and custom sampling loops with temperature control, providing a framework-agnostic understanding of RNN text generation principles.