# Neural Network Architectures in the Maths‑CS‑AI Compendium: RNNs, CNNs, Transformers, and Hybrids

> Explore RNNs, CNNs, Transformers, and hybrid neural network architectures in the Maths-CS-AI Compendium. Understand key models for machine learning and AI.

- Repository: [Henry Ndubuaku/maths-cs-ai-compendium](https://github.com/HenryNdubuaku/maths-cs-ai-compendium)
- Tags: deep-dive
- Published: 2026-07-18

---

**The HenryNdubuaku/maths-cs-ai-compendium documents RNNs, CNNs, Transformers, and hybrid architectures across dedicated chapters on machine learning, computer vision, computational linguistics, and multimodal inference.**

The open‑source Maths‑CS‑AI Compendium is a comprehensive educational repository that surveys the most influential deep‑learning architectures. It traces how **recurrent neural networks**, **convolutional neural networks**, and **Transformers** process sequences, images, and multimodal data. Each explanation is anchored to specific source files and paired with minimal, framework‑agnostic code illustrations.

## Recurrent Neural Networks (RNNs)

In `chapter 06 - machine learning/03. deep learning.md`, the compendium introduces recurrent models as networks that reuse the same weights across time steps. A companion section in `chapter 07 - computational linguistics/03. embeddings and sequence models.md` expands the discussion to practical sequence‑modeling tasks.

### Vanilla RNNs, LSTMs, and GRUs

Vanilla RNNs update a hidden state element‑by‑element, but they suffer from vanishing gradients on long sequences. **Long Short‑Term Memory (LSTM)** and **Gated Recurrent Unit (GRU)** cells solve this problem by adding gating mechanisms that regulate information flow. These variants are treated as the default choice for temporal modeling in the compendium.

### Bidirectional and Stacked Variants

**Bidirectional RNNs** process input in both forward and backward directions, exposing past and future context simultaneously. Stacking multiple recurrent layers adds depth and enables the model to learn hierarchical temporal representations. The compendium recommends these layouts for tagging and speech‑recognition pipelines.

The repository provides a minimal NumPy illustration of a single recurrent step:

```python
import numpy as np

def rnn_step(x_t, h_prev, W_xh, W_hh, b):
    """One recurrent step."""
    return np.tanh(np.dot(W_xh, x_t) + np.dot(W_hh, h_prev) + b)

```

## Convolutional Neural Networks (CNNs)

The central treatment of CNNs lives in `chapter 08 - computer vision/02. convolutional networks.md`, with cross‑references in `chapter 06 - machine learning/03. deep learning.md`. For video applications, `chapter 08 - computer vision/05. video and 3D vision.md` extends the same concepts to three dimensions.

### 1‑D, 2‑D, and 3‑D Convolutions

A convolutional layer enforces **weight sharing** and **local connectivity**, which yields translation equivariance and drastically reduces parameter counts. **1‑D CNNs** handle time‑series and raw audio, **2‑D CNNs** dominate image classification, and **3‑D CNNs** capture spatio‑temporal patterns in video volumes.

### Hierarchical Feature Learning

Stacking convolutional layers grows the **receptive field** exponentially, letting the network learn edges in early layers and complex objects in deeper ones. The compendium walks through historic architectures and modern tricks such as skip connections and dilated convolutions.

A valid 2‑D convolution is shown below:

```python
import numpy as np
from scipy.signal import correlate2d

def conv2d(input_img, kernel, bias=0):
    """Valid 2‑D convolution (single channel)."""
    return correlate2d(input_img, kernel, mode='valid') + bias

```

## Transformers and Self‑Attention

Transformer theory is introduced in `chapter 06 - machine learning/03. deep learning.md` and given language‑model focus in `chapter 07 - computational linguistics/04. transformers and language models.md`. Vision‑specific adaptations appear in `chapter 08 - computer vision/04. vision transformers and generation.md`.

### Encoder‑Decoder Architecture

**Self‑attention** replaces recurrence by computing pairwise interaction scores between every token in a sequence. This allows full parallelization during training and scales gracefully to long contexts. **Positional encodings** inject order information, because the attention mechanism itself is permutation invariant.

### Vision Transformers (ViT) and Swin

The compendium explains how **Vision Transformers (ViT)** split an image into fixed‑size patches, linearly embed each patch, and process the resulting sequence with standard Transformer blocks. **Swin Transformers** improve efficiency by introducing hierarchical feature maps and shifted windows, making attention practical for high‑resolution vision tasks.

A minimal NumPy encoder layer illustrating scaled dot‑product attention is:

```python
import numpy as np

def scaled_dot_product_attention(Q, K, V):
    """Simple attention without masking."""
    d_k = Q.shape[-1]
    scores = np.dot(Q, K.T) / np.sqrt(d_k)
    weights = np.exp(scores) / np.exp(scores).sum(axis=-1, keepdims=True)
    return np.dot(weights, V)

def transformer_encoder(x, W_q, W_k, W_v, W_o):
    """One encoder layer (no feed‑forward for brevity)."""
    Q = np.dot(x, W_q)
    K = np.dot(x, W_k)
    V = np.dot(x, W_v)
    attn = scaled_dot_product_attention(Q, K, V)
    return np.dot(attn, W_o)          # + residual & layer‑norm omitted

```

## Hybrid and Emerging Architectures

Beyond the three core families, the compendium studies models that blend their strengths. These hybrids are designed for efficiency, novel modalities, or structured data.

### RWKV: Bridging RNNs and Transformers

`chapter 17 - AI inference/02. efficient architectures.md` introduces **RWKV**, a linear‑attention architecture that trains like a Transformer but infers like an RNN. It preserves the parallelizability of self‑attention during training while eliminating the quadratic memory costs of standard Transformers at inference time.

### Temporal Graph Neural Networks

In `chapter 12 - graph neural networks/04. graph attention networks.md`, the repository covers temporal GNNs that marry graph convolutions with RNN or attention modules. This combination lets the model reason over dynamic relationships in social, traffic, or molecular graphs.

### Multimodal Encoders

`chapter 10 - multimodal learning/05. unified multimodal architectures.md` describes pipelines that fuse **CNN** or **ViT** vision encoders with language Transformers. These unified architectures power vision‑language pre‑training tasks such as image captioning and visual question answering.

## Summary

- The Maths‑CS‑AI Compendium organizes **RNNs**, **CNNs**, and **Transformers** into dedicated chapters with direct links to source files.
- Key files include `chapter 06 - machine learning/03. deep learning.md`, `chapter 08 - computer vision/02. convolutional networks.md`, and `chapter 07 - computational linguistics/04. transformers and language models.md`.
- Code illustrations use NumPy and SciPy to expose the inner mechanics of each architecture without framework overhead.
- Hybrid designs such as **RWKV**, temporal GNNs, and multimodal CNN‑ViT encoders demonstrate how classic blocks are combined for modern efficiency and cross‑modal tasks.

## Frequently Asked Questions

### What neural network architectures are covered in the compendium?

The compendium covers **RNNs**, **CNNs**, **Transformers**, and hybrid architectures including **Vision Transformers**, **Swin**, **RWKV**, and temporal graph neural networks. Each family is mapped to specific source files that explain theory, implementation, and domain‑specific adaptations.

### Which file explains Vision Transformers in detail?

Vision Transformers are detailed in `chapter 08 - computer vision/04. vision transformers and generation.md`. This file covers patch‑based tokenization, positional embeddings, and hierarchical variants such as Swin Transformer.

### Are the explanations limited to theory, or is there runnable code?

Every architecture chapter includes both conceptual discussion and minimal, runnable code. The snippets are written in NumPy and SciPy to remain framework agnostic, while full exercises in JAX appear inside the individual chapter examples.

### How do hybrid architectures differ from pure RNNs, CNNs, or Transformers?

Hybrid architectures combine building blocks from the core families to solve specific constraints. For example, **RWKV** merges Transformer‑style training with RNN‑style inference, and temporal GNNs blend graph attention with recurrent modules to handle dynamic relational data.