# How Neural Network Activation Functions Are Implemented in TheAlgorithms/Python: Mathematical Underpinnings Explained

> Discover how neural network activation functions are implemented in TheAlgorithms/Python using NumPy. Explore the mathematical underpinnings and vectorized operations explained directly from code.

- Repository: [The Algorithms/Python](https://github.com/TheAlgorithms/Python)
- Tags: deep-dive
- Published: 2026-02-24

---

**Activation functions in TheAlgorithms/Python are implemented as pure NumPy utilities under `neural_network/activation_functions/`, directly translating mathematical formulas into vectorized operations that work on any array-like input.**

The **TheAlgorithms/Python** repository provides educational implementations of classic and modern **neural network activation functions** that bridge theoretical mathematics with production-ready code. Each function resides in its own module within the `neural_network/activation_functions/` directory, offering dependency-free, O(N) performance through NumPy's element-wise broadcasting.

## Architectural Design and Code Structure

The repository follows a strict modular pattern where each activation function is isolated in a dedicated file. This design prioritizes readability, testability, and mathematical transparency.

All implementations share these characteristics:

- **Input flexibility**: Accept both Python `list[float]` and `np.ndarray` of any shape
- **Vectorized operations**: Use NumPy primitives like `np.maximum`, `np.where`, and `np.exp` to eliminate Python loops
- **Consistent returns**: Always output `np.ndarray` preserving input shape
- **Self-documenting code**: Each file includes the mathematical formula in the docstring alongside a `doctest` verification block

For example, importing specific activations follows a clean namespace pattern:

```python
from neural_network.activation_functions.rectified_linear_unit import relu
from neural_network.activation_functions.swish import sigmoid_linear_unit

```

## Mathematical Formulas and Implementation Details

Each source file maps a specific mathematical definition to its NumPy equivalent, ensuring the code accurately reflects the underlying calculus used in forward propagation.

### Rectified Linear Unit (ReLU) and Variants

**ReLU** ([`rectified_linear_unit.py`](https://github.com/TheAlgorithms/Python/blob/main/rectified_linear_unit.py)) implements the piecewise linear function \( f(x) = \max(0, x) \) using `np.maximum(0, vector)` to threshold negative values at zero.

**Leaky ReLU** ([`leaky_rectified_linear_unit.py`](https://github.com/TheAlgorithms/Python/blob/main/leaky_rectified_linear_unit.py)) extends this with a configurable negative slope \(\alpha\):

\[ f(x) = \begin{cases} x & x > 0 \\ \alpha x & x \le 0 \end{cases} \]

The implementation uses `np.where(vector > 0, vector, alpha * vector)` to apply the linear leak \( \alpha x \) to all non-positive inputs.

**Softplus** ([`softplus.py`](https://github.com/TheAlgorithms/Python/blob/main/softplus.py)) provides a smooth differentiable approximation to ReLU via \( f(x) = \ln(1 + e^{x}) \), implemented as `np.log(1 + np.exp(vector))` to avoid vanishing gradient problems during backpropagation.

### Swish and Sigmoid Linear Unit (SiLU)

The **Swish** activation ([`swish.py`](https://github.com/TheAlgorithms/Python/blob/main/swish.py)) introduces a self-gated mechanism defined as \( f(x) = x \cdot \sigma(x) \), where \( \sigma(x) = \frac{1}{1+e^{-x}} \) represents the sigmoid function. The code computes this via `1/(1+np.exp(-vector))` followed by element-wise multiplication.

A generalized variant includes a trainable scalar \( \beta \):

\[ f(x)=x\cdot\sigma(\beta x) \]

**SiLU** (Sigmoid Linear Unit) is a specific instance of Swish where \( \beta = 1 \), available through the `sigmoid_linear_unit` helper function in the same module.

### Binary Step and Specialized Functions

**Binary Step** ([`binary_step.py`](https://github.com/TheAlgorithms/Python/blob/main/binary_step.py)) implements the Heaviside step function:

\[ f(x) = \begin{cases} 1 & x \ge 0 \\ 0 & x < 0 \end{cases} \]

Using `np.where(vector >= 0, 1, 0)`, this returns a binary mask suitable for binary classification outputs or thresholding operations.

Additional functions like **Mish** and **GELU** follow identical patterns in their respective modules, combining `np.tanh`, `np.log1p`, and error functions to implement more complex non-linearities like \( f(x)=x\tanh(\ln(1+e^{x})) \).

## Practical Usage in Neural Network Layers

These activation functions integrate seamlessly into custom layer implementations. Below is a complete example demonstrating imports and a simple feed-forward computation:

```python
import numpy as np
from neural_network.activation_functions.rectified_linear_unit import relu
from neural_network.activation_functions.leaky_rectified_linear_unit import leaky_rectified_linear_unit
from neural_network.activation_functions.softplus import softplus
from neural_network.activation_functions.swish import swish, sigmoid_linear_unit

# Sample input vector

x = np.array([-2.0, -0.5, 0.0, 1.5, 3.0])

# Standard ReLU

print("ReLU:", relu(x))

# Output: [0.  0.  0.  1.5 3. ]

# Leaky ReLU with alpha=0.01

print("Leaky ReLU:", leaky_rectified_linear_unit(x, alpha=0.01))

# Output: [-0.02 -0.005  0.    1.5   3.   ]

# Softplus smooth activation

print("Softplus:", softplus(x))

# Output: [0.12692801 0.47407698 0.69314718 2.01490302 3.04858735]

# Swish with trainable parameter beta=1

print("Swish:", swish(x, trainable_parameter=1))

# Output: [-0.23840584 -0.23687862  0.5         1.30554893  3.0       ]

```

You can use these activations within custom dense layers:

```python
def dense_layer(inputs, weights, bias, activation):
    """Compute a single dense layer with activation."""
    z = inputs @ weights + bias
    return activation(z)

# Example usage with ReLU

W = np.random.randn(5, 3)
b = np.random.randn(3)
output = dense_layer(x, W, b, relu)

```

## Summary

- **TheAlgorithms/Python** implements **neural network activation functions** as standalone NumPy modules in `neural_network/activation_functions/`
- Each source file (e.g., [`rectified_linear_unit.py`](https://github.com/TheAlgorithms/Python/blob/main/rectified_linear_unit.py), [`swish.py`](https://github.com/TheAlgorithms/Python/blob/main/swish.py)) contains the mathematical formula, docstring explanation, and vectorized implementation
- Functions rely on `np.maximum`, `np.where`, and `np.exp` for O(N) performance without Python loops
- The modular architecture supports educational use, unit testing via doctests, and easy integration into larger deep learning frameworks
- Both classic functions (ReLU, Binary Step) and modern variants (Swish, Softplus) follow identical implementation patterns for consistency

## Frequently Asked Questions

### What mathematical formula does the ReLU implementation use in TheAlgorithms/Python?

The ReLU implementation in [`rectified_linear_unit.py`](https://github.com/TheAlgorithms/Python/blob/main/rectified_linear_unit.py) uses the formula \( f(x) = \max(0, x) \). It translates this directly to NumPy via `np.maximum(0, vector)`, which performs an element-wise comparison against zero and returns the input value if positive, otherwise zero.

### How does the Leaky ReLU function handle negative values?

According to [`leaky_rectified_linear_unit.py`](https://github.com/TheAlgorithms/Python/blob/main/leaky_rectified_linear_unit.py), Leaky ReLU handles negative values by multiplying them by a small constant \( \alpha \) (default 0.01). The implementation uses `np.where(vector > 0, vector, alpha * vector)` to apply a linear leak rather than zeroing out negative activations, preventing dead neurons during training.

### Are these activation functions suitable for production deep learning frameworks?

While the implementations are mathematically correct and vectorized for O(N) performance, they are designed primarily for educational purposes and lightweight prototyping. They use pure NumPy without GPU acceleration or automatic differentiation, making them ideal for learning the mathematical underpinnings of **neural network activation functions** but requiring extension via frameworks like PyTorch or TensorFlow for large-scale production use.

### What is the difference between Swish and SiLU in the repository?

As implemented in [`swish.py`](https://github.com/TheAlgorithms/Python/blob/main/swish.py), **SiLU** (Sigmoid Linear Unit) is a specific case of the **Swish** function where the trainable parameter \( \beta = 1 \). SiLU uses the formula \( f(x) = x \cdot \sigma(x) \), while the generic Swish function accepts a `trainable_parameter` argument to compute \( f(x)=x\cdot\sigma(\beta x) \), allowing learnable activation shaping during neural network training.