How Linear Transformations Connect Vectors and Matrices in Mathematics
Linear transformations provide the mathematical bridge between abstract vector operations and concrete matrix algebra, where every linear map (T:\mathbb{R}^n\rightarrow\mathbb{R}^m) can be represented as matrix multiplication (A\mathbf{x}), enabling computation, composition, and geometric visualization.
Linear transformations form the foundational link between vectors and matrices in mathematical computing. In the HenryNdubuaku/maths-cs-ai-compendium repository, this connection is explored from basic geometric operations to deep learning architectures. Understanding how matrices encode linear maps allows you to compute complex transformations through simple matrix multiplication while preserving the geometric properties of the underlying space.
The Mathematical Definition of Linear Transformations
A linear transformation is a function (T:\mathbb{R}^n\rightarrow\mathbb{R}^m) that preserves the two fundamental operations of vector spaces: addition and scalar multiplication. According to the definitions in chapter 02 - matrices/04. linear transformations.md, any transformation must satisfy:
[
T(\mathbf{u}+\mathbf{v}) = T(\mathbf{u}) + T(\mathbf{v}),\qquad
T(c,\mathbf{u}) = c,T(\mathbf{u})
]
These properties ensure that the structure of the vector space is maintained—straight lines remain straight, the origin stays fixed, and parallel lines stay parallel. This preservation of linearity is what allows us to represent these abstract functions using concrete algebraic structures.
How Matrices Encode Linear Transformations
When a basis for the domain is chosen—typically the standard basis ({\hat{\mathbf{e}}_1, \hat{\mathbf{e}}_2, \ldots, \hat{\mathbf{e}}_n})—the action of (T) on each basis vector determines the complete transformation. The image of each basis vector becomes a column in the matrix (A), creating a direct correspondence:
- Column (i) of (A) equals (T(\hat{\mathbf{e}}_i)), the transformed basis vector
- Matrix-vector multiplication (A\mathbf{x}) computes the linear combination of these transformed basis vectors using the coordinates of (\mathbf{x})
This means the matrix is the transformation. As documented in the repository's linear transformations guide, applying (A) to any vector (\mathbf{x}) yields exactly the same result as applying (T) to (\mathbf{x}): (A\mathbf{x} = T(\mathbf{x})).
Composing Transformations Through Matrix Multiplication
One of the most powerful aspects of the matrix representation is how composition translates to multiplication. When transformations are applied sequentially:
- Apply (T_1) then (T_2): The mathematical composition is (T_2 \circ T_1)
- Matrix equivalent: The product (AB) (where (A) represents (T_2) and (B) represents (T_1))
This correspondence means that complex sequences of operations—rotations, scalings, and shears—can be collapsed into a single matrix through multiplication, enabling efficient computation and analysis of invariants like determinant and trace discussed in chapter 02 - matrices/01. matrix properties.md.
Geometric Operations and Affine Extensions
Linear transformations capture fundamental geometric operations through specific matrix patterns:
- Rotation: Orthogonal matrices with determinant 1
- Scaling: Diagonal matrices stretching coordinates
- Reflection: Orthogonal matrices with determinant -1
- Shearing: Triangular matrices sliding layers of space
While pure linear transformations must fix the origin, practical applications often require translation. As noted in the compendium, affine transformations extend linear maps by adding a translation vector (\mathbf{t}), represented using homogeneous coordinates in the matrix:
[ \begin{bmatrix}A & \mathbf{t}\0^{!\top}&1\end{bmatrix} ]
This unified representation preserves the convenience of matrix multiplication while handling the full range of geometric operations including shifting coordinate systems.
Linear Transformations in Machine Learning
The connection between vectors and matrices through linear transformations is central to modern AI. In chapter 06 - machine learning/03. deep learning.md, the repository demonstrates that each neural network layer implements a linear transformation:
- Weight matrix (W) maps input features to output dimensions
- Bias term adds affine translation
- Layer stacking corresponds to composing transformations: (W_3 \cdot \sigma(W_2 \cdot \sigma(W_1 \mathbf{x})))
Without activation functions breaking linearity, deep networks would collapse to a single matrix multiplication (W_{eff}\mathbf{x}). The chapter 12 - graph neural networks/04. graph attention networks.md extends this concept, showing how shared linear transformations (W) enable attention mechanisms to process graph-structured data through learned matrix operations.
Practical Implementation in Python
The following NumPy examples from the repository demonstrate how linear transformations connect vectors and matrices in computational practice.
Basic rotation and scaling composition:
import numpy as np
# 1️⃣ Simple linear transformation: rotation + scaling
theta = np.deg2rad(30) # 30° rotation
scale = 1.5 # uniform scaling
R = np.array([[np.cos(theta), -np.sin(theta)],
[np.sin(theta), np.cos(theta)]])
S = scale * np.eye(2) # scaling matrix
A = R @ S # combined linear transformation
v = np.array([1, 0]) # original vector
v_prime = A @ v # transformed vector: A @ v = T(v)
print("Original:", v)
print("Transformed:", v_prime)
Sequential composition of shear and rotation:
import matplotlib.pyplot as plt
# 2️⃣ Composition of two linear maps (shear then rotation)
shear = np.array([[1, 0.4],
[0, 1]]) # shear in x direction
rot = np.array([[0, -1],
[1, 0]]) # 90° rotation
combined = rot @ shear # matrix product = composition T_rot ∘ T_shear
square = np.array([[0,1,1,0,0],
[0,0,1,1,0]]) # homogeneous coordinates of a unit square
transformed = combined @ square
plt.figure()
plt.plot(square[0], square[1], 'r-o', label='original')
plt.plot(transformed[0], transformed[1], 'b-o', label='shear→rotate')
plt.axis('equal'); plt.grid(True); plt.legend()
plt.show()
Neural network linear layer implementation:
# 3️⃣ Neural‑network style linear layer (weights + bias)
X = np.random.randn(5, 3) # 5 samples, 3 features each (row vectors)
W = np.random.randn(4, 3) # weight matrix mapping ℝ³ → ℝ⁴
b = np.random.randn(4) # bias (affine part)
Y = X @ W.T + b # each row of X is transformed by the same linear map
print("Output shape:", Y.shape) # (5, 4) - batch of 5 transformed vectors
Summary
-
Linear transformations preserve vector addition and scalar multiplication, making them the "structure-preserving" maps of linear algebra.
-
Matrix columns represent the images of basis vectors, establishing a concrete algebraic realization of abstract linear maps through the equation (A\mathbf{x} = T(\mathbf{x})).
-
Matrix multiplication corresponds to transformation composition, allowing complex sequences of operations to be collapsed into a single computable form.
-
Affine extensions using homogeneous coordinates enable translations, unifying all geometric operations within the matrix framework.
-
Machine learning applications rely on stacked linear transformations (weight matrices) to map data between representation spaces, from basic fully-connected layers to graph attention mechanisms.
Frequently Asked Questions
What makes a transformation "linear"?
A transformation is linear when it satisfies the superposition principle: (T(\mathbf{u} + \mathbf{v}) = T(\mathbf{u}) + T(\mathbf{v})) and (T(c\mathbf{v}) = cT(\mathbf{v})). This means the transformation respects the linear structure of the vector space, keeping the origin fixed and mapping straight lines to straight lines. As detailed in chapter 02 - matrices/04. linear transformations.md, these properties are essential for the matrix representation to exist.
Why must the columns of the matrix be the transformed basis vectors?
The columns are the transformed basis vectors because any vector (\mathbf{x}) can be written as a linear combination of basis vectors (\mathbf{x} = \sum x_i \hat{\mathbf{e}}_i). By linearity, (T(\mathbf{x}) = \sum x_i T(\hat{\mathbf{e}}_i)). When we arrange the outputs (T(\hat{\mathbf{e}}_i)) as columns of (A), the matrix-vector product (A\mathbf{x}) computes exactly this linear combination, reproducing (T(\mathbf{x})).
How does matrix multiplication relate to composing linear transformations?
Matrix multiplication corresponds to sequential application of transformations. If matrix (B) represents transformation (T_1) and matrix (A) represents (T_2), then the product (AB) represents the composition (T_2 \circ T_1) (applying (T_1) first, then (T_2)). This is why the order matters: ((AB)\mathbf{x} = A(B\mathbf{x})) means apply (B) to (\mathbf{x}), then apply (A) to the result, exactly as function composition dictates.
Why are linear transformations important in neural networks?
Every neural network layer (before activation) applies a linear transformation to its input, multiplying by a weight matrix (W) and adding a bias. Stacking layers composes these transformations, allowing the network to learn complex mappings between high-dimensional spaces. According to chapter 06 - machine learning/03. deep learning.md and chapter 12 - graph neural networks/04. graph attention networks.md, even sophisticated architectures like Graph Attention Networks rely fundamentally on shared linear transformations (W) to project vectors into new coordinate systems where attention scores are computed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →