Embedding Layers vs Linear Projections in PyTorch: Understanding the Difference in LLMs-from-scratch
Embedding layers perform integer-to-vector lookups from a learnable table, while linear projections execute dense matrix multiplications to transform continuous tensors between dimensionalities, with both serving distinct computational roles in transformer architectures.
The LLMs-from-scratch repository demonstrates fundamental transformer implementations where understanding the distinction between nn.Embedding and nn.Linear is essential for model architecture design. While both components learn parameters to map between vector spaces, they differ fundamentally in input types, memory access patterns, and gradient flow characteristics. This analysis examines the specific implementations in ch04.py and ch03.py to clarify when and why each layer type is employed.
Embedding Layers: The Token Lookup Mechanism
Embedding layers function as parameterized lookup tables that convert discrete token indices into dense vector representations. In the LLMs-from-scratch codebase, this mechanism initializes the model's input pipeline by mapping vocabulary IDs to initial hidden states.
Token and Positional Embeddings
The GPTModel class in ch04.py instantiates embeddings as the first processing step:
class GPTModel(nn.Module):
def __init__(self, cfg):
super().__init__()
# Token embedding lookup table: (vocab_size, emb_dim)
self.tok_emb = nn.Embedding(cfg["vocab_size"], cfg["emb_dim"])
# Positional embedding lookup table: (context_length, emb_dim)
self.pos_emb = nn.Embedding(cfg["context_length"], cfg["emb_dim"])
# Dropout for regularization
self.drop_emb = nn.Dropout(cfg["drop_rate"])
Source: [ch04.py lines 85-86](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch04.py#L85-L86)
Key characteristics of embedding layers:
- Input type: Accepts integer indices in the range
[0, vocab_size-1] - Parameter storage: Maintains a weight matrix of shape
(vocab_size, emb_dim)accessed via row indexing - Gradient behavior: Updates only the specific rows accessed during the forward pass (sparse gradients)
- Computational cost: O(1) lookup per token, no arithmetic operations beyond indexing
Linear Projections: Dense Transformations
Linear projections execute full matrix multiplications to transform continuous tensors from one dimensionality to another. Unlike embeddings, these layers mix information across all input dimensions and are used throughout the network for feature transformation.
Output Head Projection
The final layer mapping hidden states to vocabulary logits demonstrates a typical linear projection:
class GPTModel(nn.Module):
def __init__(self, cfg):
super().__init__()
# ... embeddings and transformer blocks ...
# Linear projection from emb_dim to vocab_size
self.out_head = nn.Linear(cfg["emb_dim"], cfg["vocab_size"], bias=False)
Source: [ch04.py lines 92-94](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch04.py#L92-L94)
Attention QKV Projections
Multi-head attention relies on linear layers to project input embeddings into query, key, and value representations. In ch03.py, the MultiHeadAttention class implements these as independent transformations:
class MultiHeadAttention(nn.Module):
def __init__(self, d_in, d_out, context_length,
dropout, num_heads, qkv_bias=False):
super().__init__()
# Projections for query, key, and value matrices
self.W_query = nn.Linear(d_in, d_out, bias=qkv_bias)
self.W_key = nn.Linear(d_in, d_out, bias=qkv_bias)
self.W_value = nn.Linear(d_in, d_out, bias=qkv_bias)
# Final output projection after attention
self.out_proj = nn.Linear(d_out, d_out)
self.dropout = nn.Dropout(dropout)
Source: [ch03.py lines 7-11](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch03.py#L7-L11)
Feed-Forward Expansions
The feed-forward network within transformer blocks uses linear projections to expand and contract dimensionality:
class FeedForward(nn.Module):
def __init__(self, cfg):
super().__init__()
self.layers = nn.Sequential(
# Expand: emb_dim -> 4*emb_dim
nn.Linear(cfg["emb_dim"], 4 * cfg["emb_dim"]),
GELU(),
# Contract: 4*emb_dim -> emb_dim
nn.Linear(4 * cfg["emb_dim"], cfg["emb_dim"]),
)
Source: [ch04.py lines 39-43](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch04.py#L39-L43)
Critical Differences: Lookup Tables vs Matrix Multiplication
The architectural distinction between these components determines their placement and function within the model:
| Aspect | nn.Embedding |
nn.Linear |
|---|---|---|
| Input domain | Discrete integers (token IDs) | Continuous real-valued tensors |
| Operation | Index-based row selection | Dense matrix multiplication (WX + b) |
| Information mixing | None (isolated row retrieval) | Full mixing across all input dimensions |
| Parameter count | O(vocab_size × emb_dim) | O(in_features × out_features) |
| Gradient flow | Sparse (only touched rows update) | Dense (entire weight matrix updates) |
| Typical position | Input layer only | Throughout network (attention, FFN, output) |
Memory layout implications: Embeddings store parameters as a lookup table where each token maintains its own vector representation. Linear layers treat weights as transformation matrices that operate on distributed representations across the entire input vector.
Architectural Integration in Transformers
The LLMs-from-scratch codebase demonstrates why both mechanisms are necessary in modern language models:
-
Embeddings provide the initial semantic grounding for discrete symbols. They are computationally efficient for the first layer because vocabulary sizes (e.g., 50,256 tokens) are manageable relative to model dimensions, and sparse gradients prevent overfitting to rare tokens.
-
Linear projections enable the model to dynamically transform representations through the network stack. Attention mechanisms require these dense transformations to compute compatibility scores between positions, while feed-forward networks use them to apply non-linear transformations at each layer.
-
Output unembedding: The linear projection in
self.out_headeffectively reverses the embedding process, mapping from the final hidden dimension back to vocabulary space for next-token prediction logits.
Summary
- Embedding layers in
ch04.pyimplementnn.Embeddingto convert integer token IDs into dense vectors via lookup tables, with parameters shaped(vocab_size, emb_dim). - Linear projections implemented as
nn.Linearinch03.pyandch04.pyperform dense matrix multiplications for attention QKV computation, feed-forward transformations, and output head mapping. - Gradient characteristics differ fundamentally: embeddings receive sparse updates (only accessed rows), while linear layers receive dense updates (full weight matrices).
- Computational patterns separate the two: embeddings use O(1) indexing whereas linear layers require O(n×m) matrix operations mixing all input dimensions.
- Architectural roles require both components: embeddings initialize the forward pass, while linear projections enable the deep feature transformations necessary for contextual understanding.
Frequently Asked Questions
Can you use a linear layer instead of an embedding layer for token inputs?
Technically yes, but it is inefficient and unconventional. You would need to one-hot encode token IDs (size vocab_size) and multiply by a weight matrix, resulting in O(vocab_size × emb_dim) operations per token versus O(1) lookup. The LLMs-from-scratch codebase uses nn.Embedding specifically in ch04.py to avoid this computational overhead.
Why does the output head use bias=False in the linear projection?
The output head self.out_head in ch04.py omits bias because the subsequent softmax operation in the loss calculation is translation-invariant. Adding a bias term would not change the probability distribution over the vocabulary, making the extra parameters redundant for next-token prediction tasks.
How do positional embeddings differ from token embeddings in implementation?
Both use nn.Embedding layers as shown in ch04.py, but they index different vocabularies. Token embeddings map from vocab_size (unique tokens), while positional embeddings map from context_length (sequence positions). Their outputs are summed element-wise to combine semantic and positional information before entering the transformer blocks.
Are the QKV projections in attention mechanisms the same as the feed-forward linear layers?
Both are instances of nn.Linear, but they serve different functional purposes. The projections in MultiHeadAttention (ch03.py) split the input into query, key, and value representations for attention computation, while feed-forward layers (ch04.py) expand and contract dimensionality to apply position-wise transformations after attention has mixed information across the sequence.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →