# Embedding Layers vs Linear Projections in PyTorch: Understanding the Difference in LLMs-from-scratch

> Understand the difference between embedding layers and linear projections in PyTorch for LLMs. Learn how these distinct components transform data within transformer architectures.

- Repository: [Sebastian Raschka/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch)
- Tags: deep-dive
- Published: 2026-05-12

---

**Embedding layers perform integer-to-vector lookups from a learnable table, while linear projections execute dense matrix multiplications to transform continuous tensors between dimensionalities, with both serving distinct computational roles in transformer architectures.**

The *LLMs-from-scratch* repository demonstrates fundamental transformer implementations where understanding the distinction between `nn.Embedding` and `nn.Linear` is essential for model architecture design. While both components learn parameters to map between vector spaces, they differ fundamentally in input types, memory access patterns, and gradient flow characteristics. This analysis examines the specific implementations in [`ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04.py) and [`ch03.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch03.py) to clarify when and why each layer type is employed.

## Embedding Layers: The Token Lookup Mechanism

Embedding layers function as parameterized lookup tables that convert discrete token indices into dense vector representations. In the *LLMs-from-scratch* codebase, this mechanism initializes the model's input pipeline by mapping vocabulary IDs to initial hidden states.

### Token and Positional Embeddings

The `GPTModel` class in [`ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04.py) instantiates embeddings as the first processing step:

```python
class GPTModel(nn.Module):
    def __init__(self, cfg):
        super().__init__()
        # Token embedding lookup table: (vocab_size, emb_dim)

        self.tok_emb = nn.Embedding(cfg["vocab_size"], cfg["emb_dim"])
        # Positional embedding lookup table: (context_length, emb_dim)

        self.pos_emb = nn.Embedding(cfg["context_length"], cfg["emb_dim"])
        # Dropout for regularization

        self.drop_emb = nn.Dropout(cfg["drop_rate"])

```

*Source:* [[`ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04.py) lines 85-86](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch04.py#L85-L86)

**Key characteristics of embedding layers:**
- **Input type**: Accepts integer indices in the range `[0, vocab_size-1]`
- **Parameter storage**: Maintains a weight matrix of shape `(vocab_size, emb_dim)` accessed via row indexing
- **Gradient behavior**: Updates only the specific rows accessed during the forward pass (sparse gradients)
- **Computational cost**: O(1) lookup per token, no arithmetic operations beyond indexing

## Linear Projections: Dense Transformations

Linear projections execute full matrix multiplications to transform continuous tensors from one dimensionality to another. Unlike embeddings, these layers mix information across all input dimensions and are used throughout the network for feature transformation.

### Output Head Projection

The final layer mapping hidden states to vocabulary logits demonstrates a typical linear projection:

```python
class GPTModel(nn.Module):
    def __init__(self, cfg):
        super().__init__()
        # ... embeddings and transformer blocks ...

        
        # Linear projection from emb_dim to vocab_size

        self.out_head = nn.Linear(cfg["emb_dim"], cfg["vocab_size"], bias=False)

```

*Source:* [[`ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04.py) lines 92-94](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch04.py#L92-L94)

### Attention QKV Projections

Multi-head attention relies on linear layers to project input embeddings into query, key, and value representations. In [`ch03.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch03.py), the `MultiHeadAttention` class implements these as independent transformations:

```python
class MultiHeadAttention(nn.Module):
    def __init__(self, d_in, d_out, context_length, 
                 dropout, num_heads, qkv_bias=False):
        super().__init__()
        # Projections for query, key, and value matrices

        self.W_query = nn.Linear(d_in, d_out, bias=qkv_bias)
        self.W_key = nn.Linear(d_in, d_out, bias=qkv_bias)
        self.W_value = nn.Linear(d_in, d_out, bias=qkv_bias)
        
        # Final output projection after attention

        self.out_proj = nn.Linear(d_out, d_out)
        self.dropout = nn.Dropout(dropout)

```

*Source:* [[`ch03.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch03.py) lines 7-11](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch03.py#L7-L11)

### Feed-Forward Expansions

The feed-forward network within transformer blocks uses linear projections to expand and contract dimensionality:

```python
class FeedForward(nn.Module):
    def __init__(self, cfg):
        super().__init__()
        self.layers = nn.Sequential(
            # Expand: emb_dim -> 4*emb_dim

            nn.Linear(cfg["emb_dim"], 4 * cfg["emb_dim"]),
            GELU(),
            # Contract: 4*emb_dim -> emb_dim

            nn.Linear(4 * cfg["emb_dim"], cfg["emb_dim"]),
        )

```

*Source:* [[`ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04.py) lines 39-43](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch04.py#L39-L43)

## Critical Differences: Lookup Tables vs Matrix Multiplication

The architectural distinction between these components determines their placement and function within the model:

| Aspect | `nn.Embedding` | `nn.Linear` |
|--------|---------------|-------------|
| **Input domain** | Discrete integers (token IDs) | Continuous real-valued tensors |
| **Operation** | Index-based row selection | Dense matrix multiplication (WX + b) |
| **Information mixing** | None (isolated row retrieval) | Full mixing across all input dimensions |
| **Parameter count** | O(vocab_size × emb_dim) | O(in_features × out_features) |
| **Gradient flow** | Sparse (only touched rows update) | Dense (entire weight matrix updates) |
| **Typical position** | Input layer only | Throughout network (attention, FFN, output) |

**Memory layout implications**: Embeddings store parameters as a lookup table where each token maintains its own vector representation. Linear layers treat weights as transformation matrices that operate on distributed representations across the entire input vector.

## Architectural Integration in Transformers

The *LLMs-from-scratch* codebase demonstrates why both mechanisms are necessary in modern language models:

1. **Embeddings** provide the initial semantic grounding for discrete symbols. They are computationally efficient for the first layer because vocabulary sizes (e.g., 50,256 tokens) are manageable relative to model dimensions, and sparse gradients prevent overfitting to rare tokens.

2. **Linear projections** enable the model to dynamically transform representations through the network stack. Attention mechanisms require these dense transformations to compute compatibility scores between positions, while feed-forward networks use them to apply non-linear transformations at each layer.

3. **Output unembedding**: The linear projection in `self.out_head` effectively reverses the embedding process, mapping from the final hidden dimension back to vocabulary space for next-token prediction logits.

## Summary

- **Embedding layers** in [`ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04.py) implement `nn.Embedding` to convert integer token IDs into dense vectors via lookup tables, with parameters shaped `(vocab_size, emb_dim)`.
- **Linear projections** implemented as `nn.Linear` in [`ch03.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch03.py) and [`ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04.py) perform dense matrix multiplications for attention QKV computation, feed-forward transformations, and output head mapping.
- **Gradient characteristics** differ fundamentally: embeddings receive sparse updates (only accessed rows), while linear layers receive dense updates (full weight matrices).
- **Computational patterns** separate the two: embeddings use O(1) indexing whereas linear layers require O(n×m) matrix operations mixing all input dimensions.
- **Architectural roles** require both components: embeddings initialize the forward pass, while linear projections enable the deep feature transformations necessary for contextual understanding.

## Frequently Asked Questions

### Can you use a linear layer instead of an embedding layer for token inputs?

Technically yes, but it is inefficient and unconventional. You would need to one-hot encode token IDs (size `vocab_size`) and multiply by a weight matrix, resulting in O(vocab_size × emb_dim) operations per token versus O(1) lookup. The *LLMs-from-scratch* codebase uses `nn.Embedding` specifically in [`ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04.py) to avoid this computational overhead.

### Why does the output head use `bias=False` in the linear projection?

The output head `self.out_head` in [`ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04.py) omits bias because the subsequent softmax operation in the loss calculation is translation-invariant. Adding a bias term would not change the probability distribution over the vocabulary, making the extra parameters redundant for next-token prediction tasks.

### How do positional embeddings differ from token embeddings in implementation?

Both use `nn.Embedding` layers as shown in [`ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04.py), but they index different vocabularies. Token embeddings map from `vocab_size` (unique tokens), while positional embeddings map from `context_length` (sequence positions). Their outputs are summed element-wise to combine semantic and positional information before entering the transformer blocks.

### Are the QKV projections in attention mechanisms the same as the feed-forward linear layers?

Both are instances of `nn.Linear`, but they serve different functional purposes. The projections in `MultiHeadAttention` ([`ch03.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch03.py)) split the input into query, key, and value representations for attention computation, while feed-forward layers ([`ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04.py)) expand and contract dimensionality to apply position-wise transformations after attention has mixed information across the sequence.