# Common Loss Functions in Neural Networks: A Complete Guide with PyTorch Implementations

> Master common loss functions in neural networks with PyTorch. Understand MSE cross-entropy focal loss and InfoNCE to improve your AI models.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: deep-dive
- Published: 2026-07-19

---

**Neural networks learn by minimizing loss functions that quantify the discrepancy between predictions and targets, with Mean-Squared Error governing regression tasks and cross-entropy variants handling classification, while specialized losses like Focal and InfoNCE address imbalanced data and contrastive learning.**

Selecting the appropriate loss function determines whether a model learns meaningful patterns or converges to suboptimal solutions. The `rohitg00/ai-engineering-from-scratch` curriculum provides a comprehensive examination of **common loss functions in neural networks**, from foundational regression criteria to advanced contrastive objectives. This article breaks down each loss type, explains when to apply them, and provides runnable PyTorch code based on the repository's reference implementations.

## Core Loss Functions: Regression vs. Classification

Neural network architectures require different optimization objectives depending on whether they predict continuous values or discrete categories. The repository's lesson documentation in [`phases/03-deep-learning-core/05-loss-functions/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/03-deep-learning-core/05-loss-functions/docs/en.md) establishes three fundamental building blocks.

### Mean-Squared Error (MSE) for Regression

Use **MSE** when your network outputs continuous values. The formula \(\frac{1}{N}\sum (y_i-\hat y_i)^2\) calculates the average squared difference between predictions \(\hat y\) and targets \(y\).

This loss is smooth and convex, penalizing large errors quadratically. However, as demonstrated in the curriculum, MSE can be "gamed" in binary tasks—a model predicting 0.5 for every sample achieves an MSE of 0.25 without learning discriminative features.

### Binary Cross-Entropy (BCE) for Two-Class Problems

For binary classification, **Binary Cross-Entropy** applies the formula \(-[y\log p+(1-y)\log(1-p)]\), where \(p\) represents the predicted probability for the positive class.

The gradients simplify to \(p - y\), providing strong signals when predictions are confident and wrong. In PyTorch, always use `F.binary_cross_entropy_with_logits` rather than manually applying a sigmoid followed by BCE, as this preserves numerical stability.

### Categorical Cross-Entropy for Multi-Class Classification

When exactly one class is correct among many, **Categorical Cross-Entropy** (implemented via softmax cross-entropy) uses \(-\log \frac{e^{z_{y}}}{\sum_j e^{z_j}}\), where \(z_y\) is the logit for the true class.

According to the source code analysis, you should call `torch.nn.functional.cross_entropy` directly rather than chaining `softmax` and `NLLLoss`, as the combined operation avoids numerical instability by computing the log-softmax in a single stable pass.

## Advanced Loss Functions for Specialized Tasks

Beyond basic classification and regression, modern deep learning requires losses that handle noisy labels, class imbalance, and representation learning.

### Label-Smoothed Cross-Entropy for Calibration

**Label smoothing** prevents models from becoming overconfident by blending hard targets (0 and 1) with a uniform distribution. Instead of targeting 1.0 for the correct class, the model optimizes toward \(1 - \epsilon\) (e.g., 0.9) and distributes the remaining \(\epsilon\) across other classes.

In PyTorch, enable this by setting `label_smoothing=0.1` in `F.cross_entropy`. This regularization technique proves particularly effective when dealing with noisy labels or when model calibration is critical.

### Focal Loss for Highly Imbalanced Datasets

**Focal Loss** addresses extreme class imbalance (common in object detection) by down-weighting easy examples. The formula \(- (1-p_t)^\gamma \log(p_t)\) introduces a modulating factor where \(p_t\) is the probability for the true class and \(\gamma\) controls the focusing strength.

When \(\gamma = 2\), the loss reduces the gradient contribution from well-classified examples by \((1-p_t)^2\), forcing the optimizer to concentrate on hard negatives and minority class instances.

### InfoNCE for Contrastive and Self-Supervised Learning

**InfoNCE** converts similarity learning into a classification problem over batch elements. The formula \(-\log \frac{e^{s_{ii}/\tau}}{\sum_j e^{s_{ij}/\tau}}\) compares the similarity \(s_{ii}\) of a positive pair against all negative pairs \(s_{ij}\), scaled by temperature \(\tau\).

This loss powers modern vision-language models and self-supervised learning frameworks by pulling matching embeddings together while pushing non-matching pairs apart. The repository notes that InfoNCE effectively uses the same optimization machinery as standard cross-entropy, making it both scalable and simple to implement.

### KL-Divergence for Knowledge Distillation

**KL-Divergence** measures how one probability distribution diverges from another, computed as \(\sum_y p(y) \log \frac{p(y)}{q(y)}\) where \(p\) is the target distribution and \(q\) is the model's output.

This loss becomes equivalent to cross-entropy when the target is one-hot, but proves essential for knowledge distillation where the "teacher" model provides soft probability distributions rather than hard labels.

## Why Loss Function Selection Determines Model Performance

The curriculum emphasizes that loss choice directly impacts learned representations. In [`phases/03-deep-learning-core/05-loss-functions/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/03-deep-learning-core/05-loss-functions/docs/en.md), the authors demonstrate that **MSE** fails for binary tasks because it plateaus when predictions approach 0.5, whereas **cross-entropy** applies an infinite penalty for confident wrong predictions through the \(-\log(p)\) term.

This mathematical property forces the network to push probabilities toward the extremes (0 or 1), learning discriminative features rather than conservative averages. Similarly, **Focal Loss** solves the dominance of background examples in detection tasks, while **InfoNCE** enables scalable contrastive learning by treating each batch as a classification problem over instance indices.

## Practical PyTorch Implementation Examples

The reference implementation in [`phases/03-deep-learning-core/05-loss-functions/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/03-deep-learning-core/05-loss-functions/code/main.py) demonstrates each loss using PyTorch's functional API. Below is a consolidated example covering all major loss types:

```python
import torch
import torch.nn.functional as F

# Dummy data: 4 samples, 3 classes

logits = torch.randn(4, 3, requires_grad=True)
labels = torch.tensor([0, 2, 1, 2])

# 1. Mean-Squared Error (regression)

target_reg = torch.randn(4, 3)
mse = F.mse_loss(logits, target_reg)
mse.backward()
print('MSE:', mse.item())

# 2. Binary Cross-Entropy

bce = F.binary_cross_entropy_with_logits(
    logits[:, 0], 
    torch.tensor([1., 0., 1., 0.])
)
bce.backward()
print('BCE:', bce.item())

# 3. Categorical Cross-Entropy

ce = F.cross_entropy(logits, labels)
ce.backward()
print('Cross-Entropy:', ce.item())

# 4. Label-Smoothed Cross-Entropy

ce_smooth = F.cross_entropy(logits, labels, label_smoothing=0.1)
ce_smooth.backward()
print('Label-smoothed CE:', ce_smooth.item())

# 5. Focal Loss (γ=2)

prob = F.softmax(logits, dim=1)
pt = prob.gather(1, labels.unsqueeze(1)).squeeze()
focal = -(1 - pt) ** 2 * torch.log(pt + 1e-12)
focal = focal.mean()
focal.backward()
print('Focal loss:', focal.item())

# 6. InfoNCE (contrastive) - batch of 4 embeddings

emb = F.normalize(torch.randn(4, 64), dim=1)
sim = emb @ emb.t() / 0.07  # temperature τ=0.07

labels_contrast = torch.arange(4)
info_nce = F.cross_entropy(sim, labels_contrast)
info_nce.backward()
print('InfoNCE:', info_nce.item())

```

Always use the high-level functional APIs (`F.mse_loss`, `F.cross_entropy`, etc.) because they fuse operations for numerical stability. For example, `F.cross_entropy` combines `log_softmax` and `NLLLoss` in a single kernel, preventing the overflow issues common when computing softmax separately.

## Key Files in the ai-engineering-from-scratch Repository

| Resource | Path | Purpose |
|----------|------|---------|
| Lesson Documentation | [`phases/03-deep-learning-core/05-loss-functions/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/03-deep-learning-core/05-loss-functions/docs/en.md) | Conceptual exposition, mathematical derivations, and pedagogical notes on each loss function. |
| Reference Implementation | [`phases/03-deep-learning-core/05-loss-functions/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/03-deep-learning-core/05-loss-functions/code/main.py) | From-scratch implementations of MSE, BCE, categorical CE, label smoothing, focal loss, and InfoNCE with gradient calculations. |
| Loss Selector Prompt | [`phases/03-deep-learning-core/05-loss-functions/outputs/prompt-loss-function-selector.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/03-deep-learning-core/05-loss-functions/outputs/prompt-loss-function-selector.md) | Interactive decision tree for selecting appropriate losses based on task characteristics. |
| Loss Debugger Prompt | [`phases/03-deep-learning-core/05-loss-functions/outputs/prompt-loss-debugger.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/03-deep-learning-core/05-loss-functions/outputs/prompt-loss-debugger.md) | Common pitfalls and debugging strategies, such as identifying duplicate softmax applications. |

These resources follow the curriculum's philosophy of building from raw mathematical principles before using high-level framework APIs.

## Summary

- Use **Mean-Squared Error** for regression tasks with continuous targets, but avoid it for classification where it permits uninformative average predictions.
- Apply **Binary Cross-Entropy** for two-class problems and **Categorical Cross-Entropy** for multi-class classification, utilizing PyTorch's fused `cross_entropy` function for numerical stability.
- Implement **Label Smoothing** (via `label_smoothing` parameter) to prevent overconfidence and improve model calibration on noisy datasets.
- Deploy **Focal Loss** with a focusing parameter \(\gamma \geq 1\) to handle severe class imbalance by down-weighting easy examples.
- Leverage **InfoNCE** for self-supervised and contrastive learning by treating similarity matching as a batch-wide classification problem.
- Consult [`phases/03-deep-learning-core/05-loss-functions/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/03-deep-learning-core/05-loss-functions/code/main.py) for complete implementations that demonstrate gradient flow through each loss computation.

## Frequently Asked Questions

### When should I use MSE versus cross-entropy in neural networks?

Use **MSE** for regression problems where outputs represent continuous values, as it minimizes squared deviations from target values. Use **cross-entropy** for classification tasks because it penalizes the log-probability of incorrect predictions, forcing the model to learn confident, discriminative features rather than converging to mean values that minimize squared error without improving accuracy.

### What is the advantage of using Focal Loss over standard cross-entropy?

**Focal Loss** adds a modulating factor \((1-p_t)^\gamma\) to the standard cross-entropy equation, which reduces the loss contribution from easy examples (those with high \(p_t\)) and focuses training on hard negatives. This prevents the vast number of easy background examples in imbalanced datasets (like object detection) from overwhelming the gradient during optimization.

### How does label smoothing prevent overconfidence in neural networks?

**Label smoothing** replaces hard targets (0 and 1) with soft targets (e.g., 0.1 and 0.9), effectively mixing the ground-truth distribution with a uniform distribution. This prevents the model from assigning extreme probabilities (0.999+) to training examples, which improves generalization and calibration when the model encounters ambiguous inputs during inference.

### Why is InfoNCE loss used in contrastive learning?

**InfoNCE** transforms the problem of learning similar representations into a classification task where the model must identify the positive pair among all other samples in the batch. By using a temperature-scaled softmax over similarity scores, it optimizes the relative ranking of embeddings, making it scalable for large-batch self-supervised learning in computer vision and multimodal AI systems.