Common Loss Functions in Neural Networks: A Complete Guide with PyTorch Implementations
Neural networks learn by minimizing loss functions that quantify the discrepancy between predictions and targets, with Mean-Squared Error governing regression tasks and cross-entropy variants handling classification, while specialized losses like Focal and InfoNCE address imbalanced data and contrastive learning.
Selecting the appropriate loss function determines whether a model learns meaningful patterns or converges to suboptimal solutions. The rohitg00/ai-engineering-from-scratch curriculum provides a comprehensive examination of common loss functions in neural networks, from foundational regression criteria to advanced contrastive objectives. This article breaks down each loss type, explains when to apply them, and provides runnable PyTorch code based on the repository's reference implementations.
Core Loss Functions: Regression vs. Classification
Neural network architectures require different optimization objectives depending on whether they predict continuous values or discrete categories. The repository's lesson documentation in phases/03-deep-learning-core/05-loss-functions/docs/en.md establishes three fundamental building blocks.
Mean-Squared Error (MSE) for Regression
Use MSE when your network outputs continuous values. The formula (\frac{1}{N}\sum (y_i-\hat y_i)^2) calculates the average squared difference between predictions (\hat y) and targets (y).
This loss is smooth and convex, penalizing large errors quadratically. However, as demonstrated in the curriculum, MSE can be "gamed" in binary tasks—a model predicting 0.5 for every sample achieves an MSE of 0.25 without learning discriminative features.
Binary Cross-Entropy (BCE) for Two-Class Problems
For binary classification, Binary Cross-Entropy applies the formula (-[y\log p+(1-y)\log(1-p)]), where (p) represents the predicted probability for the positive class.
The gradients simplify to (p - y), providing strong signals when predictions are confident and wrong. In PyTorch, always use F.binary_cross_entropy_with_logits rather than manually applying a sigmoid followed by BCE, as this preserves numerical stability.
Categorical Cross-Entropy for Multi-Class Classification
When exactly one class is correct among many, Categorical Cross-Entropy (implemented via softmax cross-entropy) uses (-\log \frac{e^{z_{y}}}{\sum_j e^{z_j}}), where (z_y) is the logit for the true class.
According to the source code analysis, you should call torch.nn.functional.cross_entropy directly rather than chaining softmax and NLLLoss, as the combined operation avoids numerical instability by computing the log-softmax in a single stable pass.
Advanced Loss Functions for Specialized Tasks
Beyond basic classification and regression, modern deep learning requires losses that handle noisy labels, class imbalance, and representation learning.
Label-Smoothed Cross-Entropy for Calibration
Label smoothing prevents models from becoming overconfident by blending hard targets (0 and 1) with a uniform distribution. Instead of targeting 1.0 for the correct class, the model optimizes toward (1 - \epsilon) (e.g., 0.9) and distributes the remaining (\epsilon) across other classes.
In PyTorch, enable this by setting label_smoothing=0.1 in F.cross_entropy. This regularization technique proves particularly effective when dealing with noisy labels or when model calibration is critical.
Focal Loss for Highly Imbalanced Datasets
Focal Loss addresses extreme class imbalance (common in object detection) by down-weighting easy examples. The formula (- (1-p_t)^\gamma \log(p_t)) introduces a modulating factor where (p_t) is the probability for the true class and (\gamma) controls the focusing strength.
When (\gamma = 2), the loss reduces the gradient contribution from well-classified examples by ((1-p_t)^2), forcing the optimizer to concentrate on hard negatives and minority class instances.
InfoNCE for Contrastive and Self-Supervised Learning
InfoNCE converts similarity learning into a classification problem over batch elements. The formula (-\log \frac{e^{s_{ii}/\tau}}{\sum_j e^{s_{ij}/\tau}}) compares the similarity (s_{ii}) of a positive pair against all negative pairs (s_{ij}), scaled by temperature (\tau).
This loss powers modern vision-language models and self-supervised learning frameworks by pulling matching embeddings together while pushing non-matching pairs apart. The repository notes that InfoNCE effectively uses the same optimization machinery as standard cross-entropy, making it both scalable and simple to implement.
KL-Divergence for Knowledge Distillation
KL-Divergence measures how one probability distribution diverges from another, computed as (\sum_y p(y) \log \frac{p(y)}{q(y)}) where (p) is the target distribution and (q) is the model's output.
This loss becomes equivalent to cross-entropy when the target is one-hot, but proves essential for knowledge distillation where the "teacher" model provides soft probability distributions rather than hard labels.
Why Loss Function Selection Determines Model Performance
The curriculum emphasizes that loss choice directly impacts learned representations. In phases/03-deep-learning-core/05-loss-functions/docs/en.md, the authors demonstrate that MSE fails for binary tasks because it plateaus when predictions approach 0.5, whereas cross-entropy applies an infinite penalty for confident wrong predictions through the (-\log(p)) term.
This mathematical property forces the network to push probabilities toward the extremes (0 or 1), learning discriminative features rather than conservative averages. Similarly, Focal Loss solves the dominance of background examples in detection tasks, while InfoNCE enables scalable contrastive learning by treating each batch as a classification problem over instance indices.
Practical PyTorch Implementation Examples
The reference implementation in phases/03-deep-learning-core/05-loss-functions/code/main.py demonstrates each loss using PyTorch's functional API. Below is a consolidated example covering all major loss types:
import torch
import torch.nn.functional as F
# Dummy data: 4 samples, 3 classes
logits = torch.randn(4, 3, requires_grad=True)
labels = torch.tensor([0, 2, 1, 2])
# 1. Mean-Squared Error (regression)
target_reg = torch.randn(4, 3)
mse = F.mse_loss(logits, target_reg)
mse.backward()
print('MSE:', mse.item())
# 2. Binary Cross-Entropy
bce = F.binary_cross_entropy_with_logits(
logits[:, 0],
torch.tensor([1., 0., 1., 0.])
)
bce.backward()
print('BCE:', bce.item())
# 3. Categorical Cross-Entropy
ce = F.cross_entropy(logits, labels)
ce.backward()
print('Cross-Entropy:', ce.item())
# 4. Label-Smoothed Cross-Entropy
ce_smooth = F.cross_entropy(logits, labels, label_smoothing=0.1)
ce_smooth.backward()
print('Label-smoothed CE:', ce_smooth.item())
# 5. Focal Loss (γ=2)
prob = F.softmax(logits, dim=1)
pt = prob.gather(1, labels.unsqueeze(1)).squeeze()
focal = -(1 - pt) ** 2 * torch.log(pt + 1e-12)
focal = focal.mean()
focal.backward()
print('Focal loss:', focal.item())
# 6. InfoNCE (contrastive) - batch of 4 embeddings
emb = F.normalize(torch.randn(4, 64), dim=1)
sim = emb @ emb.t() / 0.07 # temperature τ=0.07
labels_contrast = torch.arange(4)
info_nce = F.cross_entropy(sim, labels_contrast)
info_nce.backward()
print('InfoNCE:', info_nce.item())
Always use the high-level functional APIs (F.mse_loss, F.cross_entropy, etc.) because they fuse operations for numerical stability. For example, F.cross_entropy combines log_softmax and NLLLoss in a single kernel, preventing the overflow issues common when computing softmax separately.
Key Files in the ai-engineering-from-scratch Repository
| Resource | Path | Purpose |
|---|---|---|
| Lesson Documentation | phases/03-deep-learning-core/05-loss-functions/docs/en.md |
Conceptual exposition, mathematical derivations, and pedagogical notes on each loss function. |
| Reference Implementation | phases/03-deep-learning-core/05-loss-functions/code/main.py |
From-scratch implementations of MSE, BCE, categorical CE, label smoothing, focal loss, and InfoNCE with gradient calculations. |
| Loss Selector Prompt | phases/03-deep-learning-core/05-loss-functions/outputs/prompt-loss-function-selector.md |
Interactive decision tree for selecting appropriate losses based on task characteristics. |
| Loss Debugger Prompt | phases/03-deep-learning-core/05-loss-functions/outputs/prompt-loss-debugger.md |
Common pitfalls and debugging strategies, such as identifying duplicate softmax applications. |
These resources follow the curriculum's philosophy of building from raw mathematical principles before using high-level framework APIs.
Summary
- Use Mean-Squared Error for regression tasks with continuous targets, but avoid it for classification where it permits uninformative average predictions.
- Apply Binary Cross-Entropy for two-class problems and Categorical Cross-Entropy for multi-class classification, utilizing PyTorch's fused
cross_entropyfunction for numerical stability. - Implement Label Smoothing (via
label_smoothingparameter) to prevent overconfidence and improve model calibration on noisy datasets. - Deploy Focal Loss with a focusing parameter (\gamma \geq 1) to handle severe class imbalance by down-weighting easy examples.
- Leverage InfoNCE for self-supervised and contrastive learning by treating similarity matching as a batch-wide classification problem.
- Consult
phases/03-deep-learning-core/05-loss-functions/code/main.pyfor complete implementations that demonstrate gradient flow through each loss computation.
Frequently Asked Questions
When should I use MSE versus cross-entropy in neural networks?
Use MSE for regression problems where outputs represent continuous values, as it minimizes squared deviations from target values. Use cross-entropy for classification tasks because it penalizes the log-probability of incorrect predictions, forcing the model to learn confident, discriminative features rather than converging to mean values that minimize squared error without improving accuracy.
What is the advantage of using Focal Loss over standard cross-entropy?
Focal Loss adds a modulating factor ((1-p_t)^\gamma) to the standard cross-entropy equation, which reduces the loss contribution from easy examples (those with high (p_t)) and focuses training on hard negatives. This prevents the vast number of easy background examples in imbalanced datasets (like object detection) from overwhelming the gradient during optimization.
How does label smoothing prevent overconfidence in neural networks?
Label smoothing replaces hard targets (0 and 1) with soft targets (e.g., 0.1 and 0.9), effectively mixing the ground-truth distribution with a uniform distribution. This prevents the model from assigning extreme probabilities (0.999+) to training examples, which improves generalization and calibration when the model encounters ambiguous inputs during inference.
Why is InfoNCE loss used in contrastive learning?
InfoNCE transforms the problem of learning similar representations into a classification task where the model must identify the positive pair among all other samples in the batch. By using a temperature-scaled softmax over similarity scores, it optimizes the relative ranking of embeddings, making it scalable for large-batch self-supervised learning in computer vision and multimodal AI systems.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →