# DeepLearning-500-questions | scutan90 | Knowledge Base | Instagit

深度学习500问，以问答形式对常用的概率知识、线性代数、机器学习、深度学习、计算机视觉等热点问题进行阐述，以帮助自己及有需要的读者。 全书分为18个章节，50余万字。由于水平有限，书中不妥之处恳请广大读者批评指正。   未完待续............ 如有意合作，联系scutjy2015@163.com                     版权所有，违权必究       Tan 2018.06

GitHub Stars: 57.2k

Repository: https://github.com/scutan90/DeepLearning-500-questions

---

## Articles

### [TensorRT vs ONNX vs TFLite for Model Deployment: A Technical Comparison](/scutan90/DeepLearning-500-questions/differences-tensorrt-onnx-tflite-model-deployment)

Compare TensorRT, ONNX, and TFLite for model deployment. Understand NVIDIA GPU optimization, mobile CPU conversion, and framework-agnostic formats to choose the best solution.

- Tags: deep-dive
- Published: 2026-03-06

### [How the Attention Mechanism in Self-Supervised Models Improves Representation Learning](/scutan90/DeepLearning-500-questions/attention-mechanism-self-supervised-models-representation-learning)

Discover how attention mechanisms in self-supervised models enhance representation learning. They use dynamic weightings to focus on relevant features, improving model performance and adaptability.

- Tags: deep-dive
- Published: 2026-03-06

### [Data Augmentation in Computer Vision Deep Learning: Best Practices and Implementation Guide](/scutan90/DeepLearning-500-questions/best-practices-data-augmentation-computer-vision-deep-learning)

Master data augmentation for computer vision deep learning. Discover best practices for geometric and photometric transformations, curriculum learning, and label integrity to boost model performance.

- Tags: best-practices
- Published: 2026-03-06

### [How to Handle Class Imbalance in Object Detection and Semantic Segmentation: 6 Proven Methods](/scutan90/DeepLearning-500-questions/handle-class-imbalance-object-detection-semantic-segmentation)

Combat class imbalance in object detection and semantic segmentation with 6 proven methods. Learn to use Focal Loss, hard example mining, and data augmentation effectively for better model performance.

- Tags: how-to-guide
- Published: 2026-03-06

### [Global Average Pooling vs Fully Connected Layers in CNNs: Key Differences Explained](/scutan90/DeepLearning-500-questions/differences-global-average-pooling-fully-connected-layers-cnns)

Understand global average pooling vs fully connected layers in CNNs. Discover how GAP reduces parameters while FC layers require millions of weights for image classification.

- Tags: deep-dive
- Published: 2026-03-06

### [How to Implement Early Stopping, Checkpointing, and Model Ensembling in Deep Learning](/scutan90/DeepLearning-500-questions/implement-early-stopping-checkpointing-model-ensembling)

Implement early stopping, checkpointing, and model ensembling in deep learning to enhance accuracy and robustness. Learn to save best model weights and combine predictions effectively.

- Tags: tutorial
- Published: 2026-03-06

### [Mask R-CNN vs YOLOv3 vs EfficientDet: Key Differences for Instance Segmentation](/scutan90/DeepLearning-500-questions/differences-mask-rcnn-yolov3-efficientdet-instance-segmentation)

Explore key differences between Mask R-CNN, YOLOv3, and EfficientDet for instance segmentation. Understand their unique architectures and performance trade-offs for your next computer vision project.

- Tags: deep-dive
- Published: 2026-03-06

### [How to Optimize Deep Learning Models for GPU vs CPU Inference Performance: Architectural Strategies and Code Examples](/scutan90/DeepLearning-500-questions/optimize-deep-learning-models-gpu-vs-cpu-inference)

Optimize deep learning models for GPU and CPU inference. Discover architectural strategies and code examples for TensorRT, OpenVINO, and OneDNN performance tuning.

- Tags: performance
- Published: 2026-03-06

### [1x1 vs 3x3 vs Dilated Convolutions in CNNs: Key Differences Explained](/scutan90/DeepLearning-500-questions/differences-between-1x1-3x3-dilated-convolutions-cnns)

Understand the key differences between 1x1, 3x3, and dilated convolutions in CNNs. Learn how each convolution type extracts features and modifies dimensionality for better deep learning models.

- Tags: deep-dive
- Published: 2026-03-06

### [How GAN Generator and Discriminator Training Works: Architecture, Objectives, and Failure Modes](/scutan90/DeepLearning-500-questions/gan-generator-discriminator-training-failure-modes)

Understand GAN generator and discriminator training a complex minimax game Explore architectures objectives and failure modes like mode collapse and vanishing gradients to build stable models.

- Tags: deep-dive
- Published: 2026-03-06

### [Best Strategies for Learning Rate Scheduling and Adaptive Optimizers in Deep Learning](/scutan90/DeepLearning-500-questions/best-strategies-learning-rate-scheduling-adaptive-optimizers)

Master deep learning training with optimal learning rate scheduling and adaptive optimizers. Discover key strategies for stable convergence and improved model performance.

- Tags: best-practices
- Published: 2026-03-06

### [How to Debug a Deep Neural Network That Is Not Training Properly: A 10-Step Diagnostic Checklist](/scutan90/DeepLearning-500-questions/how-to-debug-deep-neural-network-not-training-properly)

Debug your deep neural network systematically. Follow 10 diagnostic steps to fix training issues caused by data, architecture, or hyperparamters. Improve your model performance now.

- Tags: how-to-guide
- Published: 2026-03-06

### [L1 vs L2 vs Dropout Regularization: Key Differences in Deep Learning](/scutan90/DeepLearning-500-questions/differences-between-l1-l2-dropout-regularization)

Understand L1 L2 and Dropout regularization key differences in deep learning. Learn how L1 zeros weights L2 shrinks them and Dropout prevents co-adaptation for better model performance.

- Tags: deep-dive
- Published: 2026-03-06

### [How Encoder-Decoder Architecture Differs in U-Net, FCN, and DeepLab](/scutan90/DeepLearning-500-questions/encoder-decoder-architecture-differences-unet-fcn-deeplab)

Understand the encoder-decoder architecture differences in U-Net, FCN, and DeepLab. Explore skip connections, transposed convolutions, and atrous convolutions to improve image segmentation models.

- Tags: deep-dive
- Published: 2026-03-06

### [Mathematical Differences Between Cross-Entropy, Hinge, and Contrastive Loss Functions](/scutan90/DeepLearning-500-questions/mathematical-differences-between-cross-entropy-hinge-contrastive-loss)

Explore the mathematical differences between cross-entropy, hinge, and contrastive loss functions. Understand how each optimizes model performance for distinct tasks in deep learning.

- Tags: deep-dive
- Published: 2026-03-06

### [How to Use Transfer Learning and Fine-Tuning with Limited Datasets: A Practical Guide](/scutan90/DeepLearning-500-questions/effective-transfer-learning-fine-tuning-with-limited-datasets)

Master transfer learning and fine-tuning with limited datasets. Freeze early layers and retrain later ones for efficient model adaptation and improved performance.

- Tags: how-to-guide
- Published: 2026-03-06

### [Trade‑Offs Between YOLO, SSD, and Faster R‑CNN for Real‑Time Object Detection](/scutan90/DeepLearning-500-questions/trade-offs-between-yolo-ssd-faster-rcnn-for-real-time-object-detection)

Discover YOLO SSD and Faster R-CNN trade-offs for real-time object detection. Compare speed accuracy and latency to choose the best model for your needs.

- Tags: deep-dive
- Published: 2026-03-06

### [How to Implement Model Quantization and Pruning for Mobile Deployment: A Complete Guide](/scutan90/DeepLearning-500-questions/implement-model-quantization-pruning-for-mobile-deployment)

Implement model quantization and pruning for mobile deployment to compress neural networks. Achieve faster real-time inference on ARM CPUs with minimal accuracy loss.

- Tags: tutorial
- Published: 2026-03-06

### [How to Handle Gradient Vanishing and Explosion in Deep Neural Networks: 8 Proven Mitigation Strategies](/scutan90/DeepLearning-500-questions/best-practices-for-handling-gradient-vanishing-and-explosion)

Resolve gradient vanishing and explosion in deep networks with 8 proven strategies including weight initialization, ReLU, Batch Norm, gradient clipping, and residual connections for stable gradient flow.

- Tags: best-practices
- Published: 2026-03-06

### [How Does the Attention Mechanism Work in Transformers vs LSTM/GRU](/scutan90/DeepLearning-500-questions/attention-mechanism-transformer-vs-lstm-gru)

Understand the attention mechanism in Transformers versus LSTM/GRU. Discover how Transformers use parallel self-attention and LSTM/GRU use sequential encoder-decoder attention for focused input processing.

- Tags: deep-dive
- Published: 2026-03-06

### [Key Differences Between ResNet, ResNeXt, and DenseNet Architectures](/scutan90/DeepLearning-500-questions/key-differences-between-resnet-resnext-and-densenet-architectures)

Explore ResNet, ResNeXt, and DenseNet architectures key differences. Learn how ResNet uses identity shortcuts, ResNeXt enhances parallel paths, and DenseNet maximizes feature reuse for efficient deep learning models.

- Tags: deep-dive
- Published: 2026-03-06

### [How to Choose Between Batch Normalization, Layer Normalization, and Group Normalization](/scutan90/DeepLearning-500-questions/how-to-choose-between-batch-normalization-layer-normalization-and-group-normalization)

Understand Batch Layer and Group Normalization. Learn how to choose the right normalization technique for your deep learning model to stabilize training and improve performance.

- Tags: deep-dive
- Published: 2026-03-06

