# Techniques for Handling Imbalanced Data in AI: A Complete Pipeline Guide

> Master imbalanced data in AI. Learn techniques like class weights SMOTE and threshold tuning for all imbalance levels. Optimize evaluation beyond accuracy with this complete pipeline guide.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: tutorial
- Published: 2026-07-19

---

**Use class weights for mild imbalance (≈80/20), apply SMOTE with threshold tuning for moderate (≈95/5), and combine SMOTE, class weights, and threshold tuning for severe (≈99/1) cases, while evaluating with F1, AUPRC, or MCC instead of accuracy.**

The "Handling Imbalanced Data" lesson in the `rohitg00/ai-engineering-from-scratch` repository (Phase 2, Lesson 17) provides a production-ready framework for training models when one class dominates the dataset. According to the source documentation in [`phases/02-ml-fundamentals/17-imbalanced-data/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/17-imbalanced-data/docs/en.md), traditional accuracy metrics can exceed 99% while the model completely fails to predict minority classes, making specialized techniques essential for domains like fraud detection, medical diagnosis, and intrusion detection.

## Why Accuracy Metrics Fail for Imbalanced Datasets

In imbalanced classification problems, **accuracy** becomes a deceptive metric. The curriculum demonstrates that a model can achieve greater than 99% accuracy by always predicting the majority class, effectively never catching a single fraud case or diseased patient. The confusion-matrix examples in the lesson illustrate that **Precision**, **Recall**, **F1-score**, **F-beta**, **AUPRC** (Area Under the Precision-Recall Curve), and **Matthews Correlation Coefficient (MCC)** provide truthful performance pictures when classes are heavily skewed.

## The Imbalanced-Data Pipeline: Choosing the Right Strategy

The lesson defines a decision pipeline based on the **imbalance ratio**, guiding practitioners from "Start: Imbalanced Dataset" to "Deploy with Monitoring" through specific intervention levels.

### Mild Imbalance (≈80/20): Class Weights

For datasets with roughly 80% majority and 20% minority classes, the pipeline recommends **class weighting** as the first-line solution. This technique adjusts the loss function to penalize minority misclassifications more heavily without altering the underlying data distribution.

### Moderate Imbalance (≈95/5): SMOTE and Threshold Tuning

At ratios around 95/5, the strategy escalates to **SMOTE** (Synthetic Minority Oversampling Technique) combined with **threshold tuning**. SMOTE synthesizes new minority samples by interpolating between existing examples and their k-nearest neighbors, while threshold tuning optimizes the decision boundary post-training to maximize F1 or AUPRC.

### Severe Imbalance (≈99/1): Combined Approach

For extreme cases like 99/1 ratios, the source recommends a triple strategy: **SMOTE + Class Weights + Threshold Tuning**. This maximum-intervention approach addresses the severe scarcity of minority examples through data augmentation, algorithmic weighting, and post-processing optimization, as visualized in the decision flowchart within the lesson documentation.

## Data-Level Techniques: Resampling Strategies

The lesson details three core resampling methods in [`phases/02-ml-fundamentals/17-imbalanced-data/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/17-imbalanced-data/docs/en.md):

### Random Oversampling and Undersampling

**Random oversampling** duplicates existing minority samples to balance the dataset. While simple to implement, it risks overfitting to duplicated examples. **Random undersampling** removes majority samples to achieve balance, which is computationally fast but potentially discards valuable training data and model capacity.

### SMOTE Implementation

**SMOTE** generates synthetic minority instances by drawing lines between a sample and one of its k-nearest minority neighbors, then selecting points along those lines. This avoids the overfitting risk of simple duplication while expanding the minority class boundary.

```python
from imblearn.over_sampling import SMOTE

# Apply SMOTE with k=5 neighbors

smote = SMOTE(k_neighbors=5, random_state=42)
X_resampled, y_resampled = smote.fit_resample(X, y)

```

## Algorithm-Level Techniques

### Class Weighting

Rather than changing the data, **class weighting** modifies the learning algorithm through the loss function. The lesson derives the inverse-frequency weight formula and demonstrates implementation in scikit-learn:

```python
from sklearn.linear_model import LogisticRegression

# Automatically balance weights based on class frequencies

clf = LogisticRegression(class_weight='balanced', max_iter=1000)
clf.fit(X_train, y_train)

```

### Cost-Sensitive Learning

**Cost-sensitive learning** generalizes class weighting by assigning explicit misclassification costs. For example, setting a false negative cost 100× higher than the false positive cost reflects business realities in fraud detection. The source includes cost matrix examples showing how to incorporate these asymmetric penalties directly into the training objective.

## Post-Training Optimization: Threshold Tuning

Instead of using the default 0.5 probability threshold, the lesson advocates sweeping across thresholds on validation data to maximize the chosen metric. This technique is particularly effective when combined with resampling strategies and is visualized in the curriculum with implementation code:

```python
import numpy as np
from sklearn.metrics import f1_score

# Get validation probabilities for the positive class

probs = clf.predict_proba(X_val)[:, 1]
thresholds = np.linspace(0, 1, 101)

# Find threshold maximizing F1 score

f1_scores = [f1_score(y_val, (probs >= t).astype(int)) for t in thresholds]
best_threshold = thresholds[np.argmax(f1_scores)]
print(f'Best threshold: {best_threshold:.2f}')

```

## Evaluation Metrics That Matter

The curriculum in [`phases/02-ml-fundamentals/17-imbalanced-data/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/17-imbalanced-data/docs/en.md) emphasizes abandoning accuracy in favor of **Precision**, **Recall**, **F1**, **AUPRC**, and **MCC**. These metrics account for the true positive rate on minority classes, ensuring models perform well on both classes rather than optimizing for majority-class predictions alone.

## Summary

- **Accuracy fails** in imbalanced scenarios; use **F1**, **AUPRC**, or **MCC** instead.
- **Match technique severity to imbalance ratio**: Class weights for mild (80/20), SMOTE + threshold for moderate (95/5), and combined approaches for severe (99/1).
- **SMOTE** synthesizes minority samples by interpolating between k-nearest neighbors to avoid overfitting.
- **Class weighting** and **cost-sensitive learning** adjust the loss function to penalize minority errors more heavily without data modification.
- **Threshold tuning** optimizes the decision boundary after training to maximize domain-relevant metrics, and is documented in the skill checklist at [`phases/02-ml-fundamentals/17-imbalanced-data/outputs/skill-imbalanced-data.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/17-imbalanced-data/outputs/skill-imbalanced-data.md).

## Frequently Asked Questions

### What is the best technique for handling imbalanced data with a 99/1 ratio?

For severe imbalance around 99/1, the curriculum recommends combining **SMOTE** for data augmentation, **class weights** to penalize minority misclassification, and **threshold tuning** to optimize the decision boundary. This triple intervention addresses data scarcity, algorithmic bias, and decision threshold simultaneously, followed by evaluation using **AUPRC** or **MCC** rather than accuracy.

### Should I use SMOTE before or after splitting my data into train and test sets?

Always apply SMOTE **after** splitting data, fitting the resampler only on the training set to prevent data leakage. The synthetic samples should be generated exclusively from training data, ensuring your test set reflects the true imbalanced distribution that the model will face in production.

### How do I choose between precision and recall as my optimization target?

Choose **Precision** when false positives are costly (e.g., spam filtering where legitimate emails marked as spam are problematic), and **Recall** when false negatives are more dangerous (e.g., fraud detection or disease diagnosis where missing positive cases has severe consequences). The **F-beta** metric allows weighting recall β times more important than precision when you need to balance both concerns.

### Can threshold tuning replace SMOTE for imbalanced datasets?

Threshold tuning alone can improve performance but works best when combined with other techniques. For moderate to severe imbalance, the lesson recommends using threshold tuning **alongside** SMOTE or class weights, as tuning alone cannot compensate for insufficient minority examples in the training data that prevent the model from learning minority class patterns.