Techniques for Handling Imbalanced Data in AI: A Complete Pipeline Guide

Use class weights for mild imbalance (≈80/20), apply SMOTE with threshold tuning for moderate (≈95/5), and combine SMOTE, class weights, and threshold tuning for severe (≈99/1) cases, while evaluating with F1, AUPRC, or MCC instead of accuracy.

The "Handling Imbalanced Data" lesson in the rohitg00/ai-engineering-from-scratch repository (Phase 2, Lesson 17) provides a production-ready framework for training models when one class dominates the dataset. According to the source documentation in phases/02-ml-fundamentals/17-imbalanced-data/docs/en.md, traditional accuracy metrics can exceed 99% while the model completely fails to predict minority classes, making specialized techniques essential for domains like fraud detection, medical diagnosis, and intrusion detection.

Why Accuracy Metrics Fail for Imbalanced Datasets

In imbalanced classification problems, accuracy becomes a deceptive metric. The curriculum demonstrates that a model can achieve greater than 99% accuracy by always predicting the majority class, effectively never catching a single fraud case or diseased patient. The confusion-matrix examples in the lesson illustrate that Precision, Recall, F1-score, F-beta, AUPRC (Area Under the Precision-Recall Curve), and Matthews Correlation Coefficient (MCC) provide truthful performance pictures when classes are heavily skewed.

The Imbalanced-Data Pipeline: Choosing the Right Strategy

The lesson defines a decision pipeline based on the imbalance ratio, guiding practitioners from "Start: Imbalanced Dataset" to "Deploy with Monitoring" through specific intervention levels.

Mild Imbalance (≈80/20): Class Weights

For datasets with roughly 80% majority and 20% minority classes, the pipeline recommends class weighting as the first-line solution. This technique adjusts the loss function to penalize minority misclassifications more heavily without altering the underlying data distribution.

Moderate Imbalance (≈95/5): SMOTE and Threshold Tuning

At ratios around 95/5, the strategy escalates to SMOTE (Synthetic Minority Oversampling Technique) combined with threshold tuning. SMOTE synthesizes new minority samples by interpolating between existing examples and their k-nearest neighbors, while threshold tuning optimizes the decision boundary post-training to maximize F1 or AUPRC.

Severe Imbalance (≈99/1): Combined Approach

For extreme cases like 99/1 ratios, the source recommends a triple strategy: SMOTE + Class Weights + Threshold Tuning. This maximum-intervention approach addresses the severe scarcity of minority examples through data augmentation, algorithmic weighting, and post-processing optimization, as visualized in the decision flowchart within the lesson documentation.

Data-Level Techniques: Resampling Strategies

The lesson details three core resampling methods in phases/02-ml-fundamentals/17-imbalanced-data/docs/en.md:

Random Oversampling and Undersampling

Random oversampling duplicates existing minority samples to balance the dataset. While simple to implement, it risks overfitting to duplicated examples. Random undersampling removes majority samples to achieve balance, which is computationally fast but potentially discards valuable training data and model capacity.

SMOTE Implementation

SMOTE generates synthetic minority instances by drawing lines between a sample and one of its k-nearest minority neighbors, then selecting points along those lines. This avoids the overfitting risk of simple duplication while expanding the minority class boundary.

from imblearn.over_sampling import SMOTE

# Apply SMOTE with k=5 neighbors

smote = SMOTE(k_neighbors=5, random_state=42)
X_resampled, y_resampled = smote.fit_resample(X, y)

Algorithm-Level Techniques

Class Weighting

Rather than changing the data, class weighting modifies the learning algorithm through the loss function. The lesson derives the inverse-frequency weight formula and demonstrates implementation in scikit-learn:

from sklearn.linear_model import LogisticRegression

# Automatically balance weights based on class frequencies

clf = LogisticRegression(class_weight='balanced', max_iter=1000)
clf.fit(X_train, y_train)

Cost-Sensitive Learning

Cost-sensitive learning generalizes class weighting by assigning explicit misclassification costs. For example, setting a false negative cost 100× higher than the false positive cost reflects business realities in fraud detection. The source includes cost matrix examples showing how to incorporate these asymmetric penalties directly into the training objective.

Post-Training Optimization: Threshold Tuning

Instead of using the default 0.5 probability threshold, the lesson advocates sweeping across thresholds on validation data to maximize the chosen metric. This technique is particularly effective when combined with resampling strategies and is visualized in the curriculum with implementation code:

import numpy as np
from sklearn.metrics import f1_score

# Get validation probabilities for the positive class

probs = clf.predict_proba(X_val)[:, 1]
thresholds = np.linspace(0, 1, 101)

# Find threshold maximizing F1 score

f1_scores = [f1_score(y_val, (probs >= t).astype(int)) for t in thresholds]
best_threshold = thresholds[np.argmax(f1_scores)]
print(f'Best threshold: {best_threshold:.2f}')

Evaluation Metrics That Matter

The curriculum in phases/02-ml-fundamentals/17-imbalanced-data/docs/en.md emphasizes abandoning accuracy in favor of Precision, Recall, F1, AUPRC, and MCC. These metrics account for the true positive rate on minority classes, ensuring models perform well on both classes rather than optimizing for majority-class predictions alone.

Summary

  • Accuracy fails in imbalanced scenarios; use F1, AUPRC, or MCC instead.
  • Match technique severity to imbalance ratio: Class weights for mild (80/20), SMOTE + threshold for moderate (95/5), and combined approaches for severe (99/1).
  • SMOTE synthesizes minority samples by interpolating between k-nearest neighbors to avoid overfitting.
  • Class weighting and cost-sensitive learning adjust the loss function to penalize minority errors more heavily without data modification.
  • Threshold tuning optimizes the decision boundary after training to maximize domain-relevant metrics, and is documented in the skill checklist at phases/02-ml-fundamentals/17-imbalanced-data/outputs/skill-imbalanced-data.md.

Frequently Asked Questions

What is the best technique for handling imbalanced data with a 99/1 ratio?

For severe imbalance around 99/1, the curriculum recommends combining SMOTE for data augmentation, class weights to penalize minority misclassification, and threshold tuning to optimize the decision boundary. This triple intervention addresses data scarcity, algorithmic bias, and decision threshold simultaneously, followed by evaluation using AUPRC or MCC rather than accuracy.

Should I use SMOTE before or after splitting my data into train and test sets?

Always apply SMOTE after splitting data, fitting the resampler only on the training set to prevent data leakage. The synthetic samples should be generated exclusively from training data, ensuring your test set reflects the true imbalanced distribution that the model will face in production.

How do I choose between precision and recall as my optimization target?

Choose Precision when false positives are costly (e.g., spam filtering where legitimate emails marked as spam are problematic), and Recall when false negatives are more dangerous (e.g., fraud detection or disease diagnosis where missing positive cases has severe consequences). The F-beta metric allows weighting recall β times more important than precision when you need to balance both concerns.

Can threshold tuning replace SMOTE for imbalanced datasets?

Threshold tuning alone can improve performance but works best when combined with other techniques. For moderate to severe imbalance, the lesson recommends using threshold tuning alongside SMOTE or class weights, as tuning alone cannot compensate for insufficient minority examples in the training data that prevent the model from learning minority class patterns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →