# ML Fundamentals in Phase 2 of AI Engineering from Scratch: The Complete 18-Lesson Curriculum

> Discover essential ML fundamentals in Phase 2 of ai engineering from scratch. This 18-lesson curriculum covers classical ML algorithms with hands-on Python and TypeScript examples.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: getting-started
- Published: 2026-07-23

---

**Phase 2 of the rohitg00/ai-engineering-from-scratch repository delivers a comprehensive machine learning fundamentals curriculum spanning 18 sequential lessons that teach classical ML algorithms—from linear regression to ensemble methods—through hands-on Python and TypeScript implementations without heavyweight dependencies.**

The second phase of the open-source ai-engineering-from-scratch project provides a self-contained machine learning toolbox designed to bridge theoretical concepts with production-ready code. Each lesson in `phases/02-ml-fundamentals/` combines concise documentation, reference implementations, and mastery quizzes to build classical ML pipelines from first principles.

## Curriculum Structure and Learning Philosophy

The Phase 2 directory follows a strict pedagogical contract enforced by [`scripts/audit_lessons.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/audit_lessons.py), requiring every lesson to include three components: explanatory content in [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md), executable reference code in `code/main.<lang>`, and a validation quiz in [`quiz.json`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/quiz.json). This structure aligns with the repository's philosophy documented in [`AGENTS.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/AGENTS.md), emphasizing implementations that minimize external dependencies while remaining fully unit-tested.

The entry point for this phase is [`phases/02-ml-fundamentals/README.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/README.md), which outlines the progression from basic supervised learning to advanced topics like anomaly detection.

## Supervised Learning Foundations (Lessons 01–06)

The curriculum begins with conceptual framing in **Lesson 01: What is Machine Learning?**, located at [`phases/02-ml-fundamentals/01-what-is-machine-learning/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/01-what-is-machine-learning/docs/en.md), before diving into algorithmic implementations.

**Linear Regression** (Lesson 02) demonstrates closed-form solutions and gradient descent optimization. The implementation in [`phases/02-ml-fundamentals/02-linear-regression/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/02-linear-regression/code/main.py) calculates the normal equation using NumPy:

```python
import numpy as np

# Toy dataset

X = np.array([[1], [2], [3], [4]], dtype=float)
y = np.array([2, 4, 6, 8], dtype=float)

# Add bias term

X_b = np.c_[np.ones((len(X), 1)), X]

# Closed‑form solution: θ = (XᵀX)⁻¹Xᵀy

theta = np.linalg.inv(X_b.T @ X_b) @ X_b.T @ y
print("Intercept:", theta[0], "Slope:", theta[1])

```

**Logistic Regression** (Lesson 03) covers binary classification, sigmoid activation functions, and regularization strategies. **Decision Trees** (Lesson 04) implement impurity measures and recursive partitioning with pruning strategies, providing interpretable tree-based models.

**Support Vector Machines** (Lesson 05) explore margin maximization and kernel methods, while **K-Nearest Neighbors & Distances** (Lesson 06) focus on instance-based learning and computational optimizations for distance metrics.

## Unsupervised Learning and Dimensionality Reduction (Lesson 07)

**Lesson 07: Unsupervised Learning** introduces clustering via k-means and dimensionality reduction through Principal Component Analysis (PCA). The lesson emphasizes evaluation metrics for clustering quality and variance retention, providing algorithms that function without labeled training data.

## Feature Engineering and Selection (Lessons 08, 18)

The curriculum dedicates significant attention to data preparation. **Lesson 08: Feature Engineering** addresses scaling, encoding categorical variables, and creating interaction terms. **Lesson 18: Feature Selection** implements filter, wrapper, and embedded methods to reduce dimensionality, including mutual information calculations in [`phases/02-ml-fundamentals/18-feature-selection/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/18-feature-selection/code/main.py):

```python
def mutual_info(x, y, bins=10):
    # Discretize

    x_bins = np.digitize(x, np.histogram_bin_edges(x, bins))
    y_bins = np.digitize(y, np.histogram_bin_edges(y, bins))
    # Joint histogram

    pxy, _, _ = np.histogram2d(x_bins, y_bins, bins=bins, density=True)
    px = pxy.sum(axis=1)
    py = pxy.sum(axis=0)
    # Compute MI

    mi = (pxy * np.log(pxy / (px[:, None] * py[None, :] + 1e-9) + 1e-9)).sum()
    return mi

```

## Model Evaluation and Optimization (Lessons 09–13)

Critical production skills dominate the middle section of Phase 2. **Lesson 09: Model Evaluation** establishes metrics for regression (MSE, RMSE) and classification (accuracy, F1-score), alongside cross-validation strategies and confusion matrix analysis.

**Lesson 10: Bias-Variance Trade-off** provides theoretical foundations for underfitting and overfitting, utilizing learning curves to diagnose model capacity. **Lesson 11: Ensemble Methods** covers bagging (random forests), boosting implementations, and stacking architectures to improve predictive performance.

**Lesson 12: Hyper-parameter Tuning** implements grid search, random search, and Bayesian optimization basics. **Lesson 13: ML Pipelines** teaches the construction of reproducible workflows that chain preprocessing, model fitting, and validation steps.

## Specialized ML Applications (Lessons 14–17)

The final lessons address domain-specific challenges. **Naïve Bayes** (Lesson 14) implements probabilistic classification with handling for both categorical and continuous features. **Time-Series Forecasting** (Lesson 15) covers ARIMA basics, seasonality detection, and temporal evaluation metrics located in [`phases/02-ml-fundamentals/15-time-series/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/15-time-series/docs/en.md).

**Anomaly Detection** (Lesson 16) explores statistical thresholds and model-based outlier identification. **Imbalanced Data** (Lesson 17) provides strategies including resampling techniques, class weighting adjustments, and precision-recall analysis for skewed distributions.

## Decision Tree Implementation Example

The Decision Trees lesson demonstrates recursive partitioning using Gini impurity calculations. The reference implementation in [`phases/02-ml-fundamentals/04-decision-trees/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/04-decision-trees/code/main.py) defines the tree structure as follows:

```python
class Node:
    def __init__(self, gini, samples, value, left=None, right=None):
        self.gini = gini
        self.samples = samples
        self.value = value
        self.left = left
        self.right = right

def gini_impurity(y):
    _, counts = np.unique(y, return_counts=True)
    probs = counts / counts.sum()
    return 1 - np.sum(probs ** 2)

# Very minimal recursive split (illustrative)

def grow_tree(X, y, depth=0, max_depth=3):
    if depth == max_depth or len(np.unique(y)) == 1:
        return Node(gini=gini_impurity(y), samples=len(y), value=np.bincount(y).argmax())
    # split on first feature at median

    thresh = np.median(X[:, 0])
    left_idx = X[:, 0] <= thresh
    right_idx = ~left_idx
    left = grow_tree(X[left_idx], y[left_idx], depth + 1, max_depth)
    right = grow_tree(X[right_idx], y[right_idx], depth + 1, max_depth)
    return Node(gini=gini_impurity(y), samples=len(y), value=None, left=left, right=right)

```

## Summary

Phase 2 of the ai-engineering-from-scratch curriculum provides:

- **18 sequential lessons** covering supervised, unsupervised, and specialized ML techniques from `phases/02-ml-fundamentals/01-what-is-machine-learning/` through `phases/02-ml-fundamentals/18-feature-selection/`
- **First-principles implementations** in `phases/02-ml-fundamentals/*/code/main.<lang>` without heavy framework dependencies
- **Comprehensive evaluation frameworks** including bias-variance analysis, cross-validation strategies, and confusion matrix interpretations
- **Production-ready pipelines** that chain preprocessing, hyperparameter tuning, and model validation via `phases/02-ml-fundamentals/13-ml-pipelines/`
- **Validation through unit tests** enforced by [`scripts/audit_lessons.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/audit_lessons.py) and concept mastery via [`quiz.json`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/quiz.json) files in each lesson directory

## Frequently Asked Questions

### What prerequisites are needed for Phase 2 ML fundamentals?

Learners should complete Phase 1 (Python/TypeScript basics and linear algebra fundamentals) before attempting Phase 2. The curriculum assumes familiarity with NumPy arrays and basic probability theory, though these concepts are reinforced through implementation practice in lessons like [`phases/02-ml-fundamentals/02-linear-regression/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/02-linear-regression/code/main.py).

### How long does it take to complete all 18 lessons in Phase 2?

Each lesson requires approximately 3-5 hours of study time, including reading [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md), implementing the algorithms in the `code/` directory, and passing the 6-question [`quiz.json`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/quiz.json) validation. The complete Phase 2 curriculum typically requires 2-3 weeks of dedicated study to master the ML fundamentals.

### Are the implementations in Phase 2 production-ready or educational only?

While designed for educational clarity following the repository's [`AGENTS.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/AGENTS.md) guidelines, the implementations adhere to production standards enforced by [`scripts/audit_lessons.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/audit_lessons.py). Each lesson includes unit tests and self-terminating patterns that allow direct integration into ML pipelines, though production use should include additional error handling and scaling optimizations.

### Which lesson covers cross-validation and model selection?

**Lesson 09: Model Evaluation** covers k-fold cross-validation, confusion matrices, and metric selection for both regression and classification tasks. The lesson files reside in `phases/02-ml-fundamentals/09-model-evaluation/` and include implementations of stratified sampling techniques and learning curve analysis.