ML Fundamentals in Phase 2 of AI Engineering from Scratch: The Complete 18-Lesson Curriculum
Phase 2 of the rohitg00/ai-engineering-from-scratch repository delivers a comprehensive machine learning fundamentals curriculum spanning 18 sequential lessons that teach classical ML algorithms—from linear regression to ensemble methods—through hands-on Python and TypeScript implementations without heavyweight dependencies.
The second phase of the open-source ai-engineering-from-scratch project provides a self-contained machine learning toolbox designed to bridge theoretical concepts with production-ready code. Each lesson in phases/02-ml-fundamentals/ combines concise documentation, reference implementations, and mastery quizzes to build classical ML pipelines from first principles.
Curriculum Structure and Learning Philosophy
The Phase 2 directory follows a strict pedagogical contract enforced by scripts/audit_lessons.py, requiring every lesson to include three components: explanatory content in docs/en.md, executable reference code in code/main.<lang>, and a validation quiz in quiz.json. This structure aligns with the repository's philosophy documented in AGENTS.md, emphasizing implementations that minimize external dependencies while remaining fully unit-tested.
The entry point for this phase is phases/02-ml-fundamentals/README.md, which outlines the progression from basic supervised learning to advanced topics like anomaly detection.
Supervised Learning Foundations (Lessons 01–06)
The curriculum begins with conceptual framing in Lesson 01: What is Machine Learning?, located at phases/02-ml-fundamentals/01-what-is-machine-learning/docs/en.md, before diving into algorithmic implementations.
Linear Regression (Lesson 02) demonstrates closed-form solutions and gradient descent optimization. The implementation in phases/02-ml-fundamentals/02-linear-regression/code/main.py calculates the normal equation using NumPy:
import numpy as np
# Toy dataset
X = np.array([[1], [2], [3], [4]], dtype=float)
y = np.array([2, 4, 6, 8], dtype=float)
# Add bias term
X_b = np.c_[np.ones((len(X), 1)), X]
# Closed‑form solution: θ = (XᵀX)⁻¹Xᵀy
theta = np.linalg.inv(X_b.T @ X_b) @ X_b.T @ y
print("Intercept:", theta[0], "Slope:", theta[1])
Logistic Regression (Lesson 03) covers binary classification, sigmoid activation functions, and regularization strategies. Decision Trees (Lesson 04) implement impurity measures and recursive partitioning with pruning strategies, providing interpretable tree-based models.
Support Vector Machines (Lesson 05) explore margin maximization and kernel methods, while K-Nearest Neighbors & Distances (Lesson 06) focus on instance-based learning and computational optimizations for distance metrics.
Unsupervised Learning and Dimensionality Reduction (Lesson 07)
Lesson 07: Unsupervised Learning introduces clustering via k-means and dimensionality reduction through Principal Component Analysis (PCA). The lesson emphasizes evaluation metrics for clustering quality and variance retention, providing algorithms that function without labeled training data.
Feature Engineering and Selection (Lessons 08, 18)
The curriculum dedicates significant attention to data preparation. Lesson 08: Feature Engineering addresses scaling, encoding categorical variables, and creating interaction terms. Lesson 18: Feature Selection implements filter, wrapper, and embedded methods to reduce dimensionality, including mutual information calculations in phases/02-ml-fundamentals/18-feature-selection/code/main.py:
def mutual_info(x, y, bins=10):
# Discretize
x_bins = np.digitize(x, np.histogram_bin_edges(x, bins))
y_bins = np.digitize(y, np.histogram_bin_edges(y, bins))
# Joint histogram
pxy, _, _ = np.histogram2d(x_bins, y_bins, bins=bins, density=True)
px = pxy.sum(axis=1)
py = pxy.sum(axis=0)
# Compute MI
mi = (pxy * np.log(pxy / (px[:, None] * py[None, :] + 1e-9) + 1e-9)).sum()
return mi
Model Evaluation and Optimization (Lessons 09–13)
Critical production skills dominate the middle section of Phase 2. Lesson 09: Model Evaluation establishes metrics for regression (MSE, RMSE) and classification (accuracy, F1-score), alongside cross-validation strategies and confusion matrix analysis.
Lesson 10: Bias-Variance Trade-off provides theoretical foundations for underfitting and overfitting, utilizing learning curves to diagnose model capacity. Lesson 11: Ensemble Methods covers bagging (random forests), boosting implementations, and stacking architectures to improve predictive performance.
Lesson 12: Hyper-parameter Tuning implements grid search, random search, and Bayesian optimization basics. Lesson 13: ML Pipelines teaches the construction of reproducible workflows that chain preprocessing, model fitting, and validation steps.
Specialized ML Applications (Lessons 14–17)
The final lessons address domain-specific challenges. Naïve Bayes (Lesson 14) implements probabilistic classification with handling for both categorical and continuous features. Time-Series Forecasting (Lesson 15) covers ARIMA basics, seasonality detection, and temporal evaluation metrics located in phases/02-ml-fundamentals/15-time-series/docs/en.md.
Anomaly Detection (Lesson 16) explores statistical thresholds and model-based outlier identification. Imbalanced Data (Lesson 17) provides strategies including resampling techniques, class weighting adjustments, and precision-recall analysis for skewed distributions.
Decision Tree Implementation Example
The Decision Trees lesson demonstrates recursive partitioning using Gini impurity calculations. The reference implementation in phases/02-ml-fundamentals/04-decision-trees/code/main.py defines the tree structure as follows:
class Node:
def __init__(self, gini, samples, value, left=None, right=None):
self.gini = gini
self.samples = samples
self.value = value
self.left = left
self.right = right
def gini_impurity(y):
_, counts = np.unique(y, return_counts=True)
probs = counts / counts.sum()
return 1 - np.sum(probs ** 2)
# Very minimal recursive split (illustrative)
def grow_tree(X, y, depth=0, max_depth=3):
if depth == max_depth or len(np.unique(y)) == 1:
return Node(gini=gini_impurity(y), samples=len(y), value=np.bincount(y).argmax())
# split on first feature at median
thresh = np.median(X[:, 0])
left_idx = X[:, 0] <= thresh
right_idx = ~left_idx
left = grow_tree(X[left_idx], y[left_idx], depth + 1, max_depth)
right = grow_tree(X[right_idx], y[right_idx], depth + 1, max_depth)
return Node(gini=gini_impurity(y), samples=len(y), value=None, left=left, right=right)
Summary
Phase 2 of the ai-engineering-from-scratch curriculum provides:
- 18 sequential lessons covering supervised, unsupervised, and specialized ML techniques from
phases/02-ml-fundamentals/01-what-is-machine-learning/throughphases/02-ml-fundamentals/18-feature-selection/ - First-principles implementations in
phases/02-ml-fundamentals/*/code/main.<lang>without heavy framework dependencies - Comprehensive evaluation frameworks including bias-variance analysis, cross-validation strategies, and confusion matrix interpretations
- Production-ready pipelines that chain preprocessing, hyperparameter tuning, and model validation via
phases/02-ml-fundamentals/13-ml-pipelines/ - Validation through unit tests enforced by
scripts/audit_lessons.pyand concept mastery viaquiz.jsonfiles in each lesson directory
Frequently Asked Questions
What prerequisites are needed for Phase 2 ML fundamentals?
Learners should complete Phase 1 (Python/TypeScript basics and linear algebra fundamentals) before attempting Phase 2. The curriculum assumes familiarity with NumPy arrays and basic probability theory, though these concepts are reinforced through implementation practice in lessons like phases/02-ml-fundamentals/02-linear-regression/code/main.py.
How long does it take to complete all 18 lessons in Phase 2?
Each lesson requires approximately 3-5 hours of study time, including reading docs/en.md, implementing the algorithms in the code/ directory, and passing the 6-question quiz.json validation. The complete Phase 2 curriculum typically requires 2-3 weeks of dedicated study to master the ML fundamentals.
Are the implementations in Phase 2 production-ready or educational only?
While designed for educational clarity following the repository's AGENTS.md guidelines, the implementations adhere to production standards enforced by scripts/audit_lessons.py. Each lesson includes unit tests and self-terminating patterns that allow direct integration into ML pipelines, though production use should include additional error handling and scaling optimizations.
Which lesson covers cross-validation and model selection?
Lesson 09: Model Evaluation covers k-fold cross-validation, confusion matrices, and metric selection for both regression and classification tasks. The lesson files reside in phases/02-ml-fundamentals/09-model-evaluation/ and include implementations of stratified sampling techniques and learning curve analysis.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →