Effective Feature Engineering for Machine Learning: From-Scratch Implementations

Mastering feature engineering for machine learning requires understanding the mathematical foundations of transformations like scaling, encoding, and text vectorization before applying production libraries. The open-source curriculum AI Engineering From Scratch teaches these critical skills through pure Python implementations that mirror scikit-learn's internals.

The repository rohitg00/ai-engineering-from-scratch organizes lessons into self-contained modules where each concept is built from first principles. The Feature Engineering & Selection lesson (Phase 2, Lesson 08) contained in phases/02-ml-fundamentals/08-feature-engineering/ demonstrates how raw data becomes model-ready through transparent, inspectable code rather than black-box abstractions.

Why Feature Engineering Determines Model Performance

In classical machine learning, algorithm sophistication rarely compensates for poor feature quality. The lesson documentation in phases/02-ml-fundamentals/08-feature-engineering/docs/en.md establishes that feature engineering—the process of transforming raw inputs into numerical representations—directly controls model accuracy, overfitting risk, and pipeline debuggability.

The curriculum follows a "Build It / Use It" split: students first implement algorithms manually to learn underlying mathematics, then compare their code to production libraries like scikit-learn. This approach guarantees that practitioners understand exactly what happens when fit() and transform() execute.

Core Numerical Transformations

Numerical features often require scaling to prevent magnitude dominance or to normalize distributions. The reference implementation in phases/02-ml-fundamentals/08-feature-engineering/code/features.py provides from-scratch versions of essential transforms.

Min-Max Scaling and Standardization

Min-max scaling compresses values to a [0, 1] range by subtracting the minimum and dividing by the range:

def min_max_scale(values):
    min_val = min(values)
    max_val = max(values)
    if max_val == min_val:
        return [0.0] * len(values)
    return [(v - min_val) / (max_val - min_val) for v in values]

Standardization (z-score normalization) centers data around zero with unit variance, calculated by dividing mean-centered values by the standard deviation:

def standardize(values):
    n = len(values)
    mean = sum(values) / n
    variance = sum((v - mean) ** 2 for v in values) / n
    std = math.sqrt(variance) if variance > 0 else 1.0
    return [(v - mean) / std for v in values]

The lesson also covers log transforms for skewed distributions, binning for continuous-to-categorical conversion, and polynomial features for capturing non-linear relationships.

Encoding Categorical Variables

Machine learning models require numerical inputs, necessitating encoding strategies for categorical data. The one_hot_encode() function in features.py demonstrates how to create binary indicator matrices without pandas or scikit-learn:

def one_hot_encode(values):
    categories = sorted(set(values))
    cat_to_idx = {cat: i for i, cat in enumerate(categories)}
    encoded = []
    for v in values:
        row = [0] * len(categories)
        row[cat_to_idx[v]] = 1
        encoded.append(row)
    return encoded, categories

The curriculum also implements label encoding (ordinal integer mapping) and target encoding (replacing categories with mean target values using smoothing to prevent data leakage).

Text Feature Extraction: TF-IDF From Scratch

Text data requires specialized feature engineering to convert unstructured strings into mathematical vectors. The lesson implements TF-IDF (Term Frequency-Inverse Document Frequency) without external NLP libraries:

def tfidf(documents):
    n_docs = len(documents)
    vocab = {}
    for doc in documents:
        for word in doc.lower().split():
            if word not in vocab:
                vocab[word] = len(vocab)

    doc_freq = {}
    for doc in documents:
        seen = set()
        for word in doc.lower().split():
            if word not in seen:
                doc_freq[word] = doc_freq.get(word, 0) + 1
                seen.add(word)

    vectors = []
    for doc in documents:
        words = doc.lower().split()
        word_count = len(words)
        tf_map = {}
        for word in words:
            tf_map[word] = tf_map.get(word, 0) + 1

        vec = [0.0] * len(vocab)
        for word, count in tf_map.items():
            tf = count / word_count
            idf = math.log(n_docs / doc_freq[word])
            vec[vocab[word]] = tf * idf
        vectors.append(vec)
    return vectors, vocab

This implementation calculates term frequency (word count normalized by document length) and inverse document frequency (logarithm of total documents divided by documents containing the term), mirroring scikit-learn's TfidfVectorizer behavior exactly.

Feature Selection Strategies

Dimensionality reduction through feature selection prevents overfitting and reduces training time. The repository demonstrates filter methods that rank features by statistical metrics without training a model.

The variance_threshold() function removes low-variance features that provide minimal discriminative power:

def variance_threshold(features, threshold=0.01):
    n_features = len(features[0])
    selected = []
    for j in range(n_features):
        col = [features[i][j] for i in range(len(features))]
        var = sum((v - sum(col)/len(col))**2 for v in col) / len(col)
        if var >= threshold:
            selected.append(j)
    return selected

The lesson also covers correlation filtering (removing redundant features) and mutual information (measuring dependency between variables and targets), plus an introduction to wrapper methods like L1 regularization that embed selection within model training.

From Educational Code to Production Artifacts

Each lesson in the repository generates a reusable artifact in the outputs/ directory. For feature engineering, this artifact is phases/02-ml-fundamentals/08-feature-engineering/outputs/prompt-feature-engineer.md—a structured prompt that can be fed to LLMs to generate systematic feature-engineering suggestions for new datasets.

This bridges the gap between educational implementations and production workflows: students understand the mathematics through the features.py code, then apply that knowledge using either the validated algorithms or the automated prompts for rapid iteration.

Summary

  • Feature engineering for machine learning encompasses numerical scaling, categorical encoding, text vectorization, and statistical selection.
  • The rohitg00/ai-engineering-from-scratch repository implements these transformations in phases/02-ml-fundamentals/08-feature-engineering/code/features.py using pure Python without library dependencies.
  • Key implementations include min_max_scale(), standardize(), one_hot_encode(), tfidf(), and variance_threshold().
  • Understanding the mathematical operations behind scikit-learn's transformers enables better debugging and custom pipeline development.
  • The curriculum's "Build It / Use It" approach and reusable prompt artifacts facilitate both learning and production application.

Frequently Asked Questions

What is the difference between min-max scaling and standardization?

Min-max scaling transforms values to a fixed range (typically [0, 1]) by subtracting the minimum and dividing by the range, making it sensitive to outliers but preserving zero sparsity. Standardization centers data at zero with unit variance using the mean and standard deviation, which handles outliers better but does not bound the range. According to the ai-engineering-from-scratch source code, min-max is preferred for bounded data like pixel intensities, while standardization is optimal for normally distributed features or when using algorithms assuming centered data like PCA or SVM.

How does TF-IDF improve upon simple bag-of-words counting?

TF-IDF (Term Frequency-Inverse Document Frequency) improves upon raw count vectors by downweighting common words that appear across many documents (like "the" or "and"), which the tfidf() implementation calculates using the formula log(n_docs / doc_freq[word]). This ensures that rare, discriminative terms receive higher weights than frequent but uninformative words. The from-scratch implementation in features.py mirrors scikit-learn's approach by computing term frequency (normalized by document length) multiplied by the inverse document frequency.

When should I use one-hot encoding versus target encoding?

One-hot encoding (implemented in one_hot_encode()) creates binary columns for each category and is safest when there is no ordinal relationship between categories and sufficient memory exists for the expanded feature space. Target encoding replaces categories with the mean target value and is preferable for high-cardinality features (like zip codes or device IDs), though it requires smoothing and cross-validation to prevent overfitting. The repository teaches that one-hot encoding risks the "curse of dimensionality" with many categories, while target encoding risks data leakage if not implemented carefully.

What is variance thresholding in feature selection?

Variance thresholding (implemented in variance_threshold(features, threshold)) removes features with variance below a specified cutoff, under the principle that constant or near-constant values provide no discriminative power for prediction. The function calculates population variance for each feature column and returns indices of those meeting the threshold. This unsupervised filter method is computationally efficient and serves as a preprocessing step before more expensive wrapper or embedded selection methods.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →