# Effective Feature Engineering for Machine Learning: From-Scratch Implementations

> Learn effective feature engineering for machine learning from scratch. Understand scaling, encoding, and text vectorization with pure Python implementations. Master ML transformations.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: tutorial
- Published: 2026-07-19

---

**Mastering feature engineering for machine learning requires understanding the mathematical foundations of transformations like scaling, encoding, and text vectorization before applying production libraries.** The open-source curriculum **AI Engineering From Scratch** teaches these critical skills through pure Python implementations that mirror scikit-learn's internals.

The repository `rohitg00/ai-engineering-from-scratch` organizes lessons into self-contained modules where each concept is built from first principles. The **Feature Engineering & Selection** lesson (Phase 2, Lesson 08) contained in `phases/02-ml-fundamentals/08-feature-engineering/` demonstrates how raw data becomes model-ready through transparent, inspectable code rather than black-box abstractions.

## Why Feature Engineering Determines Model Performance

In classical machine learning, algorithm sophistication rarely compensates for poor feature quality. The lesson documentation in [`phases/02-ml-fundamentals/08-feature-engineering/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/08-feature-engineering/docs/en.md) establishes that **feature engineering**—the process of transforming raw inputs into numerical representations—directly controls model accuracy, overfitting risk, and pipeline debuggability.

The curriculum follows a **"Build It / Use It"** split: students first implement algorithms manually to learn underlying mathematics, then compare their code to production libraries like scikit-learn. This approach guarantees that practitioners understand exactly what happens when `fit()` and `transform()` execute.

## Core Numerical Transformations

Numerical features often require scaling to prevent magnitude dominance or to normalize distributions. The reference implementation in [`phases/02-ml-fundamentals/08-feature-engineering/code/features.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/08-feature-engineering/code/features.py) provides from-scratch versions of essential transforms.

### Min-Max Scaling and Standardization

**Min-max scaling** compresses values to a [0, 1] range by subtracting the minimum and dividing by the range:

```python
def min_max_scale(values):
    min_val = min(values)
    max_val = max(values)
    if max_val == min_val:
        return [0.0] * len(values)
    return [(v - min_val) / (max_val - min_val) for v in values]

```

**Standardization** (z-score normalization) centers data around zero with unit variance, calculated by dividing mean-centered values by the standard deviation:

```python
def standardize(values):
    n = len(values)
    mean = sum(values) / n
    variance = sum((v - mean) ** 2 for v in values) / n
    std = math.sqrt(variance) if variance > 0 else 1.0
    return [(v - mean) / std for v in values]

```

The lesson also covers **log transforms** for skewed distributions, **binning** for continuous-to-categorical conversion, and **polynomial features** for capturing non-linear relationships.

## Encoding Categorical Variables

Machine learning models require numerical inputs, necessitating encoding strategies for categorical data. The `one_hot_encode()` function in [`features.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/features.py) demonstrates how to create binary indicator matrices without pandas or scikit-learn:

```python
def one_hot_encode(values):
    categories = sorted(set(values))
    cat_to_idx = {cat: i for i, cat in enumerate(categories)}
    encoded = []
    for v in values:
        row = [0] * len(categories)
        row[cat_to_idx[v]] = 1
        encoded.append(row)
    return encoded, categories

```

The curriculum also implements **label encoding** (ordinal integer mapping) and **target encoding** (replacing categories with mean target values using smoothing to prevent data leakage).

## Text Feature Extraction: TF-IDF From Scratch

Text data requires specialized feature engineering to convert unstructured strings into mathematical vectors. The lesson implements **TF-IDF** (Term Frequency-Inverse Document Frequency) without external NLP libraries:

```python
def tfidf(documents):
    n_docs = len(documents)
    vocab = {}
    for doc in documents:
        for word in doc.lower().split():
            if word not in vocab:
                vocab[word] = len(vocab)

    doc_freq = {}
    for doc in documents:
        seen = set()
        for word in doc.lower().split():
            if word not in seen:
                doc_freq[word] = doc_freq.get(word, 0) + 1
                seen.add(word)

    vectors = []
    for doc in documents:
        words = doc.lower().split()
        word_count = len(words)
        tf_map = {}
        for word in words:
            tf_map[word] = tf_map.get(word, 0) + 1

        vec = [0.0] * len(vocab)
        for word, count in tf_map.items():
            tf = count / word_count
            idf = math.log(n_docs / doc_freq[word])
            vec[vocab[word]] = tf * idf
        vectors.append(vec)
    return vectors, vocab

```

This implementation calculates term frequency (word count normalized by document length) and inverse document frequency (logarithm of total documents divided by documents containing the term), mirroring scikit-learn's `TfidfVectorizer` behavior exactly.

## Feature Selection Strategies

Dimensionality reduction through **feature selection** prevents overfitting and reduces training time. The repository demonstrates **filter methods** that rank features by statistical metrics without training a model.

The `variance_threshold()` function removes low-variance features that provide minimal discriminative power:

```python
def variance_threshold(features, threshold=0.01):
    n_features = len(features[0])
    selected = []
    for j in range(n_features):
        col = [features[i][j] for i in range(len(features))]
        var = sum((v - sum(col)/len(col))**2 for v in col) / len(col)
        if var >= threshold:
            selected.append(j)
    return selected

```

The lesson also covers **correlation filtering** (removing redundant features) and **mutual information** (measuring dependency between variables and targets), plus an introduction to **wrapper methods** like L1 regularization that embed selection within model training.

## From Educational Code to Production Artifacts

Each lesson in the repository generates a reusable artifact in the `outputs/` directory. For feature engineering, this artifact is [`phases/02-ml-fundamentals/08-feature-engineering/outputs/prompt-feature-engineer.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/08-feature-engineering/outputs/prompt-feature-engineer.md)—a structured prompt that can be fed to LLMs to generate systematic feature-engineering suggestions for new datasets.

This bridges the gap between educational implementations and production workflows: students understand the mathematics through the [`features.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/features.py) code, then apply that knowledge using either the validated algorithms or the automated prompts for rapid iteration.

## Summary

- **Feature engineering for machine learning** encompasses numerical scaling, categorical encoding, text vectorization, and statistical selection.
- The `rohitg00/ai-engineering-from-scratch` repository implements these transformations in [`phases/02-ml-fundamentals/08-feature-engineering/code/features.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/02-ml-fundamentals/08-feature-engineering/code/features.py) using pure Python without library dependencies.
- Key implementations include `min_max_scale()`, `standardize()`, `one_hot_encode()`, `tfidf()`, and `variance_threshold()`.
- Understanding the mathematical operations behind scikit-learn's transformers enables better debugging and custom pipeline development.
- The curriculum's "Build It / Use It" approach and reusable prompt artifacts facilitate both learning and production application.

## Frequently Asked Questions

### What is the difference between min-max scaling and standardization?

**Min-max scaling** transforms values to a fixed range (typically [0, 1]) by subtracting the minimum and dividing by the range, making it sensitive to outliers but preserving zero sparsity. **Standardization** centers data at zero with unit variance using the mean and standard deviation, which handles outliers better but does not bound the range. According to the `ai-engineering-from-scratch` source code, min-max is preferred for bounded data like pixel intensities, while standardization is optimal for normally distributed features or when using algorithms assuming centered data like PCA or SVM.

### How does TF-IDF improve upon simple bag-of-words counting?

**TF-IDF** (Term Frequency-Inverse Document Frequency) improves upon raw count vectors by downweighting common words that appear across many documents (like "the" or "and"), which the `tfidf()` implementation calculates using the formula `log(n_docs / doc_freq[word])`. This ensures that rare, discriminative terms receive higher weights than frequent but uninformative words. The from-scratch implementation in [`features.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/features.py) mirrors scikit-learn's approach by computing term frequency (normalized by document length) multiplied by the inverse document frequency.

### When should I use one-hot encoding versus target encoding?

**One-hot encoding** (implemented in `one_hot_encode()`) creates binary columns for each category and is safest when there is no ordinal relationship between categories and sufficient memory exists for the expanded feature space. **Target encoding** replaces categories with the mean target value and is preferable for high-cardinality features (like zip codes or device IDs), though it requires smoothing and cross-validation to prevent overfitting. The repository teaches that one-hot encoding risks the "curse of dimensionality" with many categories, while target encoding risks data leakage if not implemented carefully.

### What is variance thresholding in feature selection?

**Variance thresholding** (implemented in `variance_threshold(features, threshold)`) removes features with variance below a specified cutoff, under the principle that constant or near-constant values provide no discriminative power for prediction. The function calculates population variance for each feature column and returns indices of those meeting the threshold. This unsupervised filter method is computationally efficient and serves as a preprocessing step before more expensive wrapper or embedded selection methods.