How to Use a KNN Classifier for Face Recognition with Python

You can implement a KNN classifier for face recognition by training a scikit-learn KNeighborsClassifier on 128-dimensional face encodings extracted from labeled images, then predicting identities based on weighted distance to the nearest neighbors in the embedding space.

The face_recognition library provides a complete, production-ready K-Nearest Neighbors implementation that transforms face recognition into a classical machine learning classification problem. This approach leverages the library's pre-trained deep learning model to generate face encodings, then applies scikit-learn's efficient neighbor search algorithms to identify individuals without requiring GPU acceleration during inference.

How the KNN Face Recognition System Works

The system operates on face encodings—128-dimensional vectors generated by the library's underlying dlib model. Each encoding represents the unique facial geometry of a detected face in a compact, comparable format.

During training, the train() function in examples/face_recognition_knn.py (lines 46–100) walks your dataset directory, extracts encodings, and stores them as the feature matrix X with corresponding identity labels y. The classifier uses scikit-learn's KNeighborsClassifier with algorithm='ball_tree' and weights='distance' to enable efficient high-dimensional search and distance-weighted voting.

During prediction, the predict() function (lines 111–147) encodes faces in query images and queries the trained model for the k nearest neighbors. If the Euclidean distance to the closest match exceeds the distance_threshold (default 0.6), the face is labeled "unknown" to prevent false positives.

Preparing Your Training Dataset

The KNN implementation expects a specific directory structure where folder names serve as identity labels.

Create a root training directory with subfolders for each person:

training_dir/
├── alice/
│   ├── alice1.jpg
│   └── alice2.jpg
├── bob/
│   ├── bob1.jpg
│   └── bob2.jpg
└── charlie/
    └── charlie1.jpg

Critical constraint: Training images must contain exactly one detectable face. The train() function automatically skips images with zero or multiple faces to ensure unambiguous label assignment. For prediction, however, the system correctly handles images containing multiple faces.

Training the KNN Classifier

Import the training function from the example file and fit your model:

from face_recognition_knn import train

# Train the classifier

knn_clf = train(
    train_dir="training_dir",
    model_save_path="trained_knn_model.clf",  # Optional: persists via pickle

    n_neighbors=2,                            # Optional: auto-selects sqrt(n) if omitted

    verbose=True
)

Key implementation details from examples/face_recognition_knn.py:

  • If n_neighbors is omitted, the function calculates the square root of the training sample count as a statistically sound default.
  • The classifier uses algorithm='ball_tree' for efficient nearest neighbor search in the 128-dimensional encoding space.
  • Distance weighting (weights='distance') ensures closer neighbors exert more influence on predictions than distant ones.
  • The trained model is serialized using Python's pickle module when model_save_path is provided.

Making Predictions with the Trained Model

Use the predict() function to identify faces in new images:

from face_recognition_knn import predict

# Predict identities in a new image

predictions = predict(
    X_img_path="test_images/group_photo.jpg",
    model_path="trained_knn_model.clf",  # Load the saved model

    distance_threshold=0.6                 # Faces farther than this are "unknown"

)

# Display results: name and bounding box coordinates (top, right, bottom, left)

for name, (top, right, bottom, left) in predictions:
    print(f"Found {name} at location {(left, top, right, bottom)}")

Distance threshold mechanics:

The distance_threshold represents the maximum Euclidean distance in the 128-dimensional encoding space for a match to be considered valid. Lower values (e.g., 0.4) make the system more conservative, reducing false positives but potentially misclassifying known faces as "unknown." Higher values (e.g., 0.8) increase recall but risk false identifications. For security-critical applications, start with 0.5 and adjust based on your specific dataset's intra-class variance.

Complete End-to-End Example

Combine training and prediction into a single workflow:

from face_recognition_knn import train, predict
import os

# Configuration

TRAIN_DIR = "my_faces"
MODEL_FILE = "face_recognition_knn_model.clf"
TEST_IMAGE = "unknown_group.jpg"

# Step 1: Train the classifier

if not os.path.exists(MODEL_FILE):
    print("Training KNN classifier...")
    train(
        train_dir=TRAIN_DIR,
        model_save_path=MODEL_FILE,
        n_neighbors=2,
        verbose=True
    )

# Step 2: Predict faces in new image

print("Identifying faces...")
results = predict(
    X_img_path=TEST_IMAGE,
    model_path=MODEL_FILE,
    distance_threshold=0.6
)

for person, (top, right, bottom, left) in results:
    print(f"- {person} found at coordinates {(left, top, right, bottom)}")

Key Implementation Details

Understanding the underlying architecture helps optimize your face recognition pipeline:

  • Source location: The complete KNN logic resides in examples/face_recognition_knn.py (functions at lines 46–100 and 111–147). This file is not installed with the base package and must be copied from the repository into your project.

  • Dependencies: Requires scikit-learn in addition to the base face_recognition library. Install via pip install scikit-learn.

  • Algorithm choice: Uses algorithm='ball_tree' by default, offering superior performance for nearest neighbor search in high-dimensional spaces compared to brute-force methods.

  • Helper utilities: The script imports image_files_in_folder from face_recognition/face_recognition_cli.py to recursively discover image files during training.

  • Training constraints: Images must contain exactly one face. The train() function filters out images with zero or multiple detections to prevent ambiguous label assignments.

Summary

  • The face_recognition repository provides a complete KNN classifier example in examples/face_recognition_knn.py that wraps scikit-learn's KNeighborsClassifier around 128-dimensional face encodings.

  • Training requires a directory structure where subfolder names serve as identity labels, with each training image containing exactly one face. The train() function automatically selects optimal hyperparameters and persists models using pickle.

  • Prediction uses the predict() function to load saved models, detect multiple faces in query images, and return identity labels with bounding box coordinates. Faces exceeding the distance threshold are labeled "unknown".

  • The implementation relies on scikit-learn and uses ball-tree algorithm with distance weighting for efficient, accurate nearest neighbor search in high-dimensional encoding space.

Frequently Asked Questions

What is the optimal value for n_neighbors in face recognition KNN?

If you omit the n_neighbors parameter, the train() function automatically calculates the square root of the number of training samples, providing a statistically sound baseline. For small datasets with fewer than 20 samples per person, use n_neighbors=1 to avoid over-smoothing, while larger datasets benefit from values between 3 and 5 to reduce noise sensitivity without sacrificing accuracy.

How does the distance threshold affect recognition accuracy?

The distance_threshold parameter (default 0.6) acts as a confidence gate in the 128-dimensional encoding space. Lower values (e.g., 0.4) make the system more conservative, reducing false positives but potentially misclassifying known faces as "unknown." Higher values (e.g., 0.8) increase recall but risk false identifications. For security-critical applications, start with 0.5 and adjust based on your specific dataset's intra-class variance.

Can I use images with multiple faces for training the KNN classifier?

No. The train() function in examples/face_recognition_knn.py explicitly skips any training image that does not contain exactly one face. This constraint ensures unambiguous label assignment—since the folder name becomes the label, the system cannot determine which face in a multi-face image corresponds to that identity. For training, ensure each image contains only the target person; for prediction, the predict() function correctly handles images containing multiple faces.

What dependencies are required to run the KNN face recognition example?

The KNN implementation requires scikit-learn in addition to the base face_recognition library and its dependencies (dlib, numpy, Pillow). Install the required package via pip install scikit-learn. The example also uses Python's built-in pickle module for model serialization and os/pathlib for file system operations, which require no additional installation. Note that the face_recognition_knn.py example file itself must be obtained from the repository's examples/ directory, as it is not included in the standard package installation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →