How to Cluster Similar Faces in a Large Image Collection Using face_recognition

You can cluster similar faces by extracting 128-dimensional embeddings with face_encodings, computing L2 distances via face_distance, and applying DBSCAN with an eps of approximately 0.6 to automatically group images by identity.

The face_recognition library converts every detected face into a compact numerical signature that lives in Euclidean space, making it ideal for unsupervised grouping algorithms. By leveraging these embeddings generated by the underlying dlib model, you can organize thousands of images into person-specific clusters without manual labeling or prior training data. This guide demonstrates the exact pipeline implemented in ageitgey/face_recognition to cluster similar faces at scale.

How Face Embeddings Enable Clustering

In face_recognition/api.py, the face_encodings function (lines 203–215) wraps the dlib face recognition model (dlib.face_recognition_model_v1) to produce a fixed-length, 128-dimensional vector for each face. Because these embeddings exist in a metric space, the L2 distance between two vectors directly correlates with facial similarity. The library exposes this calculation through face_distance (lines 63–71), which returns a NumPy array of distances suitable for any clustering algorithm that accepts pairwise metrics.

The Clustering Pipeline

Follow these four steps to transform a directory of photos into organized face clusters.

Step 1: Extract 128-Dimensional Encodings

Process each image to generate a single representative vector. Use load_image_file to read the image, then call face_encodings to obtain the embedding. If an image contains multiple faces, select the first encoding or filter by location depending on your use case.

import face_recognition
import numpy as np
import glob

image_paths = glob.glob("photos/**/*.jpg", recursive=True)
encodings = []
valid_paths = []

for path in image_paths:
    img = face_recognition.load_image_file(path)
    face_enc = face_recognition.face_encodings(img)
    if face_enc:                     # Skip images with no detectable face

        encodings.append(face_enc[0])
        valid_paths.append(path)

encodings = np.stack(encodings)   # Shape: (n_faces, 128)

Step 2: Apply DBSCAN Clustering

DBSCAN is the preferred algorithm for face clustering because it does not require pre-specifying the number of clusters. Set the eps parameter to approximately 0.6, which matches the default tolerance used internally by compare_faces for determining face matches. This threshold effectively separates distinct identities while grouping similar faces into the same cluster.

from sklearn.cluster import DBSCAN

clusterer = DBSCAN(metric="euclidean", eps=0.6, min_samples=1)
labels = clusterer.fit_predict(encodings)   # -1 indicates outliers

Step 3: Map Cluster Labels Back to Files

The algorithm returns an integer label for each face. Labels 0 to n represent distinct individuals, while -1 marks outliers (faces that do not fit any cluster). Map these labels back to your original file paths to organize images into folders.

clusters = {}
for label, path in zip(labels, valid_paths):
    clusters.setdefault(label, []).append(path)

for label, files in clusters.items():
    if label == -1:
        print(f"⚠️  {len(files)} images could not be assigned to any cluster")
    else:
        print(f"👤 Cluster {label}: {len(files)} images")
        for f in files[:3]:
            print(f"   - {f}")

Scaling to Large Collections

Processing thousands or millions of images requires optimizations to manage memory and compute efficiently.

GPU-Accelerated Batch Detection

Instead of processing images individually, use batch_face_locations (defined in face_recognition/api.py, lines 135–145) to run the CNN face detector on GPU across multiple images simultaneously. This significantly accelerates the bottleneck of face detection before encoding.

import glob
import face_recognition
from sklearn.cluster import DBSCAN

paths = glob.glob("large_collection/**/*.jpg", recursive=True)
images = [face_recognition.load_image_file(p) for p in paths]

# Process 64 images at once on GPU (requires CUDA-enabled dlib)

locations_batch = face_recognition.batch_face_locations(
    images, number_of_times_to_upsample=1, batch_size=64
)

# Extract encodings using the first detected face in each image

encodings = []
valid_paths = []
for img, locations, path in zip(images, locations_batch, paths):
    if locations:
        enc = face_recognition.face_encodings(
            img, known_face_locations=[locations[0]]
        )
        if enc:
            encodings.append(enc[0])
            valid_paths.append(path)

encodings = np.stack(encodings)
labels = DBSCAN(metric="euclidean", eps=0.6, min_samples=1).fit_predict(encodings)

Memory-Efficient Streaming

For collections exceeding available RAM, encode images in modest batches and write vectors to disk (CSV or NumPy .npy format) before clustering. This prevents loading all images into memory simultaneously while still allowing the final clustering step to operate on the full embedding matrix.

Parallel Encoding Across CPU Cores

The CPU-bound encoding process can be parallelized across cores using Python's multiprocessing module. Since face_encodings releases the GIL during computation, you can achieve near-linear speedup on multi-core machines by mapping the encoding function across a process pool.

Alternative Clustering Approaches

While DBSCAN excels at discovering unknown group sizes, other algorithms suit specific constraints.

K-Means is effective when you know the exact number of distinct people in your collection. Supply the expected count as n_clusters and fit the model directly to your encoding matrix.

from sklearn.cluster import KMeans

n_people = 25  # Known number of individuals

kmeans = KMeans(n_clusters=n_people, random_state=0).fit(encodings)
labels = kmeans.labels_

Agglomerative Clustering with Ward linkage can also be applied to face embeddings when you prefer hierarchical grouping or need to enforce a specific number of clusters post-hoc.

Key Source Files for Reference

Understanding these specific files in ageitgey/face_recognition helps customize your pipeline:

Summary

  • The face_recognition library generates 128-dimensional embeddings via face_encodings in api.py (lines 203–215), enabling mathematical comparison of facial similarity through L2 distance.
  • DBSCAN with eps=0.6 provides robust unsupervised clustering without requiring the number of people to be specified in advance, matching the tolerance used in compare_faces.
  • Use batch_face_locations (api.py lines 135–145) to leverage GPU acceleration when detecting faces in large datasets.
  • Stream encoding results to disk when working with millions of images to maintain constant memory usage.
  • Map resulting cluster labels back to original filenames to organize photos into person-specific groups.

Frequently Asked Questions

What distance metric should I use for clustering faces?

Use Euclidean (L2) distance as implemented by face_distance in face_recognition/api.py (lines 63–71). The 128-dimensional embeddings are optimized for this metric during the training of the underlying dlib model, making L2 distance directly proportional to facial dissimilarity.

How do I handle images with multiple faces?

For clustering purposes, extract one encoding per image or treat each detected face as a separate data point. When calling face_encodings, the function returns a list of vectors for all faces found. You can either select the first face (face_enc[0]) for single-person-per-image assumptions, or flatten all faces from all images into a single matrix and track parent filenames via a parallel list.

Can I cluster faces without knowing the number of people in advance?

Yes. DBSCAN is specifically designed for this scenario. By setting eps to roughly 0.6 (matching the tolerance in compare_faces) and min_samples=1, the algorithm automatically discovers the optimal number of clusters based on density in the embedding space, labeling sparse outliers as -1.

Why is my clustering producing too many small groups?

If DBSCAN creates excessive micro-clusters, increase the eps parameter slightly (e.g., to 0.7 or 0.8) to relax the similarity threshold. Conversely, if unrelated faces are merging into single clusters, decrease eps to enforce stricter matching criteria. The examples/face_recognition_knn.py file demonstrates techniques for tuning distance thresholds in similar applications.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →