How to Cluster Similar Faces in a Large Image Collection Using face_recognition
You can cluster similar faces by extracting 128-dimensional embeddings with face_encodings, computing L2 distances via face_distance, and applying DBSCAN with an eps of approximately 0.6 to automatically group images by identity.
The face_recognition library converts every detected face into a compact numerical signature that lives in Euclidean space, making it ideal for unsupervised grouping algorithms. By leveraging these embeddings generated by the underlying dlib model, you can organize thousands of images into person-specific clusters without manual labeling or prior training data. This guide demonstrates the exact pipeline implemented in ageitgey/face_recognition to cluster similar faces at scale.
How Face Embeddings Enable Clustering
In face_recognition/api.py, the face_encodings function (lines 203–215) wraps the dlib face recognition model (dlib.face_recognition_model_v1) to produce a fixed-length, 128-dimensional vector for each face. Because these embeddings exist in a metric space, the L2 distance between two vectors directly correlates with facial similarity. The library exposes this calculation through face_distance (lines 63–71), which returns a NumPy array of distances suitable for any clustering algorithm that accepts pairwise metrics.
The Clustering Pipeline
Follow these four steps to transform a directory of photos into organized face clusters.
Step 1: Extract 128-Dimensional Encodings
Process each image to generate a single representative vector. Use load_image_file to read the image, then call face_encodings to obtain the embedding. If an image contains multiple faces, select the first encoding or filter by location depending on your use case.
import face_recognition
import numpy as np
import glob
image_paths = glob.glob("photos/**/*.jpg", recursive=True)
encodings = []
valid_paths = []
for path in image_paths:
img = face_recognition.load_image_file(path)
face_enc = face_recognition.face_encodings(img)
if face_enc: # Skip images with no detectable face
encodings.append(face_enc[0])
valid_paths.append(path)
encodings = np.stack(encodings) # Shape: (n_faces, 128)
Step 2: Apply DBSCAN Clustering
DBSCAN is the preferred algorithm for face clustering because it does not require pre-specifying the number of clusters. Set the eps parameter to approximately 0.6, which matches the default tolerance used internally by compare_faces for determining face matches. This threshold effectively separates distinct identities while grouping similar faces into the same cluster.
from sklearn.cluster import DBSCAN
clusterer = DBSCAN(metric="euclidean", eps=0.6, min_samples=1)
labels = clusterer.fit_predict(encodings) # -1 indicates outliers
Step 3: Map Cluster Labels Back to Files
The algorithm returns an integer label for each face. Labels 0 to n represent distinct individuals, while -1 marks outliers (faces that do not fit any cluster). Map these labels back to your original file paths to organize images into folders.
clusters = {}
for label, path in zip(labels, valid_paths):
clusters.setdefault(label, []).append(path)
for label, files in clusters.items():
if label == -1:
print(f"⚠️ {len(files)} images could not be assigned to any cluster")
else:
print(f"👤 Cluster {label}: {len(files)} images")
for f in files[:3]:
print(f" - {f}")
Scaling to Large Collections
Processing thousands or millions of images requires optimizations to manage memory and compute efficiently.
GPU-Accelerated Batch Detection
Instead of processing images individually, use batch_face_locations (defined in face_recognition/api.py, lines 135–145) to run the CNN face detector on GPU across multiple images simultaneously. This significantly accelerates the bottleneck of face detection before encoding.
import glob
import face_recognition
from sklearn.cluster import DBSCAN
paths = glob.glob("large_collection/**/*.jpg", recursive=True)
images = [face_recognition.load_image_file(p) for p in paths]
# Process 64 images at once on GPU (requires CUDA-enabled dlib)
locations_batch = face_recognition.batch_face_locations(
images, number_of_times_to_upsample=1, batch_size=64
)
# Extract encodings using the first detected face in each image
encodings = []
valid_paths = []
for img, locations, path in zip(images, locations_batch, paths):
if locations:
enc = face_recognition.face_encodings(
img, known_face_locations=[locations[0]]
)
if enc:
encodings.append(enc[0])
valid_paths.append(path)
encodings = np.stack(encodings)
labels = DBSCAN(metric="euclidean", eps=0.6, min_samples=1).fit_predict(encodings)
Memory-Efficient Streaming
For collections exceeding available RAM, encode images in modest batches and write vectors to disk (CSV or NumPy .npy format) before clustering. This prevents loading all images into memory simultaneously while still allowing the final clustering step to operate on the full embedding matrix.
Parallel Encoding Across CPU Cores
The CPU-bound encoding process can be parallelized across cores using Python's multiprocessing module. Since face_encodings releases the GIL during computation, you can achieve near-linear speedup on multi-core machines by mapping the encoding function across a process pool.
Alternative Clustering Approaches
While DBSCAN excels at discovering unknown group sizes, other algorithms suit specific constraints.
K-Means is effective when you know the exact number of distinct people in your collection. Supply the expected count as n_clusters and fit the model directly to your encoding matrix.
from sklearn.cluster import KMeans
n_people = 25 # Known number of individuals
kmeans = KMeans(n_clusters=n_people, random_state=0).fit(encodings)
labels = kmeans.labels_
Agglomerative Clustering with Ward linkage can also be applied to face embeddings when you prefer hierarchical grouping or need to enforce a specific number of clusters post-hoc.
Key Source Files for Reference
Understanding these specific files in ageitgey/face_recognition helps customize your pipeline:
face_recognition/api.py– Contains the core public API includingface_encodings,face_distance, andbatch_face_locations.examples/face_recognition_knn.py– Demonstrates building a K-Nearest Neighbors model over encodings; useful for extending clustering to classification tasks.examples/face_recognition_svm.py– Shows supervised training on face embeddings when you have labeled data instead of purely unsupervised clusters.examples/find_faces_in_batches.py– Illustrates efficient batch processing patterns essential for large-scale implementations.
Summary
- The
face_recognitionlibrary generates 128-dimensional embeddings viaface_encodingsinapi.py(lines 203–215), enabling mathematical comparison of facial similarity through L2 distance. - DBSCAN with
eps=0.6provides robust unsupervised clustering without requiring the number of people to be specified in advance, matching the tolerance used incompare_faces. - Use
batch_face_locations(api.py lines 135–145) to leverage GPU acceleration when detecting faces in large datasets. - Stream encoding results to disk when working with millions of images to maintain constant memory usage.
- Map resulting cluster labels back to original filenames to organize photos into person-specific groups.
Frequently Asked Questions
What distance metric should I use for clustering faces?
Use Euclidean (L2) distance as implemented by face_distance in face_recognition/api.py (lines 63–71). The 128-dimensional embeddings are optimized for this metric during the training of the underlying dlib model, making L2 distance directly proportional to facial dissimilarity.
How do I handle images with multiple faces?
For clustering purposes, extract one encoding per image or treat each detected face as a separate data point. When calling face_encodings, the function returns a list of vectors for all faces found. You can either select the first face (face_enc[0]) for single-person-per-image assumptions, or flatten all faces from all images into a single matrix and track parent filenames via a parallel list.
Can I cluster faces without knowing the number of people in advance?
Yes. DBSCAN is specifically designed for this scenario. By setting eps to roughly 0.6 (matching the tolerance in compare_faces) and min_samples=1, the algorithm automatically discovers the optimal number of clusters based on density in the embedding space, labeling sparse outliers as -1.
Why is my clustering producing too many small groups?
If DBSCAN creates excessive micro-clusters, increase the eps parameter slightly (e.g., to 0.7 or 0.8) to relax the similarity threshold. Conversely, if unrelated faces are merging into single clusters, decrease eps to enforce stricter matching criteria. The examples/face_recognition_knn.py file demonstrates techniques for tuning distance thresholds in similar applications.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →