# How to Cluster Similar Faces in a Large Image Collection Using face_recognition

> Cluster similar faces in large image collections using face_recognition. Extract embeddings, compute distances, and apply DBSCAN for automatic identity grouping. Learn how now.

- Repository: [Adam Geitgey/face_recognition](https://github.com/ageitgey/face_recognition)
- Tags: how-to-guide
- Published: 2026-03-06

---

**You can cluster similar faces by extracting 128-dimensional embeddings with `face_encodings`, computing L2 distances via `face_distance`, and applying DBSCAN with an `eps` of approximately 0.6 to automatically group images by identity.**

The `face_recognition` library converts every detected face into a compact numerical signature that lives in Euclidean space, making it ideal for unsupervised grouping algorithms. By leveraging these embeddings generated by the underlying dlib model, you can organize thousands of images into person-specific clusters without manual labeling or prior training data. This guide demonstrates the exact pipeline implemented in `ageitgey/face_recognition` to cluster similar faces at scale.

## How Face Embeddings Enable Clustering

In [`face_recognition/api.py`](https://github.com/ageitgey/face_recognition/blob/main/face_recognition/api.py), the `face_encodings` function (lines 203–215) wraps the **dlib face recognition model** (`dlib.face_recognition_model_v1`) to produce a fixed-length, 128-dimensional vector for each face. Because these embeddings exist in a metric space, the **L2 distance** between two vectors directly correlates with facial similarity. The library exposes this calculation through `face_distance` (lines 63–71), which returns a NumPy array of distances suitable for any clustering algorithm that accepts pairwise metrics.

## The Clustering Pipeline

Follow these four steps to transform a directory of photos into organized face clusters.

### Step 1: Extract 128-Dimensional Encodings

Process each image to generate a single representative vector. Use `load_image_file` to read the image, then call `face_encodings` to obtain the embedding. If an image contains multiple faces, select the first encoding or filter by location depending on your use case.

```python
import face_recognition
import numpy as np
import glob

image_paths = glob.glob("photos/**/*.jpg", recursive=True)
encodings = []
valid_paths = []

for path in image_paths:
    img = face_recognition.load_image_file(path)
    face_enc = face_recognition.face_encodings(img)
    if face_enc:                     # Skip images with no detectable face

        encodings.append(face_enc[0])
        valid_paths.append(path)

encodings = np.stack(encodings)   # Shape: (n_faces, 128)

```

### Step 2: Apply DBSCAN Clustering

**DBSCAN** is the preferred algorithm for face clustering because it does not require pre-specifying the number of clusters. Set the `eps` parameter to approximately **0.6**, which matches the default tolerance used internally by `compare_faces` for determining face matches. This threshold effectively separates distinct identities while grouping similar faces into the same cluster.

```python
from sklearn.cluster import DBSCAN

clusterer = DBSCAN(metric="euclidean", eps=0.6, min_samples=1)
labels = clusterer.fit_predict(encodings)   # -1 indicates outliers

```

### Step 3: Map Cluster Labels Back to Files

The algorithm returns an integer label for each face. Labels `0` to `n` represent distinct individuals, while `-1` marks outliers (faces that do not fit any cluster). Map these labels back to your original file paths to organize images into folders.

```python
clusters = {}
for label, path in zip(labels, valid_paths):
    clusters.setdefault(label, []).append(path)

for label, files in clusters.items():
    if label == -1:
        print(f"⚠️  {len(files)} images could not be assigned to any cluster")
    else:
        print(f"👤 Cluster {label}: {len(files)} images")
        for f in files[:3]:
            print(f"   - {f}")

```

## Scaling to Large Collections

Processing thousands or millions of images requires optimizations to manage memory and compute efficiently.

### GPU-Accelerated Batch Detection

Instead of processing images individually, use `batch_face_locations` (defined in [`face_recognition/api.py`](https://github.com/ageitgey/face_recognition/blob/main/face_recognition/api.py), lines 135–145) to run the CNN face detector on GPU across multiple images simultaneously. This significantly accelerates the bottleneck of face detection before encoding.

```python
import glob
import face_recognition
from sklearn.cluster import DBSCAN

paths = glob.glob("large_collection/**/*.jpg", recursive=True)
images = [face_recognition.load_image_file(p) for p in paths]

# Process 64 images at once on GPU (requires CUDA-enabled dlib)

locations_batch = face_recognition.batch_face_locations(
    images, number_of_times_to_upsample=1, batch_size=64
)

# Extract encodings using the first detected face in each image

encodings = []
valid_paths = []
for img, locations, path in zip(images, locations_batch, paths):
    if locations:
        enc = face_recognition.face_encodings(
            img, known_face_locations=[locations[0]]
        )
        if enc:
            encodings.append(enc[0])
            valid_paths.append(path)

encodings = np.stack(encodings)
labels = DBSCAN(metric="euclidean", eps=0.6, min_samples=1).fit_predict(encodings)

```

### Memory-Efficient Streaming

For collections exceeding available RAM, encode images in modest batches and write vectors to disk (CSV or NumPy `.npy` format) before clustering. This prevents loading all images into memory simultaneously while still allowing the final clustering step to operate on the full embedding matrix.

### Parallel Encoding Across CPU Cores

The CPU-bound encoding process can be parallelized across cores using Python's `multiprocessing` module. Since `face_encodings` releases the GIL during computation, you can achieve near-linear speedup on multi-core machines by mapping the encoding function across a process pool.

## Alternative Clustering Approaches

While DBSCAN excels at discovering unknown group sizes, other algorithms suit specific constraints.

**K-Means** is effective when you know the exact number of distinct people in your collection. Supply the expected count as `n_clusters` and fit the model directly to your encoding matrix.

```python
from sklearn.cluster import KMeans

n_people = 25  # Known number of individuals

kmeans = KMeans(n_clusters=n_people, random_state=0).fit(encodings)
labels = kmeans.labels_

```

**Agglomerative Clustering** with Ward linkage can also be applied to face embeddings when you prefer hierarchical grouping or need to enforce a specific number of clusters post-hoc.

## Key Source Files for Reference

Understanding these specific files in `ageitgey/face_recognition` helps customize your pipeline:

- [`face_recognition/api.py`](https://github.com/ageitgey/face_recognition/blob/main/face_recognition/api.py) – Contains the core public API including `face_encodings`, `face_distance`, and `batch_face_locations`.
- [`examples/face_recognition_knn.py`](https://github.com/ageitgey/face_recognition/blob/main/examples/face_recognition_knn.py) – Demonstrates building a K-Nearest Neighbors model over encodings; useful for extending clustering to classification tasks.
- [`examples/face_recognition_svm.py`](https://github.com/ageitgey/face_recognition/blob/main/examples/face_recognition_svm.py) – Shows supervised training on face embeddings when you have labeled data instead of purely unsupervised clusters.
- [`examples/find_faces_in_batches.py`](https://github.com/ageitgey/face_recognition/blob/main/examples/find_faces_in_batches.py) – Illustrates efficient batch processing patterns essential for large-scale implementations.

## Summary

- The `face_recognition` library generates **128-dimensional embeddings** via `face_encodings` in [`api.py`](https://github.com/ageitgey/face_recognition/blob/main/api.py) (lines 203–215), enabling mathematical comparison of facial similarity through L2 distance.
- **DBSCAN** with `eps=0.6` provides robust unsupervised clustering without requiring the number of people to be specified in advance, matching the tolerance used in `compare_faces`.
- Use `batch_face_locations` (api.py lines 135–145) to leverage GPU acceleration when detecting faces in large datasets.
- Stream encoding results to disk when working with millions of images to maintain constant memory usage.
- Map resulting cluster labels back to original filenames to organize photos into person-specific groups.

## Frequently Asked Questions

### What distance metric should I use for clustering faces?

Use **Euclidean (L2) distance** as implemented by `face_distance` in [`face_recognition/api.py`](https://github.com/ageitgey/face_recognition/blob/main/face_recognition/api.py) (lines 63–71). The 128-dimensional embeddings are optimized for this metric during the training of the underlying dlib model, making L2 distance directly proportional to facial dissimilarity.

### How do I handle images with multiple faces?

For clustering purposes, extract one encoding per image or treat each detected face as a separate data point. When calling `face_encodings`, the function returns a list of vectors for all faces found. You can either select the first face (`face_enc[0]`) for single-person-per-image assumptions, or flatten all faces from all images into a single matrix and track parent filenames via a parallel list.

### Can I cluster faces without knowing the number of people in advance?

Yes. **DBSCAN** is specifically designed for this scenario. By setting `eps` to roughly 0.6 (matching the tolerance in `compare_faces`) and `min_samples=1`, the algorithm automatically discovers the optimal number of clusters based on density in the embedding space, labeling sparse outliers as `-1`.

### Why is my clustering producing too many small groups?

If DBSCAN creates excessive micro-clusters, increase the `eps` parameter slightly (e.g., to 0.7 or 0.8) to relax the similarity threshold. Conversely, if unrelated faces are merging into single clusters, decrease `eps` to enforce stricter matching criteria. The [`examples/face_recognition_knn.py`](https://github.com/ageitgey/face_recognition/blob/main/examples/face_recognition_knn.py) file demonstrates techniques for tuning distance thresholds in similar applications.