How Deep-Live-Cam Performs Cluster Analysis for Face Embedding Matching with find_closest_centroid
Deep-Live-Cam clusters face embeddings using K-Means to generate representative centroids, then employs find_closest_centroid with dot-product similarity to match detected faces against these centroids in constant time relative to cluster count.
Deep-Live-Cam is an open-source face-swapping application that processes both static images and video streams. To handle the computational cost of matching faces across hundreds or thousands of video frames, the repository implements a cluster analysis pipeline that compresses redundant embeddings into a compact set of reference centroids. This article examines the implementation in modules/cluster_analysis.py, explaining how the system builds these centroids and uses find_closest_centroid to perform efficient face embedding matching at runtime.
Building Reference Centroids via K-Means Clustering
When processing video sources with many target faces, Deep-Live-Cam avoids brute-force comparison against every frame’s embeddings. Instead, it leverages find_cluster_centroids in modules/cluster_analysis.py to distill the full embedding set into a handful of representative vectors.
The function executes an elbow method search across k values from 1 to max_k (default 10):
# modules/cluster_analysis.py
def find_cluster_centroids(embeddings, max_k=10) -> Any:
inertia = []
cluster_centroids = []
K = range(1, max_k+1)
for k in K:
kmeans = KMeans(n_clusters=k, random_state=0)
kmeans.fit(embeddings)
inertia.append(kmeans.inertia_)
cluster_centroids.append({"k": k, "centroids": kmeans.cluster_centers_})
# Choose the k that gives the largest drop in inertia
diffs = [inertia[i] - inertia[i+1] for i in range(len(inertia)-1)]
optimal_centroids = cluster_centroids[diffs.index(max(diffs)) + 1]['centroids']
return optimal_centroids
Key implementation details:
- Inertia tracking – The algorithm records the sum of squared distances for each k, identifying the "elbow" where diminishing returns on cluster compactness indicate the optimal k.
- Deterministic initialization –
random_state=0ensures reproducible centroid generation across runs. - Dimensional reduction – The output centroids (typically 512-dimensional for InsightFace embeddings) act as a compressed signature of the target identity, reducing subsequent matching complexity from O(n) to O(k).
Matching Embeddings with find_closest_centroid
Once centroids are established, Deep-Live-Cam matches new face detections using find_closest_centroid. This function exploits the fact that InsightFace embeddings are L2-normalized, allowing cosine similarity to be computed as a simple dot product.
# modules/cluster_analysis.py
def find_closest_centroid(centroids: list, normed_face_embedding) -> list:
try:
centroids = np.array(centroids)
normed_face_embedding = np.array(normed_face_embedding)
similarities = np.dot(centroids, normed_face_embedding)
closest_centroid_index = np.argmax(similarities)
return closest_centroid_index, centroids[closest_centroid_index]
except ValueError:
return None
Cosine similarity via dot product – Because both the centroid matrix and query embedding are unit vectors, np.dot(centroids, normed_face_embedding) yields the cosine similarity without explicit normalization. The function returns the index of the maximum similarity and the centroid vector itself, enabling O(k) lookup regardless of the original video length.
Pipeline Integration: Video Analysis and Live Swapping
The cluster analysis operates in two distinct contexts within the Deep-Live-Cam architecture: preprocessing for video targets and real-time matching during face swapping.
Video Preprocessing in face_analyser.py
During target video analysis (lines 31-44 of modules/face_analyser.py), the system collects all embeddings from extracted frames. It then:
- Invokes
find_cluster_centroidsto generate centroids from the accumulated embeddings. - Tags each face with its nearest centroid index using
find_closest_centroid. - Builds a
source_target_mapthat groups faces by their assigned centroid for efficient retrieval during swapping.
Runtime Matching in face_swapper.py
In live mode or when using a simple target map (lines 492-511 of modules/processors/frame/face_swapper.py), the system calls find_closest_centroid to pair detected faces with target embeddings:
# Simplified excerpt from modules/processors/frame/face_swapper.py
if simple_map:
target_embeddings = simple_map["target_embeddings"]
for detected_face in detected_faces:
closest_idx, _ = find_closest_centroid(
target_embeddings,
detected_face.normed_embedding
)
source_target_pairs.append(
(source_faces[closest_idx], detected_face)
)
This logic handles asymmetric scenarios—whether there are more detections than targets or vice versa—by always finding the nearest centroid in the embedding space.
Code Implementation Examples
Direct Usage of find_closest_centroid
To implement the matching logic in custom scripts:
from modules.cluster_analysis import find_cluster_centroids, find_closest_centroid
import numpy as np
# Assume embeddings is a list of 512-dim numpy arrays from InsightFace
centroids = find_cluster_centroids(embeddings, max_k=10)
# Match a new detection
query = detected_face.normed_embedding # Shape: (512,)
idx, matched_centroid = find_closest_centroid(centroids, query)
print(f"Matched to centroid {idx} with similarity {np.dot(matched_centroid, query):.4f}")
Full Video Processing Pipeline
Replicating the video preprocessing workflow from modules/face_analyser.py:
from modules.cluster_analysis import find_cluster_centroids, find_closest_centroid
from modules.face_analyser import get_many_faces
import cv2
# Collect all embeddings from frame sequence
all_embeddings = []
frame_data = []
for frame_path in temp_frame_paths:
frame = cv2.imread(frame_path)
faces = get_many_faces(frame)
for face in faces:
all_embeddings.append(face.normed_embedding)
frame_data.append({"path": frame_path, "faces": faces})
# Cluster to find optimal centroids
centroids = find_cluster_centroids(all_embeddings, max_k=10)
# Assign each face to its nearest centroid
for frame in frame_data:
for face in frame["faces"]:
centroid_idx, _ = find_closest_centroid(centroids, face.normed_embedding)
face["target_centroid"] = centroid_idx
Summary
- Cluster compression –
find_cluster_centroidsinmodules/cluster_analysis.pyuses K-Means with an elbow-method optimizer to compress video embeddings into 1-10 representative centroids. - Efficient matching –
find_closest_centroidleverages normalized embeddings to compute cosine similarity vianp.dot, achieving O(k) complexity where k is the cluster count. - Dual-stage pipeline – Video preprocessing (
face_analyser.py) builds centroid maps, while runtime swapping (face_swapper.py) uses these maps for instantaneous face-to-target matching. - Numerical stability – The system assumes L2-normalized inputs from InsightFace, eliminating the need for explicit distance calculations beyond matrix multiplication.
Frequently Asked Questions
How does Deep-Live-Cam determine the optimal number of clusters for face embeddings?
The find_cluster_centroids function tests k values from 1 to 10, fitting a K-Means model for each and recording the inertia (within-cluster sum of squares). It calculates the difference in inertia between consecutive k values and selects the k with the largest drop—the "elbow" point—indicating the best trade-off between compression and accuracy.
Why does find_closest_centroid use dot product instead of Euclidean distance?
Because InsightFace outputs L2-normalized embeddings (unit vectors), the dot product between two vectors equals their cosine similarity. This is mathematically equivalent to measuring angular distance but computationally cheaper than Euclidean distance, requiring only a single matrix multiplication via np.dot.
When is cluster analysis triggered in the Deep-Live-Cam pipeline?
Cluster analysis occurs during the preprocessing phase for video targets in modules/face_analyser.py, where the system aggregates embeddings from all frames to build centroids. It is also used dynamically in modules/processors/frame/face_swapper.py during live processing when the number of target embeddings exceeds practical linear search thresholds.
What happens if find_closest_centroid receives malformed embeddings?
The function wraps its logic in a try-except block that catches ValueError. If the centroid list or query embedding has incompatible shapes—such as mismatched dimensions between the 512-dim InsightFace output and the centroid matrix—the function returns None to prevent runtime crashes during face swapping.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →