How face_distance Calculates the Distance Between Faces in Python
face_distance computes the Euclidean (L2) distance between a candidate face encoding and a list of known face encodings using vectorized NumPy operations.
The face_distance function in the ageitgey/face_recognition library provides a fast, mathematically precise way to quantify facial similarity by comparing 128-dimensional embeddings generated by dlib’s face recognition model. Understanding how face_distance calculates these distances enables developers to optimize recognition pipelines and set appropriate matching thresholds for their specific use cases.
What Is face_distance?
face_distance is a utility function defined in face_recognition/api.py that measures the numerical distance between face encodings. Unlike higher-level functions such as compare_faces that return boolean match results, face_distance returns raw floating-point distances, giving developers fine-grained control over recognition sensitivity and custom threshold logic.
How face_distance Works Under the Hood
The implementation relies entirely on pure NumPy operations for performance, avoiding Python loops and leveraging optimized linear algebra libraries.
Input Validation
Before performing calculations, face_distance validates its inputs to prevent downstream errors. If the face_encodings list is empty, the function returns an empty NumPy array immediately. This safety check allows calling code to iterate over results without additional null checks. The validation logic resides at lines 63-74 in face_recognition/api.py.
Vectorized Euclidean Distance Calculation
The core computation occurs in two vectorized steps at lines 75-76 of face_recognition/api.py:
-
Broadcasting subtraction: The function receives
face_encodingsas an n × 128 NumPy array (where n is the number of known faces) andface_to_compareas a single 128-dimensional vector. It subtracts the candidate encoding from every row of the known encodings matrix using NumPy broadcasting, producing an n × 128 difference matrix. -
L2 norm computation: It then computes the Euclidean (L2) norm of each row using
np.linalg.norm(..., axis=1), resulting in a 1-dimensional array of n distances—one floating-point value per known face.
This vectorized approach leverages optimized BLAS libraries under NumPy, making face_distance extremely fast even when comparing a candidate against thousands of known faces.
Interpreting the Results
Smaller distances indicate higher facial similarity. While face_distance returns raw numerical values, the typical threshold used by higher-level helpers such as compare_faces is 0.6, where distances below this value indicate a match. However, applications can implement custom thresholds: lower values (e.g., 0.5) reduce false positives for security applications, while higher values (e.g., 0.7) may suit casual photo organization where missing a match is worse than an occasional false positive.
Practical Code Examples
Single Face Comparison
Calculate the distance between one known face and one unknown face:
import face_recognition
# Load images and extract 128-dimensional encodings
known_image = face_recognition.load_image_file("known_person.jpg")
unknown_image = face_recognition.load_image_file("candidate.jpg")
known_encoding = face_recognition.face_encodings(known_image)[0]
unknown_encoding = face_recognition.face_encodings(unknown_image)[0]
# Calculate Euclidean distance
distance = face_recognition.face_distance([known_encoding], unknown_encoding)[0]
print(f"Euclidean distance: {distance:.4f}")
# Apply standard 0.6 threshold
is_match = distance < 0.6
print(f"Match: {is_match}")
Batch Processing Multiple Faces
Compare one candidate against an entire gallery efficiently using vectorization:
import face_recognition
import numpy as np
# Gallery of known face encodings
gallery_encodings = [encoding1, encoding2, encoding3, encoding4]
# Single candidate encoding
candidate = unknown_encoding
# Vectorized distance array, shape = (len(gallery_encodings),)
distances = face_recognition.face_distance(gallery_encodings, candidate)
# Find best match
best_idx = np.argmin(distances)
best_distance = distances[best_idx]
print(f"Closest match is person #{best_idx} with distance {best_distance:.4f}")
This batch approach corresponds to the reference script shipped with the library at examples/face_distance.py.
Source Code Reference
The face_distance function resides in face_recognition/api.py at lines 63-76 of the ageitgey/face_recognition repository. The implementation is a thin wrapper around NumPy’s linear algebra operations, handling input validation at lines 63-74 and the core vectorized calculation at lines 75-76. Additional usage examples are available in examples/face_distance.py, and unit tests verifying the calculation behavior are located in tests/test_face_recognition.py at lines 162-195.
Summary
face_distancecomputes Euclidean (L2) distance between 128-dimensional face encodings using vectorized NumPy operations inface_recognition/api.py.- Input validation ensures empty encoding lists return empty arrays safely at lines 63-74.
- Vectorized calculation subtracts the candidate encoding from all known encodings and computes
np.linalg.normalong axis 1 at lines 75-76. - Typical threshold of 0.6 distinguishes matches from non-matches, though raw distances allow custom logic.
- Batch processing is fully supported and significantly outperforms iterative comparisons.
Frequently Asked Questions
What distance threshold should I use with face_distance?
The default tolerance used by compare_faces is 0.6, where distances below this value indicate a match. For security-critical applications, use a lower threshold (e.g., 0.5) to minimize false positives. For photo organization tools where missing a match is worse than occasional false positives, consider a higher threshold (e.g., 0.7).
Is face_distance the same as cosine similarity?
No. face_distance calculates Euclidean (L2) distance, which measures the straight-line distance between two points in 128-dimensional space. Cosine similarity measures the angle between vectors regardless of their magnitude. The face_recognition library uses Euclidean distance because the dlib model generates normalized embeddings where L2 distance directly correlates with perceptual facial similarity.
How does face_distance handle multiple faces in one image?
face_distance operates on face encodings, not raw images. You must first extract encodings using face_encodings(), which returns a list of 128-dimensional arrays—one per detected face. If an image contains multiple faces, iterate over each encoding and call face_distance individually, or batch process by passing the entire list of known encodings against each candidate encoding separately.
Where can I find the source code for face_distance?
The implementation resides in face_recognition/api.py at lines 63-76 of the ageitgey/face_recognition repository. The function is a thin wrapper around NumPy’s linalg.norm that handles input validation and vectorized Euclidean distance calculation. You can also view usage examples in examples/face_distance.py and unit tests in tests/test_face_recognition.py that verify the calculation behavior.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →