How Face Recognition Using Embeddings Works in the Multi-Cam Face Tracker

Face recognition using embeddings in this system works by generating 512-dimensional vectors from detected faces via InsightFace, then comparing live embeddings against a stored database using cosine similarity to identify matches above a configurable threshold.

The multi-cam-face-tracker repository implements a complete face recognition pipeline that converts facial images into mathematical representations called embeddings. These high-dimensional vectors capture unique facial features in a format that enables fast, accurate comparison across multiple camera streams in real time.

Generating Face Embeddings with InsightFace

The system relies on the InsightFace library to transform raw face images into compact numerical representations. This process happens in core/face_detection.py, where the FaceDetector class orchestrates model loading and inference.

Model Initialization and Detection

The embedding generation begins with model initialization in FaceDetector._load_model(). This method creates an insightface.app.FaceAnalysis instance with detection, recognition, and attribute modules enabled:


# From core/face_detection.py

def _load_model(self):
    self.model = insightface.app.FaceAnalysis(
        name='buffalo_l',
        root='./models',
        providers=['CUDAExecutionProvider', 'CPUExecutionProvider']
    )
    self.model.prepare(ctx_id=0, det_size=(640, 640))

When processing a frame, detect_faces(image) calls self.model.get(image), which returns face objects containing pre-computed embeddings from the model's recognition head.

The 512-Dimensional Embedding Vector

Each detected face is wrapped in a Face dataclass that stores the raw embedding—a NumPy array with shape (512,) representing a 512-dimensional vector:

faces = detector.detect_faces(frame)           # list of Face objects

live_embedding = faces[0].embedding           # np.ndarray, shape (512,)

These embeddings encode facial geometry and texture patterns in a high-dimensional space where similar faces cluster together, enabling mathematical comparison regardless of lighting, pose, or minor appearance changes.

Building the Known Face Database

Before recognition can occur, the system must establish a reference library of known individuals. This happens through load_known_faces() in core/face_detection.py and persistent storage in core/database.py.

Loading and Processing Reference Images

When the application starts or when users add new persons via the UI, load_known_faces(dir) iterates over image files in the specified directory:


# From core/face_detection.py

def load_known_faces(self, known_faces_dir):
    for filename in os.listdir(known_faces_dir):
        if filename.endswith(('.jpg', '.png')):
            image_path = os.path.join(known_faces_dir, filename)
            img = cv2.imread(image_path)
            faces = self.detect_faces(img)
            
            if faces:
                name = os.path.splitext(filename)[0]
                self.known_faces.append(KnownFace(
                    name=name,
                    embedding=faces[0].embedding,
                    image_path=image_path
                ))

Each valid face detection yields a KnownFace object containing the person's name and their 512-dimensional embedding vector.

Persistent Storage in SQLite

For durability across sessions, embeddings are stored in core/database.py via FaceDatabase.add_known_face():


# From core/database.py

def add_known_face(self, name, embedding, image_path):
    embedding_blob = embedding.tobytes()
    self.cursor.execute(
        "INSERT INTO known_faces (name, embedding, image_path) VALUES (?, ?, ?)",
        (name, embedding_blob, image_path)
    )
    self.conn.commit()

The known_faces table stores the binary embedding blob, enabling the system to rebuild the known face library on startup without reprocessing reference images.

Face Recognition Using Embeddings Comparison

The core recognition logic resides in FaceDetector.recognize_faces(), which implements vector comparison using cosine similarity.

Cosine Similarity Calculation

When processing live video frames, the system compares detected face embeddings against the known face library:


# From core/face_detection.py

def recognize_faces(self, faces):
    results = []
    if not self.known_faces:
        return [(face, None, 0.0) for face in faces]
    
    known_embeddings = np.array([kf.embedding for kf in self.known_faces])
    
    for face in faces:
        # Cosine similarity calculation

        similarities = np.dot(known_embeddings, face.embedding) / (
            np.linalg.norm(known_embeddings, axis=1) * np.linalg.norm(face.embedding)
        )
        max_idx = np.argmax(similarities)
        max_similarity = similarities[max_idx]
        
        if max_similarity >= self.recognition_threshold:
            results.append((face, self.known_faces[max_idx], max_similarity))
        else:
            results.append((face, None, max_similarity))
    
    return results

This implementation computes the cosine similarity between the live embedding and all known embeddings simultaneously using vectorized NumPy operations. Cosine similarity measures the angle between vectors rather than their absolute distance, making it robust to variations in lighting and scale.

Threshold-Based Matching

The system uses a configurable threshold defined in config.yaml to determine valid matches. When max_similarity exceeds self.recognition_threshold (typically set between 0.5 and 0.7), the system labels the face with the corresponding known person's name. Otherwise, the face remains marked as "Unknown."

This threshold provides a tunable balance between precision (avoiding false positives) and recall (catching all instances of known persons).

Real-Time Recognition Pipeline

The end-to-end flow in main.py demonstrates how these components work together for live multi-camera processing:

from core.face_detection import FaceDetector
import cv2
import yaml

# Load configuration

with open("config/config.yaml", "r") as f:
    cfg = yaml.safe_load(f)

detector = FaceDetector(cfg)
detector.load_known_faces("assets/known_faces")

cap = cv2.VideoCapture(0)

while True:
    ret, frame = cap.read()
    if not ret:
        break

    faces = detector.detect_faces(frame)
    results = detector.recognize_faces(faces)

    for face, known, sim in results:
        label = known.name if known else "Unknown"
        x1, y1, x2, y2 = map(int, face.bbox)
        cv2.rectangle(frame, (x1, y1), (x2, y2), (0, 255, 0), 2)
        cv2.putText(frame, f"{label} ({sim:.2f})", (x1, y1 - 10),
                    cv2.FONT_HERSHEY_SIMPLEX, 0.6, (0, 255, 0), 2)

    cv2.imshow("Multi-Cam Face Tracker", frame)
    if cv2.waitKey(1) & 0xFF == 27:
        break

cap.release()
cv2.destroyAllWindows()

Recognized events are logged via FaceDatabase.log_face_event() in core/database.py, creating a permanent record linking names, confidence scores, timestamps, and camera IDs.

Summary

  • Face recognition using embeddings in this system converts facial images into 512-dimensional vectors using the InsightFace library.
  • The FaceDetector class in core/face_detection.py handles detection, embedding extraction, and cosine similarity comparison.
  • Known faces are stored as embedding vectors in memory and persisted to SQLite via core/database.py.
  • Recognition occurs through vectorized cosine similarity calculations between live embeddings and the known face library.
  • A configurable threshold in config.yaml determines whether a match is accepted, balancing precision and recall for real-time multi-camera deployments.

Frequently Asked Questions

What is a face embedding?

A face embedding is a mathematical representation of a facial image as a high-dimensional vector—specifically 512 dimensions in this system. The embedding encodes unique facial features such as the distance between eyes, nose shape, and jawline contours into a format that enables mathematical comparison. Similar faces produce embeddings that cluster together in vector space, making it possible to quantify facial similarity using geometric distance metrics.

Why does the system use cosine similarity instead of Euclidean distance?

The system uses cosine similarity because it measures the angle between embedding vectors rather than their absolute distance. This approach is more robust to variations in lighting conditions, image scale, and minor pose changes that might shift the vector's magnitude but preserve its directional orientation. By normalizing the comparison to the unit sphere, cosine similarity ensures that recognition depends on facial structure rather than image brightness or camera distance.

How accurate is the face recognition using embeddings in this system?

Accuracy depends primarily on the quality of reference images and the configured recognition_threshold in config.yaml. The InsightFace model generates highly discriminative 512-dimensional embeddings that achieve state-of-the-art performance on standard benchmarks. In practice, the system achieves reliable recognition when reference images are clear, frontal, and well-lit. The configurable threshold allows administrators to tune the trade-off between false positives (incorrectly identifying strangers as known) and false negatives (failing to recognize enrolled individuals).

Can the system handle multiple faces per frame?

Yes, the architecture is designed for multi-face detection and recognition per camera frame. The detect_faces() method returns a list of all detected faces in an image, and recognize_faces() processes each embedding independently against the known face library. This enables the system to track and identify multiple individuals simultaneously across multi-camera deployments, with each recognized face receiving its own similarity score and identity label.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →