How to Improve Face Recognition Performance on Raspberry Pi: A Complete Optimization Guide

Set your Pi camera to 320×240 resolution, use the HOG detection model with zero upsampling, and limit encoding jitter to achieve real-time face recognition on Raspberry Pi hardware.

The face_recognition library by ageitgey wraps high-performance C++ dlib algorithms in a thin Python API. While the heavy computation happens in native code, you can significantly improve face recognition performance on Raspberry Pi by optimizing how the Python wrapper processes input frames and configures detection parameters in face_recognition/api.py.

How the Library Processes Images

Understanding the pipeline helps identify optimization points. The library executes four main steps for every frame according to the source in face_recognition/api.py:

  1. Load image – load_image_file uses Pillow to read files into NumPy arrays (lines 78-90).
  2. Detect faces – face_locations calls dlib's HOG or CNN detector via _raw_face_locations (lines 108-122).
  3. Encode faces – face_encodings runs the 128-dimensional ResNet model on each detected face (lines 203-215).
  4. Compare – compare_faces calculates Euclidean distance between encodings (lines 217-226).

Each function exposes optional arguments that directly impact CPU usage on resource-constrained devices.

Critical Performance Parameters

The following arguments in face_recognition/api.py control computational workload:

Argument Function Effect Recommended Pi Value
model face_locations "hog" uses CPU only; "cnn" requires GPU acceleration "hog"
number_of_times_to_upsample face_locations Higher values detect smaller faces but multiply processing time 0 or 1
num_jitters face_encodings More jitter improves accuracy but increases encoding time linearly 1 (default)
Image resolution Input preparation Smaller frames reduce pixels scanned by dlib 320 × 240

Raspberry Pi-Specific Optimizations

Capture Low-Resolution Frames

The most effective optimization is reducing camera resolution. The official demo in examples/facerec_on_raspberry_pi.py (lines 16-18) configures the PiCamera to capture 320×240 frames into a pre-allocated NumPy array:

camera = picamera.PiCamera()
camera.resolution = (320, 240)
output = np.empty((240, 320, 3), dtype=np.uint8)

This resolution reduces the pixel count by over 90% compared to 1080p, directly decreasing detection time in dlib's HOG scanner.

Disable Upsampling

When calling face_locations, set number_of_times_to_upsample=0 to prevent the detector from creating enlarged copies of the image. This eliminates expensive image preprocessing passes critical for real-time performance on the Pi's limited CPU.

Avoid CNN Model

Do not use model="cnn" on Raspberry Pi. The CNN detector requires CUDA-capable GPUs to achieve performance benefits. On Pi hardware, the CNN model runs slower than HOG because it lacks GPU acceleration. Use model="hog" exclusively for CPU-only deployment.

Minimize Encoding Jitter

Keep num_jitters=1 (the default) in face_encodings. While higher values (10-100) improve robustness to lighting variations, they create prohibitive latency on the Pi. The linear performance cost outweighs accuracy gains for real-time applications.

Pre-compute Known Encodings

Load and encode your reference images once at startup, then reuse those 128-dimensional vectors for comparison. Never recompute known face encodings inside the main processing loop.

Complete Optimized Implementation

This implementation combines all optimizations from face_recognition/api.py and examples/facerec_on_raspberry_pi.py:

import face_recognition
import picamera
import numpy as np

# Initialize camera at low resolution

camera = picamera.PiCamera()
camera.resolution = (320, 240)
frame_buffer = np.empty((240, 320, 3), dtype=np.uint8)

# Load known faces once (pre-compute encodings)

known_image = face_recognition.load_image_file("obama_small.jpg")
known_encoding = face_recognition.face_encodings(known_image)[0]

# Real-time processing loop

while True:
    # Capture directly into NumPy array (no conversion overhead)

    camera.capture(frame_buffer, format="rgb")
    
    # Fast HOG detection, zero upsampling

    face_locations = face_recognition.face_locations(
        frame_buffer, 
        number_of_times_to_upsample=0, 
        model="hog"
    )
    
    # One-jitter encoding (default) - cheap & fast

    face_encodings = face_recognition.face_encodings(
        frame_buffer, 
        known_face_locations=face_locations,
        num_jitters=1
    )
    
    # Compare against known encodings

    for encoding in face_encodings:
        matches = face_recognition.compare_faces([known_encoding], encoding)
        name = "Barack Obama" if matches[0] else "Unknown"
        print(f"I see someone named {name}!")

Key differences from stock implementations:

  • number_of_times_to_upsample=0 eliminates expensive image preprocessing
  • 320×240 resolution reduces pixel count by 90% versus 1080p
  • Pre-allocated NumPy array avoids memory allocation overhead during capture

Key Source Files for Reference

Summary

To improve face recognition performance on Raspberry Pi:

  • Capture at 320×240 resolution to reduce pixel processing by 90% compared to 1080p.
  • Use model="hog" exclusively; avoid the CNN model without GPU acceleration.
  • Set number_of_times_to_upsample=0 to eliminate expensive image upsampling passes.
  • Keep num_jitters=1 (default) to minimize encoding computation time.
  • Pre-compute known face encodings once at startup rather than in the main loop.

These optimizations, derived directly from the face_recognition source code in face_recognition/api.py and examples/facerec_on_raspberry_pi.py, enable real-time face recognition on Raspberry Pi hardware with minimal latency.

Frequently Asked Questions

How do I make face recognition run faster on Raspberry Pi?

Reduce the camera resolution to 320×240, use the HOG detection model with number_of_times_to_upsample=0, and ensure num_jitters=1 when encoding faces. These settings minimize CPU load according to the implementation in face_recognition/api.py, allowing the Pi to process frames in real-time.

Should I use the CNN model on Raspberry Pi 4?

No. The CNN model requires CUDA-capable GPUs to achieve performance benefits. On Raspberry Pi hardware, the CNN model runs slower than the HOG model because the Pi lacks GPU acceleration for these operations. Use model="hog" exclusively for CPU-only deployment on the Pi.

What is the optimal camera resolution for real-time face recognition on Pi?

320×240 pixels provides the best balance between detection accuracy and processing speed. The official Raspberry Pi demo in examples/facerec_on_raspberry_pi.py uses this resolution, which reduces the pixel count by over 90% compared to 1080p while maintaining sufficient detail for the HOG face detector to locate faces accurately.

Does reducing num_jitters affect recognition accuracy?

Reducing num_jitters from higher values (10-100) to the default of 1 may slightly reduce robustness to extreme lighting variations, but the impact is minimal for typical Pi use cases. The performance gain—linear reduction in encoding time—far outweighs the minor accuracy trade-off for real-time applications on resource-constrained hardware.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →