How to Handle Different Image Sizes and File Formats in face_recognition
The face_recognition library uses Pillow to load virtually any image format and provides parameters like number_of_times_to_upsample to handle varying image dimensions without manual preprocessing.
Working with real-world image data means dealing with JPEGs from smartphones, PNGs from web scraping, and high-resolution frames from video streams. The face_recognition library, maintained in the ageitgey/face_recognition repository, simplifies how you handle different image sizes and file formats by leveraging Pillow for I/O and exposing explicit controls for scaling during face detection.
Loading Images in Any Format with Pillow
The library delegates all file format handling to Pillow (PIL), which supports dozens of formats including JPEG, PNG, BMP, GIF, TIFF, and WebP.
The load_image_file Function
In face_recognition/api.py, the load_image_file function wraps Pillow’s image opener and converts the result to a NumPy array suitable for dlib processing:
def load_image_file(file, mode='RGB'):
"""Loads an image file (.jpg, .png, etc) into a numpy array"""
im = PIL.Image.open(file) # ← Pillow handles many formats
if mode:
im = im.convert(mode) # 'RGB' (default) or 'L' (grayscale)
return np.array(im) # → NumPy array (H × W × C)
Source: [face_recognition/api.py lines 78‑89](https://github.com/ageitgey/face_recognition/blob/master/face_recognition/api.py#L78-L89)
Key details:
- Supported modes: Only
'RGB'(3‑channel color) and'L'(grayscale) are accepted, matching dlib’s expectations. - GIF handling: Animated GIFs are read as a single static frame.
Handling Corrupt or Truncated Images
Network downloads or incomplete transfers often produce truncated files. The library sets Pillow’s safety flag at module initialization in face_recognition/api.py (line 15) to prevent crashes:
ImageFile.LOAD_TRUNCATED_IMAGES = True
This allows load_image_file to process partially downloaded images rather than raising an exception, though visual artifacts may occur in the loaded data.
Managing Different Image Sizes and Resolutions
High‑resolution images from modern cameras (12MP+) can strain CPU detection pipelines, while tiny faces in wide‑angle shots may be missed entirely. The library provides two complementary strategies.
Detecting Small Faces in Large Images
The face_locations function accepts a number_of_times_to_upsample parameter that controls how many times the image pyramid is upsampled before running the HOG or CNN detector:
def face_locations(img, number_of_times_to_upsample=1, model="hog"):
# ...
Source: [face_recognition/api.py lines 96‑99](https://github.com/ageitgey/face_recognition/blob/master/face_recognition/api.py#L96-L99)
Increasing this value (e.g., to 2 or 3) forces the detector to search at finer scales, improving recall on small faces at the cost of slower execution.
Explicit Resizing for Performance
For real‑time applications such as webcam streams, the examples demonstrate explicit downsampling with OpenCV before processing. In examples/blink_detection.py (lines 27‑28):
small_frame = cv2.resize(frame, (0, 0), fx=0.25, fy=0.25)
Similarly, examples/facerec_from_webcam_faster.py shows converting the BGR OpenCV format to RGB and processing at quarter resolution to maintain high frame rates.
Complete Code Examples
Loading a JPEG and Detecting Small Faces
import face_recognition
# Load any format Pillow supports (JPEG, PNG, BMP, etc.)
image = face_recognition.load_image_file("group_photo.jpg")
# Upsample twice to find faces that occupy <1% of image area
face_locations = face_recognition.face_locations(
image,
number_of_times_to_upsample=2
)
print(f"Found {len(face_locations)} faces")
Resizing Large Images with Pillow Before Detection
from PIL import Image
import numpy as np
import face_recognition
# Open with Pillow (supports dozens of formats)
pil_img = Image.open("huge_image.png")
# Reduce to max 800px on longest side while keeping aspect ratio
pil_img.thumbnail((800, 800))
# Convert to NumPy array for face_recognition API
np_img = np.array(pil_img)
# Detect faces (default upsample is sufficient for this size)
locations = face_recognition.face_locations(np_img)
print(locations)
Real-Time Webcam Processing with OpenCV Resizing
import cv2
import face_recognition
video_capture = cv2.VideoCapture(0)
while True:
ret, frame = video_capture.read()
# Downscale to 25% of original size for faster processing
small_frame = cv2.resize(frame, (0, 0), fx=0.25, fy=0.25)
# Convert BGR (OpenCV) to RGB (face_recognition expects RGB)
rgb_small_frame = small_frame[:, :, ::-1]
# Find faces on the smaller frame
face_locations = face_recognition.face_locations(rgb_small_frame)
# Exit on 'q' key
if cv2.waitKey(1) & 0xFF == ord('q'):
break
video_capture.release()
cv2.destroyAllWindows()
Summary
- Universal format support:
load_image_fileinface_recognition/api.pyleverages Pillow to read JPEG, PNG, BMP, GIF, TIFF, WebP, and other formats, converting them to RGB NumPy arrays. - Corruption resilience: The library sets
ImageFile.LOAD_TRUNCATED_IMAGES = Trueto handle incomplete downloads gracefully. - Small face detection: Increase
number_of_times_to_upsampleinface_locationsto detect tiny faces in high‑resolution images without manual resizing. - Performance optimization: For real‑time video, explicitly resize frames with OpenCV (as shown in
examples/blink_detection.py) before passing them to the recognition pipeline.
Frequently Asked Questions
What image formats does face_recognition support?
The library supports any format that Pillow (PIL) can open, including JPEG, PNG, BMP, GIF, TIFF, and WebP. The load_image_file function in face_recognition/api.py uses PIL.Image.open() internally, so if Pillow recognizes the file extension and codec, face_recognition can process it.
How do I detect small faces in a high-resolution photo?
Pass a higher value to the number_of_times_to_upsample parameter in face_locations. The default value of 1 works for most medium‑resolution images, but setting it to 2 or 3 forces the HOG or CNN detector to search at finer scales, improving detection of faces that occupy only a small percentage of the total image area.
Should I resize images before passing them to face_recognition?
For batch processing of static photos, you can rely on the number_of_times_to_upsample parameter to handle scaling internally. However, for real‑time video streams or webcam feeds, explicit resizing with OpenCV (e.g., cv2.resize(frame, (0, 0), fx=0.25, fy=0.25)) is recommended to reduce CPU load and maintain high frame rates, as demonstrated in the examples/blink_detection.py file.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →