How to Handle Different Image Sizes and File Formats in face_recognition

The face_recognition library uses Pillow to load virtually any image format and provides parameters like number_of_times_to_upsample to handle varying image dimensions without manual preprocessing.

Working with real-world image data means dealing with JPEGs from smartphones, PNGs from web scraping, and high-resolution frames from video streams. The face_recognition library, maintained in the ageitgey/face_recognition repository, simplifies how you handle different image sizes and file formats by leveraging Pillow for I/O and exposing explicit controls for scaling during face detection.

Loading Images in Any Format with Pillow

The library delegates all file format handling to Pillow (PIL), which supports dozens of formats including JPEG, PNG, BMP, GIF, TIFF, and WebP.

The load_image_file Function

In face_recognition/api.py, the load_image_file function wraps Pillow’s image opener and converts the result to a NumPy array suitable for dlib processing:

def load_image_file(file, mode='RGB'):
    """Loads an image file (.jpg, .png, etc) into a numpy array"""
    im = PIL.Image.open(file)                 # ← Pillow handles many formats

    if mode:
        im = im.convert(mode)                  # 'RGB' (default) or 'L' (grayscale)

    return np.array(im)                       # → NumPy array (H × W × C)

Source: [face_recognition/api.py lines 78‑89](https://github.com/ageitgey/face_recognition/blob/master/face_recognition/api.py#L78-L89)

Key details:

  • Supported modes: Only 'RGB' (3‑channel color) and 'L' (grayscale) are accepted, matching dlib’s expectations.
  • GIF handling: Animated GIFs are read as a single static frame.

Handling Corrupt or Truncated Images

Network downloads or incomplete transfers often produce truncated files. The library sets Pillow’s safety flag at module initialization in face_recognition/api.py (line 15) to prevent crashes:

ImageFile.LOAD_TRUNCATED_IMAGES = True

This allows load_image_file to process partially downloaded images rather than raising an exception, though visual artifacts may occur in the loaded data.

Managing Different Image Sizes and Resolutions

High‑resolution images from modern cameras (12MP+) can strain CPU detection pipelines, while tiny faces in wide‑angle shots may be missed entirely. The library provides two complementary strategies.

Detecting Small Faces in Large Images

The face_locations function accepts a number_of_times_to_upsample parameter that controls how many times the image pyramid is upsampled before running the HOG or CNN detector:

def face_locations(img, number_of_times_to_upsample=1, model="hog"):
    # ...

Source: [face_recognition/api.py lines 96‑99](https://github.com/ageitgey/face_recognition/blob/master/face_recognition/api.py#L96-L99)

Increasing this value (e.g., to 2 or 3) forces the detector to search at finer scales, improving recall on small faces at the cost of slower execution.

Explicit Resizing for Performance

For real‑time applications such as webcam streams, the examples demonstrate explicit downsampling with OpenCV before processing. In examples/blink_detection.py (lines 27‑28):

small_frame = cv2.resize(frame, (0, 0), fx=0.25, fy=0.25)

Similarly, examples/facerec_from_webcam_faster.py shows converting the BGR OpenCV format to RGB and processing at quarter resolution to maintain high frame rates.

Complete Code Examples

Loading a JPEG and Detecting Small Faces

import face_recognition

# Load any format Pillow supports (JPEG, PNG, BMP, etc.)

image = face_recognition.load_image_file("group_photo.jpg")

# Upsample twice to find faces that occupy <1% of image area

face_locations = face_recognition.face_locations(
    image, 
    number_of_times_to_upsample=2
)
print(f"Found {len(face_locations)} faces")

Resizing Large Images with Pillow Before Detection

from PIL import Image
import numpy as np
import face_recognition

# Open with Pillow (supports dozens of formats)

pil_img = Image.open("huge_image.png")

# Reduce to max 800px on longest side while keeping aspect ratio

pil_img.thumbnail((800, 800))

# Convert to NumPy array for face_recognition API

np_img = np.array(pil_img)

# Detect faces (default upsample is sufficient for this size)

locations = face_recognition.face_locations(np_img)
print(locations)

Real-Time Webcam Processing with OpenCV Resizing

import cv2
import face_recognition

video_capture = cv2.VideoCapture(0)

while True:
    ret, frame = video_capture.read()
    
    # Downscale to 25% of original size for faster processing

    small_frame = cv2.resize(frame, (0, 0), fx=0.25, fy=0.25)
    
    # Convert BGR (OpenCV) to RGB (face_recognition expects RGB)

    rgb_small_frame = small_frame[:, :, ::-1]
    
    # Find faces on the smaller frame

    face_locations = face_recognition.face_locations(rgb_small_frame)
    
    # Exit on 'q' key

    if cv2.waitKey(1) & 0xFF == ord('q'):
        break

video_capture.release()
cv2.destroyAllWindows()

Summary

  • Universal format support: load_image_file in face_recognition/api.py leverages Pillow to read JPEG, PNG, BMP, GIF, TIFF, WebP, and other formats, converting them to RGB NumPy arrays.
  • Corruption resilience: The library sets ImageFile.LOAD_TRUNCATED_IMAGES = True to handle incomplete downloads gracefully.
  • Small face detection: Increase number_of_times_to_upsample in face_locations to detect tiny faces in high‑resolution images without manual resizing.
  • Performance optimization: For real‑time video, explicitly resize frames with OpenCV (as shown in examples/blink_detection.py) before passing them to the recognition pipeline.

Frequently Asked Questions

What image formats does face_recognition support?

The library supports any format that Pillow (PIL) can open, including JPEG, PNG, BMP, GIF, TIFF, and WebP. The load_image_file function in face_recognition/api.py uses PIL.Image.open() internally, so if Pillow recognizes the file extension and codec, face_recognition can process it.

How do I detect small faces in a high-resolution photo?

Pass a higher value to the number_of_times_to_upsample parameter in face_locations. The default value of 1 works for most medium‑resolution images, but setting it to 2 or 3 forces the HOG or CNN detector to search at finer scales, improving detection of faces that occupy only a small percentage of the total image area.

Should I resize images before passing them to face_recognition?

For batch processing of static photos, you can rely on the number_of_times_to_upsample parameter to handle scaling internally. However, for real‑time video streams or webcam feeds, explicit resizing with OpenCV (e.g., cv2.resize(frame, (0, 0), fx=0.25, fy=0.25)) is recommended to reduce CPU load and maintain high frame rates, as demonstrated in the examples/blink_detection.py file.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →