How to Use the face_recognition CLI Tool for Batch Processing

The face_recognition CLI tool batch processes images by scanning a folder of known faces and comparing them against a directory of unknown images, outputting CSV-style results that can be parallelized across multiple CPU cores.

The ageitgey/face_recognition library includes a powerful command-line interface for automating face recognition workflows. When you need to identify faces across hundreds or thousands of images, the face_recognition CLI tool for batch processing eliminates the need to write custom Python scripts. The tool automatically walks directories, leverages multiprocessing for performance, and produces machine-readable output suitable for integration with shell pipelines or downstream applications.

How the Batch Processing Pipeline Works

The CLI implementation in face_recognition/face_recognition_cli.py follows a strict pipeline architecture designed for high-throughput processing. When you point the tool at a directory of unknown images, it distributes the workload across available CPU cores and emits standardized results for each file.

Scanning Known Faces

The scan_known_people() function (lines 14-32 of face_recognition_cli.py) walks your supplied folder of reference images and builds parallel lists of identities and face encodings. For each image file matching the *.jpg, *.jpeg, or *.png pattern, the function extracts the first face encoding detected and stores it alongside the filename (minus extension) as the person's identifier.

Processing Individual Images

For every unknown image, the test_image() function (lines 42-65) loads the file, optionally down-scales large pictures to improve speed, and computes face encodings using the underlying dlib-based API from face_recognition/api.py. Each encoding is compared against the known set using Euclidean distance, and the function prints CSV-formatted results: filename,name or filename,name,distance when using the --show-distance flag.

Parallel Execution with Multiprocessing

When you specify --cpus with a value other than 1, the process_images_in_process_pool() function (lines 71-94) creates a multiprocessing.Pool using the "forkserver" start method on macOS to avoid memory issues. The tool discovers all image files via image_files_in_folder(), splits the list across the pool, and maps the test_image function to each worker process. This design allows the CLI to saturate all available cores without blocking on I/O or face encoding computation.

Running Batch Jobs from the Command Line

The CLI accepts a folder of known people as the first argument and either a single image or folder of images as the second. By default, processing runs single-threaded, but production workloads should leverage the parallel execution features.

Basic Directory Comparison

To compare every image in an unknown_images/ folder against your known_people/ reference set:

python -m face_recognition.face_recognition_cli \
    known_people/ \
    unknown_images/

Output appears as comma-separated values:


unknown_images/img001.jpg,alice
unknown_images/img002.jpg,unknown_person
unknown_images/img003.jpg,no_persons_found

Parallel Processing Across All CPUs

For maximum throughput on batch jobs, use the --cpus flag with -1 to automatically detect and utilize all available cores:

python -m face_recognition.face_recognition_cli \
    known_people/ \
    unknown_images/ \
    --cpus -1

This triggers the multiprocessing path in process_images_in_process_pool(), significantly reducing processing time for large image collections.

Tuning Recognition Tolerance and Distance

The default tolerance threshold of 0.6 controls how strict the face matching is—lower values reduce false positives but may miss subtle matches. Combine this with --show-distance true to output the numeric distance for debugging:

python -m face_recognition.face_recognition_cli \
    known_people/ \
    unknown_images/ \
    --tolerance 0.45 \
    --show-distance true

Results include the distance metric:


unknown_images/img001.jpg,alice,0.312
unknown_images/img002.jpg,unknown_person,

Integrating Batch Output with External Tools

The CLI intentionally emits lightweight CSV-style output to support Unix-style pipelines and automated workflows.

Filtering Results with Shell Tools

Pipe the output directly into awk or other text processors to filter matches based on confidence scores:

python -m face_recognition.face_recognition_cli \
    known_people/ \
    unknown_images/ \
    --show-distance true | \
    awk -F, '$3 < 0.5 {print $1, $2}'

This command prints only filenames and names where the face distance is below 0.5, effectively filtering for high-confidence matches.

Calling the CLI from Python Scripts

You can invoke the batch processor programmatically using the subprocess module when you need to integrate face recognition into larger Python applications:

import subprocess
import shlex

cmd = (
    "python -m face_recognition.face_recognition_cli "
    "known_people/ unknown_images/ "
    "--cpus 4 --show-distance true"
)

output = subprocess.check_output(shlex.split(cmd)).decode()
for line in output.splitlines():
    filename, name, distance = line.split(',')
    print(f"Match: {filename} → {name} (distance={distance})")

This pattern allows you to leverage the CLI's multiprocessing optimizations while retaining the flexibility of Python for downstream data handling.

Summary

  • The face_recognition CLI tool processes directories recursively, automatically discovering .jpg, .jpeg, and .png files via image_files_in_folder().
  • Batch processing is handled by scan_known_people() for reference loading and test_image() for comparison, with process_images_in_process_pool() managing parallel execution.
  • Use --cpus -1 to enable multiprocessing across all available cores, significantly accelerating large batch jobs.
  • The tolerance parameter (default 0.6) and show-distance flag allow fine-tuning of matching strictness and debugging of borderline cases.
  • CSV-style output enables seamless integration with shell pipelines, awk, or Python subprocess calls for automated workflows.

Frequently Asked Questions

What image file formats does the face_recognition CLI support?

The CLI supports JPEG and PNG images exclusively. The image_files_in_folder() function uses a regular expression to match files ending in .jpg, .jpeg, or .png, skipping all other formats during directory traversal.

How does the CLI handle images containing no faces or multiple faces?

When processing unknown images, the test_image() function outputs no_persons_found if no faces are detected. If faces are found but none match the known set within the specified tolerance, it prints unknown_person. For known people folders, the tool only extracts the first face encoding per image, so reference photos should ideally contain a single subject.

Can I use GPU acceleration with the face_recognition CLI batch processor?

The standard CLI implementation in face_recognition_cli.py uses CPU-based processing via the face_recognition API. For GPU-accelerated batch processing, refer to examples/find_faces_in_batches.py in the repository, which demonstrates how to use the batch_face_locations() function with CUDA-enabled dlib for significantly faster inference on compatible hardware.

How do I improve accuracy when the CLI produces too many false positives?

Reduce the tolerance value below the default 0.6 using the --tolerance flag. According to the source code in face_recognition_cli.py, lower values make the matching more strict by requiring a smaller Euclidean distance between face encodings. Use --show-distance true to analyze the distribution of distances in your dataset and select an optimal threshold for your specific use case.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →