FaceSwap Multi-Process vs Single-Process Mode: Performance Impact and Resource Trade-offs

FaceSwap's multi-process mode runs extraction, training, and conversion stages concurrently across CPU cores for maximum throughput, while single-process mode executes them sequentially to minimize VRAM consumption.

The deepfakes/faceswap repository provides two distinct execution strategies for its computer vision pipelines. Understanding how multi-process vs single-process mode affects performance allows you to optimize for speed on high-end GPUs or stability on memory-constrained systems.

Architecture of Multi-Process vs Single-Process Mode

Pipeline Implementation in plugins/extract/pipeline.py

The execution strategy is determined by the multiprocess boolean passed to the ExtractionPipeline class. When enabled, the system spawns independent Python Process instances for each phase—detection, alignment, masking, and cropping—allowing these stages to run simultaneously. The pipeline coordinates these processes through inter-process communication, keeping both CPU cores and the GPU active.

In single-process mode, the same phases execute in a simple sequential loop within the parent process. The system waits for detection to complete before starting alignment, ensuring only one set of model weights and GPU buffers resides in memory at any given time.

CLI Argument Definition in lib/cli/args_extract_convert.py

The user-facing flag is defined in lib/cli/args_extract_convert.py with support for NVIDIA, ROCm, and Apple Silicon backends:

argument_list.append({
    "opts": ("-P", "--singleprocess"),
    "action": "store_true",
    "default": False,
    "backend": ("nvidia", "rocm", "apple_silicon"),
    "group": _("settings"),
    "help": _(
        "Don't run extraction in parallel. Will run each part of the extraction process "
        "separately (one after the other) rather than all at the same time. Useful if "
        "VRAM is at a premium.")
})

Flag Processing in scripts/extract.py

During initialization, the extraction script evaluates the flag and inverts the logic to determine the multiprocessing state, as seen in scripts/extract.py:


# In scripts/extract.py

multiprocess = not self._args.singleprocess

# later:

pipeline = ExtractionPipeline(..., multiprocess=multiprocess)

Performance Comparison: Speed vs Memory

Multi-process mode maximizes throughput by dividing work across available CPU cores via lib/multithreading.py utilities and maintaining constant GPU utilization. This parallelism reduces wall-clock time significantly but requires sufficient VRAM—typically 8 GB or more—to accommodate simultaneous model instances and buffer allocations across all active phases.

Single-process mode prioritizes memory efficiency over speed. By executing phases sequentially in the same Python process, GPU buffers are released immediately after each stage completes. This approach reduces VRAM usage substantially, making it viable for GPUs with limited memory, CPU-only configurations, or when running multiple FaceSwap jobs simultaneously on the same hardware.

Configuration and Usage Examples

Default Multi-Process Execution

Run the extraction pipeline with maximum parallelism for fastest processing on well-equipped GPUs:

python faceswap.py extract -i ./src -a ./alignments -r ./extract

Single-Process Mode for Memory Constraints

Enable sequential processing with the -sp or --singleprocess flag to reduce VRAM footprint:

python faceswap.py extract -i ./src -a ./alignments -r ./extract -sp

Limiting Parallelism with the Jobs Flag

You can fine-tune resource usage without fully disabling multiprocessing by capping concurrent jobs with the -j or --jobs argument defined in the same CLI arguments file:

python faceswap.py extract -i ./src -a ./alignments -r ./extract -j 2

This restricts the pipeline to two concurrent processes, offering a middle ground between raw speed and memory conservation.

Platform-Specific Considerations

Windows users often prefer single-process mode due to known fragility in Python's multiprocessing implementation on that platform, which can cause crashes or zombie processes when spawning multiple workers. Linux and macOS systems typically handle the default multi-process configuration more robustly. Additionally, single-process mode eliminates process-spawning overhead on systems where the gain from parallelism is outweighed by inter-process communication costs.

Summary

  • Multi-process mode trades VRAM for speed, running detection, alignment, and masking concurrently across separate Python processes.
  • Single-process mode minimizes memory usage by executing phases sequentially in one process, essential for GPUs with <8 GB VRAM.
  • The switch is controlled via the -P/--singleprocess flag defined in lib/cli/args_extract_convert.py and implemented in plugins/extract/pipeline.py.
  • Windows users should consider single-process mode for improved stability.
  • The -j parameter allows granular control over concurrency without fully disabling parallelism.

Frequently Asked Questions

Does single-process mode affect output quality?

No, the choice between multi-process and single-process mode impacts only execution speed and resource utilization. The underlying neural network models, alignment algorithms, and output formats remain identical regardless of the processing mode selected.

How much VRAM is required for multi-process mode?

According to the source code documentation in lib/cli/args_extract_convert.py, multi-process mode typically requires GPUs with at least 8 GB of VRAM, though actual consumption varies based on the detection model, input resolution, and whether you are running extraction, training, or conversion pipelines simultaneously.

Can I use multi-process mode on a CPU-only system?

Yes, the backend supports CPU execution, but the performance benefits diminish since CPU-bound tasks often contend for the same compute resources. Single-process mode is frequently more efficient on CPU-only systems, particularly on Windows where multiprocessing overhead can exceed the gains from parallel execution.

Where exactly is the multi-process logic implemented in the codebase?

The primary implementation resides in plugins/extract/pipeline.py, which builds either parallel or sequential processing phases based on the multiprocess argument. The CLI interface is defined in lib/cli/args_extract_convert.py, while scripts/extract.py and scripts/convert.py handle the runtime flag interpretation and pass the boolean to the respective pipeline constructors.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →