How to Configure Batch Processing and Queue Sizes for Optimal Performance in Faceswap

Enable batch mode with the -B flag, set --convert-batchsize based on your GPU VRAM (4–8 for ≤4GB, 8–16 for 8GB, 16–32 for ≥12GB), and cap queue sizes to 2–4× the batch size using -Q to prevent memory bloat while maintaining pipeline throughput.

The Faceswap deep learning pipeline processes thousands of frames through extraction, alignment, and conversion stages. To maximize throughput and prevent bottlenecks when working with the deepfakes/faceswap repository, you must configure batch processing and queue sizes for optimal performance based on your hardware constraints and the specific stage being executed.

Understanding Batch Processing in Faceswap

Batch processing controls how many items are collected before a model is invoked, such as a face-swap or conversion network. By default, Faceswap processes frames individually, which underutilizes GPU parallelism. Enabling batch mode groups frames together for simultaneous inference.

Where Batch Mode is Defined

All high-level tools share the same flag definition across the CLI tools. You can find the -B or --batch-mode argument in files such as tools/alignments/cli.py, tools/mask/cli.py, and tools/sort/cli.py:


# Example from tools/alignments/cli.py

{
    "opts": ("-B", "--batch-mode"),
    "dest": "batch_mode",
    "action": "store_true",
    "help": "Run the job in batch mode (collect many frames before processing)."
}

Enabling Batch Mode via CLI

Add -B (or --batch-mode) when launching any job to activate batch processing. When active, tools internally call read_image_meta_batch from lib/image.py to collect N frames before inference:

python scripts/convert.py -i input_folder -o output_folder -B
python tools/alignments/cli.py -i video.mp4 -B
python tools/mask/cli.py -i aligned_folder -B

Configuring Batch Sizes for GPU Memory

The effective batch size determines how many frames travel through the model simultaneously. This value is constrained by available VRAM and configured through the convert_batchsize setting.

The convert_batchsize Configuration

The maximum batch size per GPU is stored in the training configuration as convert_batchsize. In plugins/train/train_config.py, this is defined as a ConfigItem:


# plugins/train/train_config.py

convert_batchsize = ConfigItem(
    "convert_batchsize",
    default=8,
    doc="Maximum number of frames to feed to the conversion model at once."
)

You can override this at runtime with the --convert-batchsize CLI option exposed by scripts/convert.py.

GPU-Aware Batch Size Calculation

The actual batch size used at runtime is clipped to available device memory. In scripts/convert.py, the _get_batchsize() method implements this logic:


# scripts/convert.py

def _get_batchsize(self, queue_size: int) -> int:
    batchsize = 1 if is_cpu else mod_cfg.convert_batchsize()
    batchsize = min(queue_size, batchsize)
    return batchsize

This ensures the batch size never exceeds the queue size or available VRAM.

Select a --convert-batchsize value that fits comfortably within your GPU memory to avoid out-of-memory errors while maximizing utilization:

  • ≤4 GB VRAM: Use batch size 4–8. This fits comfortably within limited memory and prevents OOM crashes during high-resolution conversions.
  • 8 GB VRAM: Use batch size 8–16. This utilizes most of the available memory without exceeding typical overhead limits.
  • ≥12 GB VRAM: Use batch size 16–32. Larger batches improve GPU utilization and reduce kernel launch overhead on high-end cards.

Set the value in config.ini or pass it on the command line:

python scripts/convert.py -i in -o out -B --convert-batchsize 16

Managing Queue Sizes for Pipeline Throughput

Queue sizes control the maximum number of items that can sit in intermediate buffers before producers block. Proper queue sizing prevents memory bloat while ensuring the GPU never starves for data.

How Queues Work in Faceswap

The pipeline uses named EventQueue instances created via lib/queue_manager._QueueManager.add_queue(). By default, queues are created with maxsize=0 (unlimited), which can consume excessive RAM when consumers are slower than producers.

Key files implementing queue management include:

Setting Queue Size Limits

Most tools expose the -Q or --queue-size CLI flag, which propagates to the internal _queue_size attribute used when calling add_queue():


# Example from scripts/convert.py

queue_manager.add_queue(qname, self._queue_size)  # self._queue_size defaults to 32

Set the queue size at runtime:

python scripts/convert.py -i in -o out -B -Q 64

Optimal Queue Sizing Strategy

Size your queues relative to the batch size to balance memory usage against pipeline stalls:

  • Input reader stage: Set to 2× the batch size. This prevents the GPU from starving while the CPU reads and decodes images.
  • Model inference stage: Set to 1× the batch size. The model consumes batches immediately, so minimal buffering is required.
  • Post-processing stage: Set to 2–4× the batch size. Disk I/O operations can be slower than GPU inference; extra buffering smooths out write spikes.

If you encounter "MemoryError" or "Queue full" warnings, reduce the queue size (-Q) or increase the batch size if VRAM permits. If the GPU utilization drops to zero between batches, increase the queue size to allow the CPU to stay ahead of the GPU.

Complete Configuration Example

The following workflow demonstrates optimal settings for a system with 8 GB VRAM:


# 1️⃣ Extract faces (high-resolution, GPU-heavy)

python scripts/extract.py -i video.mp4 -o faces_dir -B \
    --convert-batchsize 24 -Q 48

# 2️⃣ Align the extracted faces (CPU-bound, can use many workers)

python tools/alignments/cli.py -i faces_dir -o aligned_dir -B \
    --processes 8 -Q 64

# 3️⃣ Convert (swap) using the trained model

python scripts/convert.py -i aligned_dir -o swapped_dir -B \
    --convert-batchsize 16 -Q 32
  • Batch mode (-B) ensures each step processes groups of frames rather than one-by-one.
  • --convert-batchsize tailors the GPU workload to your hardware capabilities.
  • -Q caps each intermediate queue to prevent runaway RAM consumption while maintaining pipeline flow.

Key Source Code References

The following files implement the batch and queue management systems described above:

Summary

  • Enable batch processing using the -B flag to group frames for parallel GPU inference rather than processing them sequentially.
  • Size batches to VRAM by setting --convert-batchsize to 4–8 for ≤4GB GPUs, 8–16 for 8GB GPUs, or 16–32 for ≥12GB GPUs.
  • Bound queue sizes with -Q set to 2–4× the batch size to prevent memory bloat while ensuring the GPU never starves for data.
  • Scale CPU workers using --processes for CPU-bound stages like alignment, independent of GPU batch settings.

Frequently Asked Questions

What is the default batch size in Faceswap?

The default batch size is 1 when running on CPU, or the value specified in convert_batchsize (default 8) when using a GPU. However, the actual value is clipped by the _get_batchsize() method in scripts/convert.py to ensure it fits within available VRAM and does not exceed the queue size.

How do I prevent out-of-memory errors during conversion?

Reduce the --convert-batchsize value to match your GPU's VRAM capacity. For GPUs with 4GB or less, use values between 4 and 8. Additionally, ensure queue sizes are bounded using -Q to prevent unlimited memory growth in intermediate buffers. If errors persist, disable batch mode entirely by omitting the -B flag.

Should I use the same queue size for all pipeline stages?

No, different stages require different queue sizes relative to the batch size. Set input reader queues to 2× the batch size to prevent GPU starvation, model inference queues to 1× the batch size since consumption is immediate, and post-processing queues to 2–4× the batch size to accommodate slower disk I/O operations.

Does batch mode work with CPU-only setups?

Yes, batch mode functions on CPU-only systems, but the batch size defaults to 1 regardless of configuration settings, as seen in the _get_batchsize() logic in scripts/convert.py. CPU processing does not benefit from larger batches in the same way GPUs do, so enabling batch mode primarily helps standardize pipeline behavior across hardware types.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →