How FaceSwap Handles Mask Prediction: Enabling and Using Learned Masks

FaceSwap learns facial masks during training as a fourth output channel and applies them during conversion by enabling the Learn Mask option in training and selecting --mask-type predicted in conversion.

FaceSwap, the open-source deepfake framework maintained by deepfakes/faceswap, includes a powerful mask prediction system that learns to identify facial boundaries during the training process. This FaceSwap mask prediction capability allows the model to generate an alpha channel that seamlessly blends swapped faces into target images. Understanding how to enable and configure this feature is essential for producing high-quality, natural-looking face swaps.

Understanding FaceSwap Mask Prediction Architecture

FaceSwap mask prediction operates as an integrated component of the training and conversion pipeline. The system treats the mask as a fourth channel alongside the standard RGB output, enabling the neural network to learn facial segmentation implicitly through the training process.

Training Phase Mask Learning

During training, the mask prediction capability is controlled through the learn_mask configuration parameter found in plugins/train/train_config.py. When enabled, the trainer incorporates a dedicated mask loss function that teaches the network to predict accurate facial boundaries.

The training interface handles mask visualization through plugins/train/trainer/_display.py, where the _display_mask property controls whether the mask overlay appears in training previews. This allows you to monitor mask quality in real-time as the model learns to distinguish facial features from background elements.

Model Output Structure

When mask learning is enabled, the generator network returns a tensor with shape (H, W, 4) instead of the standard RGB (H, W, 3). The first three channels contain the color information for the swapped face, while the fourth channel contains the predicted mask values.

The MaskLoss class in lib/model/losses/loss.py processes this fourth channel during training, comparing the predicted mask against the ground truth mask derived from alignment data. This loss ensures that the model learns to generate masks that accurately reflect facial boundaries and features.

Enabling FaceSwap Mask Prediction

Activating mask prediction requires configuration changes in both the training and conversion phases. You must first train a model with mask learning enabled, then explicitly request the predicted mask during conversion.

Step 1: Configure Training for Mask Learning

To enable mask prediction, activate the Learn Mask option in the Training GUI's Loss tab. Alternatively, set learn_mask=True in your training configuration file.

You must also specify a Mask Type (such as components or extended) that matches the masks stored in your alignment files. This ensures the training process has ground truth data against which to calculate the mask loss.


# Configuration structure from plugins/train/train_config.py

learn_mask = ConfigItem(
    default=True,
    help=_("Enable mask learning during training."),
)
mask_loss_function = ConfigItem(
    default=LossFunction.MSE,
    help=_("Loss function to use when learning a mask."),
)

Step 2: Convert Using the Predicted Mask

After training completes, invoke the conversion process with the --mask-type predicted flag to utilize the model-generated mask. In the GUI, navigate to the Convert panel, expand Advanced, and set Mask Type to Predicted.

faceswap convert -i INPUT_DIR -o OUTPUT_DIR \
    --model-dir MODEL_DIR \
    --mask-type predicted

The conversion script validates this request in scripts/convert.py. If the model lacks mask prediction capabilities, it logs a warning and falls back to the first available mask:


# Validation from scripts/convert.py

if self._args.mask_type == "predicted" and not self._predictor.has_predicted_mask:
    logger.warning("Selected mask type 'predicted' but model does not output a mask.")

Step 3: Validate Mask Output

During conversion, the Convert class in lib/convert.py extracts the predicted mask and appends it to the face's alpha channel. The _get_image_mask() method handles this extraction, while _pre_warp_adjustments() prepares the mask for blending.


# Mask extraction from lib/convert.py

mask, raw_mask = self._adjustments.mask.run(
    detected_face,
    reference_face.pose.offset[mask_centering],
    reference_face.pose.offset[self._centering],
    self._centering,
    predicted_mask=predicted_mask,
)
new_face = np.concatenate((new_face, mask), -1)   # add as 4th channel

You can verify mask generation using the preview tool in tools/preview/preview.py, which provides a drop-down interface to toggle between mask types and visualize the predicted overlay in real-time.

Technical Implementation Details

The mask prediction system integrates deeply with FaceSwap's neural network architecture. When enabled, the generator network learns to predict masks through auxiliary loss calculations that compare the fourth output channel against ground truth alignment masks.

The MaskLoss implementation in lib/model/losses/loss.py applies the configured loss function (typically MSE) to the mask channel, ensuring that the network learns to distinguish facial regions from background. This loss operates alongside the standard reconstruction losses, creating a multi-objective optimization problem where the network must simultaneously generate realistic faces and accurate segmentation masks.

During conversion, the predicted mask provides superior blending compared to static alignment masks because it adapts to the specific facial features and expressions generated by the model, rather than relying on pre-computed geometric masks.

Summary

  • FaceSwap mask prediction generates an alpha channel during training when Learn Mask is enabled in the training configuration.
  • The model outputs a four-channel tensor (H, W, 4) where the fourth channel contains the predicted mask.
  • Enable mask prediction during conversion using --mask-type predicted in the CLI or selecting Predicted in the GUI's Advanced mask settings.
  • The conversion pipeline in lib/convert.py attaches the predicted mask to the face's alpha channel for seamless blending.
  • If the model lacks mask training, FaceSwap falls back to alignment masks and logs a warning in scripts/convert.py.

Frequently Asked Questions

What is the difference between learned masks and alignment masks in FaceSwap?

Alignment masks are pre-computed geometric masks derived from facial landmarks stored in alignment files, such as components or extended types. Learned masks, conversely, are generated by the neural network during training when the Learn Mask option is enabled, producing a fourth channel that adapts to the specific facial features and expressions the model generates. According to the source code in lib/convert.py, learned masks provide superior blending because they reflect the model's actual output rather than static geometric assumptions.

Can I use mask prediction with any FaceSwap model architecture?

Mask prediction requires that the specific model architecture supports a four-channel output and that you have trained with learn_mask=True enabled. The validation logic in scripts/convert.py explicitly checks self._predictor.has_predicted_mask when --mask-type predicted is requested. If this property returns false, the system logs a warning and falls back to the first available alignment mask. Therefore, you must verify that your specific model plugin supports mask learning before attempting to use predicted masks in conversion.

Why does FaceSwap warn that the model does not output a mask when using --mask-type predicted?

This warning occurs when you request --mask-type predicted in the CLI (or select Predicted in the GUI) but the model checkpoint was not trained with the Learn Mask option enabled. As implemented in scripts/convert.py, the code checks self._predictor.has_predicted_mask and emits the warning: "Selected mask type 'predicted' but model does not output a mask." To resolve this, you must retrain your model with learn_mask=True in the training configuration, ensuring the network learns to output the fourth channel containing the predicted mask.

How do I visualize the predicted mask during conversion?

You can visualize the predicted mask using the preview tool located in tools/preview/preview.py. During conversion setup, the GUI provides a drop-down interface in the preview window (controlled by tools/preview/control_panels.py) that allows you to select Predicted from the mask type options. This overlays the fourth channel output onto the face visualization in real-time, letting you verify mask quality before running the full conversion batch. In the CLI workflow, you can inspect the alpha channel of output images using standard image editing tools to verify the predicted mask has been correctly attached.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →