How the Model Backup and Restore System Works in Faceswap

The Faceswap model backup and restore system automatically creates .bk file backups when training loss improves and provides CLI tools to restore corrupted models or create full snapshots.

The deepfakes/faceswap repository implements a comprehensive model backup and restore system to protect trained models from corruption and allow users to roll back to known-good states. At the heart of this system is the Backup class located in lib/model/backup_restore.py, which integrates with the training pipeline in plugins/train/model/_base/io.py and the Model tool in tools/model/model.py.

Core Architecture of the Backup System

The backup and restore functionality is distributed across three primary components that handle automatic protection during training and manual recovery operations.

The Backup Class

The Backup class in lib/model/backup_restore.py contains the core logic for file-level backups, snapshot creation, and restoration. It manages backup files by appending the .bk extension to model files and handles archiving of corrupted states during restore operations.

Integration with Training IO

The IO class in plugins/train/model/_base/io.py instantiates the backup system during model initialization:

self._backup = Backup(self._model_dir, self._plugin.name)

This integration point triggers automatic backups during the training lifecycle and manages snapshot creation at specific iteration intervals.

Automatic Backup Workflow During Training

The system creates protective backups automatically during training based on model performance metrics rather than time intervals.

Loss-Based Backup Decisions

When IO.save() is called after each training iteration, it invokes _maybe_backup() only if IO._should_backup() determines that the current average loss is lower than any previously recorded value. The method compares the current loss against self._plugin.state.lowest_avg_loss to ensure backups capture only improving model states.

File-Level Backup Creation

When the loss condition is met, the system executes Backup.backup_model() which copies the current .keras model file and the associated state file to backup versions with the .bk extension. This operation preserves the model weights and training state before any new save operation can overwrite them.

Creating Manual Snapshots

Beyond automatic backups, the system supports full directory snapshots for major milestones. The IO.snapshot() method calls Backup.snapshot_models() to create a complete copy of the entire model folder under a name following the pattern modeldir_snapshot_<iters>_iters. This captures the complete training state at a specific iteration count, useful before hyperparameter changes or experimental modifications.

Restoring Models from Backup

The Model tool (tools/model/model.py) provides the restore command to recover from corrupted training states. The Restore.process() method instantiates a Backup object and calls its restore() method, which executes a three-phase recovery process.

Archive Current State

The restore operation first moves all current model files (except logs) into a timestamped archive folder named <model_name>_archived_<timestamp>. This preserves the corrupted state for forensic analysis while clearing the workspace for restoration.

Replace with Backups

The system then copies each *.bk file back to its original name by removing the .bk suffix, effectively rolling back the model weights and state files to their last backed-up versions.

Restore Logs

Finally, the operation recreates log folders that existed up to the last backup point, ensuring that training history and tensorboard data remain consistent with the restored model state.

Command Line Usage Examples

The model backup and restore system exposes three primary operations through the Model tool CLI:


# Automatic backups occur during training - no manual intervention needed

python3 scripts/train.py --model-dir ./my_model ...

# Create a manual snapshot before major hyperparameter changes

python3 tools/model/model.py --model-dir ./my_model --job snapshot

# Restore the most recent backup after corruption or crashes

python3 tools/model/model.py --model-dir ./my_model --job restore

The Model tool dispatches these commands through a job selection mechanism in tools/model/model.py:

jobs = {"inference": Inference, "nan-scan": NaNScan, "restore": Restore}
return jobs[arguments.job](arguments)

Summary

  • The model backup and restore system centers on the Backup class in lib/model/backup_restore.py, which manages .bk file creation and restoration workflows.
  • Automatic backups trigger only when training loss improves, protecting good models from being overwritten by corrupted states during the save process.
  • Snapshots create complete directory copies at specific iterations, providing full checkpoints independent of the incremental backup system.
  • Restoration archives the current corrupted state, replaces active files with .bk backups, and rebuilds log directories to maintain training history consistency.
  • The Model tool CLI exposes snapshot and restore functionality through tools/model/model.py for manual intervention when automatic protection is insufficient.

Frequently Asked Questions

When does Faceswap create an automatic backup?

Faceswap creates automatic backups only when the current training iteration achieves a lower average loss than any previous save, as determined by IO._should_backup() comparing against self._plugin.state.lowest_avg_loss. This loss-based approach ensures that backup files always represent the best-performing model state rather than creating backups at fixed time intervals.

What files are included in a model restore?

The restore process handles three categories of files: model weights (.keras files), training state files, and log directories. The system archives existing files (excluding logs), restores .bk backup files to their original names, and recreates log folders that existed at the time of the last backup, ensuring complete recovery of the training environment.

How do I create a manual checkpoint before changing hyperparameters?

Run the snapshot command provided by the Model tool: python3 tools/model/model.py --model-dir ./my_model --job snapshot. This invokes Backup.snapshot_models() to create a complete copy of your model directory named modeldir_snapshot_<iters>_iters, preserving the exact state before experimental changes without disrupting the automatic backup workflow.

What's the difference between a backup and a snapshot in Faceswap?

Backups are incremental file copies (with .bk extensions) created automatically when loss improves, protecting individual model files from corruption during saves. Snapshots are complete directory copies created manually or at specific iterations that capture the entire model folder state, providing full checkpoints for major milestones or before significant configuration changes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →