How to Configure Multi-GPU and Device Mapping for Heretic: A Complete Guide

Configure multi-GPU and device mapping for Heretic by setting device_map and max_memory in the Settings model, accessible via TOML config files, CLI flags, or environment variables.

Heretic, the open-source evaluation framework hosted at p-e-w/heretic, supports distributed inference across multiple GPUs and accelerators through tight integration with Hugging Face Accelerate. By configuring device mapping (which layers reside on which GPU) and memory limits (per-device VRAM caps), you can optimize throughput for models that exceed single-GPU memory or balance workloads across heterogeneous hardware.

Understanding Device Mapping and Memory Management in Heretic

The Settings Model

All multi-GPU configuration in Heretic flows through the Settings class defined in src/heretic/config.py. Two fields control device placement:

  • device_map (lines 103-106): Accepts "auto" to let Accelerate calculate optimal placement, or an explicit dictionary mapping layer indices to device IDs (e.g., {"0": 0, "1": 1}).
  • max_memory (lines 108-111): Optional per-device memory constraints formatted as strings (e.g., {"0": "20GB", "1": "22GB"}). Accelerate uses these limits when partitioning model layers.

How Accelerate Handles Device Placement

When Heretic loads a model, it passes your device_map and max_memory values directly to accelerate.init_empty_weights() and accelerate.load_checkpoint_and_dispatch(). This happens transparently—no manual tensor movement required. The startup logic in src/heretic/main.py (lines 176-188) enumerates available CUDA, XPU, MLU, SDAA, MUSA, and MPS devices, printing their names and VRAM to help you construct valid device maps.

Configuration Methods for Multi-GPU Setups

Heretic supports three hierarchical configuration layers: environment variables (lowest priority), TOML config files, and CLI flags (highest priority).

Using a TOML Configuration File

Create a config.toml in the repository root for persistent multi-GPU settings:


# config.toml

device_map = { "0" = 0, "1" = 1 }
max_memory = { "0" = "22GB", "1" = "14GB" }

This configuration splits layers between GPU 0 and GPU 1 while capping memory usage at 22 GB and 14 GB respectively. The Settings.settings_customise_sources method in src/heretic/config.py automatically loads this file on startup.

Command-Line Interface Overrides

Override config files or set device mapping ad-hoc using JSON-formatted CLI arguments:

heretic run \
  --device_map '{"0":0,"1":1}' \
  --max_memory '{"0":"22GB","1":"14GB"}'

Heretic parses these strings and injects them into the Settings model before model initialization.

Environment Variable Configuration

For containerized deployments or quick experiments, use HERETIC_ prefixed environment variables:

export HERETIC_DEVICE_MAP='{"0":0,"1":1}'
export HERETIC_MAX_MEMORY='{"0":"22GB","1":"14GB"}'
heretic run

Environment variables are processed after CLI flags but before TOML files, allowing flexible layering of configurations.

Practical Multi-GPU Configuration Examples

Balancing Unequal GPUs

Consider a system with two NVIDIA GPUs: GPU 0 has 24 GB VRAM, GPU 1 has 16 GB VRAM. To maximize throughput without out-of-memory errors:

device_map = { "0" = 0, "1" = 1 }
max_memory = { "0" = "22GB", "1" = "14GB" }

The max_memory values leave headroom for activation tensors and system overhead. Heretic's utility functions in src/heretic/utils.py (specifically get_device_stats()) report actual VRAM utilization during runtime to help you fine-tune these caps.

Extending to Three or More GPUs

Adding a third GPU requires only extending the dictionaries:

device_map = { "0" = 0, "1" = 1, "2" = 2 }
max_memory = { "0" = "22GB", "1" = "14GB", "2" = "16GB" }

Heretic automatically detects the additional device during startup (as implemented in src/heretic/main.py lines 176-188) and validates that your device map indices exist in the enumerated hardware list.

Verifying Your Device Mapping Configuration

After configuring, confirm correct placement by observing the startup logs. Heretic prints enumerated devices in src/heretic/main.py (lines 176-188), showing:


Detected devices:
  cuda:0: NVIDIA GeForce RTX 3090 (24GB)
  cuda:1: NVIDIA GeForce RTX 3080 (16GB)

If your device_map references non-existent indices, Heretic raises a validation error before loading weights, preventing silent CPU fallback or out-of-memory crashes.

Summary

  • Configure multi-GPU and device mapping for Heretic through the Settings model in src/heretic/config.py, specifically the device_map and max_memory fields.
  • Use TOML files for persistent configurations, CLI flags for ad-hoc overrides, or environment variables for containerized deployments.
  • Reference actual GPU indices detected by the startup logic in src/heretic/main.py (lines 176-188) when building explicit device maps.
  • Set conservative max_memory values to leave headroom for activations and system overhead, especially when balancing heterogeneous GPUs.

Frequently Asked Questions

What is the difference between device_map and max_memory in Heretic?

The device_map parameter controls which layers are placed on which device, accepting either "auto" for automatic distribution or a dictionary mapping layer names to device indices. The max_memory parameter sets per-device VRAM caps (e.g., {"0": "20GB"}), instructing Accelerate to respect these limits when partitioning the model. Use device_map for explicit placement control and max_memory to prevent out-of-memory errors on GPUs with limited VRAM.

Can I use automatic device mapping instead of manual configuration?

Yes. Set device_map = "auto" in your config.toml or pass --device_map '"auto"' via CLI. When set to "auto", Heretic delegates placement decisions to Hugging Face Accelerate, which analyzes layer sizes and available VRAM to generate an optimal mapping. This is ideal for homogeneous multi-GPU setups where you want maximum memory utilization without manual tuning.

How do I check if Heretic detected my GPUs correctly?

Heretic enumerates all available accelerators during startup in src/heretic/main.py (lines 176-188). The startup log displays a "Detected devices" list showing device indices (e.g., cuda:0), model names, and VRAM capacity. If a GPU is missing from this list, verify PyTorch installation (torch.cuda.is_available()) and driver compatibility. Heretic validates device_map indices against this detected list, raising errors for invalid references before model loading begins.

What file contains the device detection logic in Heretic?

The device detection and reporting logic resides in src/heretic/main.py, specifically lines 176-188. This code iterates through supported backends (CUDA, XPU, MLU, SDAA, MUSA, MPS), retrieves device properties, and prints a formatted summary of available hardware. For runtime VRAM statistics aggregation, see src/heretic/utils.py, which contains get_device_stats() used by the CLI to display memory utilization during inference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →