How the SUMMON Stage Handles Large Models on Multiple GPUs in OBLITERATUS

The SUMMON stage automatically distributes large language models across multiple GPUs by detecting available CUDA devices, calculating memory requirements, and utilizing Hugging Face Accelerate's device_map to shard model layers while respecting per-GPU memory limits.

The OBLITERATUS pipeline begins with the SUMMON stage, which is responsible for loading models and tokenizers into memory. When working with large models that exceed the capacity of a single GPU, this stage implements sophisticated sharding strategies to distribute workloads across available hardware according to the source code in the elder-plinius/OBLITERATUS repository.

GPU Detection and Telemetry

Before loading begins, the SUMMON stage queries the system hardware to determine available compute resources. The _detect_gpu() function in obliteratus/telemetry.py scans the runtime for CUDA-capable devices and returns a structured list containing each GPU's index, name, and available VRAM.

This detection mechanism enables the pipeline to make informed decisions about model placement. The function returns data structures such as [{index:0, name:"A100", vram_gb:40}, ...], which feeds directly into the device selection logic. According to the source code around line 45 of obliteratus/telemetry.py, this telemetry data forms the foundation for all subsequent memory calculations and device mapping decisions.

Memory Estimation and Device Mapping

The core orchestration occurs in AbliterationPipeline._summon() within obliteratus/abliterate.py (around line 406). This method initiates the loading sequence by calling load_model() with the user-specified gpu_memory_utilization parameter, which defaults to 0.9 to maintain a safety margin against out-of-memory errors.

The actual sharding logic resides in obliteratus/models/loader.py within the load_model() function (approximately line 630). This implementation performs several critical operations:

  • Parameter Counting: Calculates the total memory required for model weights based on the specified torch_dtype
  • Budget Allocation: Compares model requirements against available VRAM across detected GPUs while respecting gpu_memory_utilization limits
  • Device Map Construction: Builds a balanced mapping that assigns specific transformer layers to different GPUs to maximize parallel utilization

When a model fits within a single GPU's budget, the system uses device_map="auto" or the explicitly specified device string. For models exceeding single-GPU capacity, the loader constructs a custom device map that distributes layers across all available GPUs in the detected set.

Distributed Loading with Accelerate

The actual model instantiation leverages the Hugging Face ecosystem for reliable multi-GPU support. The load_model() function invokes transformers.PretrainedModel.from_pretrained() with specific parameters enabling automatic sharding:

  • device_map=device_map: Specifies the layer-to-GPU allocation strategy calculated earlier
  • offload_folder=tmp_dir: Provides a fallback location for layers that cannot fit in GPU memory, enabling CPU or disk offloading
  • max_memory: Enforces per-GPU memory constraints to prevent allocation beyond available VRAM

This approach utilizes the accelerate library under the hood, which handles the low-level details of splitting model weights across devices without requiring manual torch.nn.DataParallel or torch.distributed configuration. After loading completes, the pipeline logs the GPU allocation summary via self._emit("summon", ...) so the interface can display which GPUs hold which portions of the model.

Configuration and CLI Integration

Users control multi-GPU behavior through arguments parsed in obliteratus/cli.py (around line 210). The --device or --gpus parameters allow explicit specification of which GPUs to utilize, while --gpu-memory-utilization adjusts the safety margin for memory allocation.

When device="auto" is specified (the default), the pipeline automatically utilizes all detected CUDA devices. Alternatively, users can restrict processing to specific GPUs using comma-separated identifiers such as device="cuda:0,1" to limit the cluster to the first two GPUs only, while still allowing the SUMMON stage to shard the model across that subset.

Practical Implementation Examples

The following examples demonstrate how to configure the SUMMON stage for various multi-GPU scenarios:


# Automatic multi-GPU detection and sharding

from obliteratus import AbliterationPipeline

pipeline = AbliterationPipeline(
    model_name="meta-llama/Llama-3.1-70B-Instruct",
    device="auto",                # Detect and use all available GPUs

    gpu_memory_utilization=0.85,  # Reserve 15% VRAM buffer

)
pipeline.run()  # SUMMON distributes model across GPUs automatically

# Manual GPU selection for specific hardware constraints

from obliteratus import AbliterationPipeline

pipeline = AbliterationPipeline(
    model_name="meta-llama/Llama-3.1-70B-Instruct",
    device="cuda:0,1",           # Restrict to GPU 0 and GPU 1 only

    gpu_memory_utilization=0.90,
)
pipeline.run()

Both configurations trigger the same internal loading flow described above. The second example simply narrows the device list that load_model() can allocate to, ensuring predictable resource usage on shared multi-tenant systems.

Summary

  • The SUMMON stage automatically detects all CUDA GPUs and queries their available VRAM using _detect_gpu() in obliteratus/telemetry.py (around line 45)
  • Memory requirements are calculated against available resources in obliteratus/abliterate.py via the _summon() method (around line 406), respecting the gpu_memory_utilization safety margin
  • Model sharding is handled by load_model() in obliteratus/models/loader.py (around line 630), which constructs a balanced device_map for distributing layers across GPUs
  • Actual loading utilizes Hugging Face Transformers with the Accelerate library via from_pretrained(), supporting automatic offloading to CPU or disk when GPU memory is insufficient
  • Users control GPU selection through CLI arguments parsed in obliteratus/cli.py (around line 210), supporting both automatic detection and manual device specification

Frequently Asked Questions

How does the SUMMON stage decide whether to use multiple GPUs?

The SUMMON stage compares the estimated memory footprint of the target model against the available VRAM of each detected GPU. If the model exceeds the capacity of any single GPU while respecting the gpu_memory_utilization threshold, load_model() automatically constructs a device_map that splits the model across all available GPUs to distribute the memory load according to the balancing logic in obliteratus/models/loader.py.

Can I restrict the SUMMON stage to specific GPUs on my system?

Yes. While the default device="auto" setting utilizes all detected CUDA devices, you can specify exact GPUs using the device parameter. Passing device="cuda:0,1" (parsed in obliteratus/cli.py around line 210) limits the pipeline to GPUs 0 and 1 only, allowing you to reserve other GPUs for different workloads while still enabling multi-GPU sharding across the selected subset.

What happens if my model is too large for all my GPUs combined?

When the model exceeds the total available VRAM across all specified GPUs, the SUMMON stage activates CPU and disk offloading mechanisms. The load_model() function passes an offload_folder parameter to from_pretrained(), allowing the Hugging Face Accelerate library to move excess layers to system RAM or temporary disk storage. This enables loading models larger than your total GPU memory, though with reduced inference speed due to CPU-GPU transfer overhead.

Does the SUMMON stage support manual sharding configurations?

The current implementation automatically calculates the optimal device_map based on layer sizes and available memory. While the source code in obliteratus/models/loader.py handles the balancing logic internally around line 630, the architecture supports the underlying Accelerate library's device mapping capabilities. Advanced users can modify the load_model() function to pass custom device maps if specific layer placement is required for their hardware topology.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →