# How the SUMMON Stage Handles Large Models on Multiple GPUs in OBLITERATUS

> Learn how OBLITERATUS SUMMON stage efficiently distributes large models across multiple GPUs using Hugging Face Accelerate, optimizing memory and performance for your AI projects.

- Repository: [pliny/OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS)
- Tags: internals
- Published: 2026-08-22

---

**The SUMMON stage automatically distributes large language models across multiple GPUs by detecting available CUDA devices, calculating memory requirements, and utilizing Hugging Face Accelerate's device_map to shard model layers while respecting per-GPU memory limits.**

The OBLITERATUS pipeline begins with the SUMMON stage, which is responsible for loading models and tokenizers into memory. When working with large models that exceed the capacity of a single GPU, this stage implements sophisticated sharding strategies to distribute workloads across available hardware according to the source code in the elder-plinius/OBLITERATUS repository.

## GPU Detection and Telemetry

Before loading begins, the SUMMON stage queries the system hardware to determine available compute resources. The `_detect_gpu()` function in [`obliteratus/telemetry.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/telemetry.py) scans the runtime for CUDA-capable devices and returns a structured list containing each GPU's index, name, and available VRAM.

This detection mechanism enables the pipeline to make informed decisions about model placement. The function returns data structures such as `[{index:0, name:"A100", vram_gb:40}, ...]`, which feeds directly into the device selection logic. According to the source code around line 45 of [`obliteratus/telemetry.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/telemetry.py), this telemetry data forms the foundation for all subsequent memory calculations and device mapping decisions.

## Memory Estimation and Device Mapping

The core orchestration occurs in `AbliterationPipeline._summon()` within [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py) (around line 406). This method initiates the loading sequence by calling `load_model()` with the user-specified `gpu_memory_utilization` parameter, which defaults to 0.9 to maintain a safety margin against out-of-memory errors.

The actual sharding logic resides in [`obliteratus/models/loader.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/models/loader.py) within the `load_model()` function (approximately line 630). This implementation performs several critical operations:

- **Parameter Counting**: Calculates the total memory required for model weights based on the specified `torch_dtype`
- **Budget Allocation**: Compares model requirements against available VRAM across detected GPUs while respecting `gpu_memory_utilization` limits
- **Device Map Construction**: Builds a balanced mapping that assigns specific transformer layers to different GPUs to maximize parallel utilization

When a model fits within a single GPU's budget, the system uses `device_map="auto"` or the explicitly specified device string. For models exceeding single-GPU capacity, the loader constructs a custom device map that distributes layers across all available GPUs in the detected set.

## Distributed Loading with Accelerate

The actual model instantiation leverages the Hugging Face ecosystem for reliable multi-GPU support. The `load_model()` function invokes `transformers.PretrainedModel.from_pretrained()` with specific parameters enabling automatic sharding:

- `device_map=device_map`: Specifies the layer-to-GPU allocation strategy calculated earlier
- `offload_folder=tmp_dir`: Provides a fallback location for layers that cannot fit in GPU memory, enabling CPU or disk offloading
- `max_memory`: Enforces per-GPU memory constraints to prevent allocation beyond available VRAM

This approach utilizes the **accelerate** library under the hood, which handles the low-level details of splitting model weights across devices without requiring manual `torch.nn.DataParallel` or `torch.distributed` configuration. After loading completes, the pipeline logs the GPU allocation summary via `self._emit("summon", ...)` so the interface can display which GPUs hold which portions of the model.

## Configuration and CLI Integration

Users control multi-GPU behavior through arguments parsed in [`obliteratus/cli.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/cli.py) (around line 210). The `--device` or `--gpus` parameters allow explicit specification of which GPUs to utilize, while `--gpu-memory-utilization` adjusts the safety margin for memory allocation.

When `device="auto"` is specified (the default), the pipeline automatically utilizes all detected CUDA devices. Alternatively, users can restrict processing to specific GPUs using comma-separated identifiers such as `device="cuda:0,1"` to limit the cluster to the first two GPUs only, while still allowing the SUMMON stage to shard the model across that subset.

## Practical Implementation Examples

The following examples demonstrate how to configure the SUMMON stage for various multi-GPU scenarios:

```python

# Automatic multi-GPU detection and sharding

from obliteratus import AbliterationPipeline

pipeline = AbliterationPipeline(
    model_name="meta-llama/Llama-3.1-70B-Instruct",
    device="auto",                # Detect and use all available GPUs

    gpu_memory_utilization=0.85,  # Reserve 15% VRAM buffer

)
pipeline.run()  # SUMMON distributes model across GPUs automatically

```

```python

# Manual GPU selection for specific hardware constraints

from obliteratus import AbliterationPipeline

pipeline = AbliterationPipeline(
    model_name="meta-llama/Llama-3.1-70B-Instruct",
    device="cuda:0,1",           # Restrict to GPU 0 and GPU 1 only

    gpu_memory_utilization=0.90,
)
pipeline.run()

```

Both configurations trigger the same internal loading flow described above. The second example simply narrows the device list that `load_model()` can allocate to, ensuring predictable resource usage on shared multi-tenant systems.

## Summary

- The SUMMON stage automatically detects all CUDA GPUs and queries their available VRAM using `_detect_gpu()` in [`obliteratus/telemetry.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/telemetry.py) (around line 45)
- Memory requirements are calculated against available resources in [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py) via the `_summon()` method (around line 406), respecting the `gpu_memory_utilization` safety margin
- Model sharding is handled by `load_model()` in [`obliteratus/models/loader.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/models/loader.py) (around line 630), which constructs a balanced `device_map` for distributing layers across GPUs
- Actual loading utilizes Hugging Face Transformers with the Accelerate library via `from_pretrained()`, supporting automatic offloading to CPU or disk when GPU memory is insufficient
- Users control GPU selection through CLI arguments parsed in [`obliteratus/cli.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/cli.py) (around line 210), supporting both automatic detection and manual device specification

## Frequently Asked Questions

### How does the SUMMON stage decide whether to use multiple GPUs?

The SUMMON stage compares the estimated memory footprint of the target model against the available VRAM of each detected GPU. If the model exceeds the capacity of any single GPU while respecting the `gpu_memory_utilization` threshold, `load_model()` automatically constructs a `device_map` that splits the model across all available GPUs to distribute the memory load according to the balancing logic in [`obliteratus/models/loader.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/models/loader.py).

### Can I restrict the SUMMON stage to specific GPUs on my system?

Yes. While the default `device="auto"` setting utilizes all detected CUDA devices, you can specify exact GPUs using the device parameter. Passing `device="cuda:0,1"` (parsed in [`obliteratus/cli.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/cli.py) around line 210) limits the pipeline to GPUs 0 and 1 only, allowing you to reserve other GPUs for different workloads while still enabling multi-GPU sharding across the selected subset.

### What happens if my model is too large for all my GPUs combined?

When the model exceeds the total available VRAM across all specified GPUs, the SUMMON stage activates CPU and disk offloading mechanisms. The `load_model()` function passes an `offload_folder` parameter to `from_pretrained()`, allowing the Hugging Face Accelerate library to move excess layers to system RAM or temporary disk storage. This enables loading models larger than your total GPU memory, though with reduced inference speed due to CPU-GPU transfer overhead.

### Does the SUMMON stage support manual sharding configurations?

The current implementation automatically calculates the optimal `device_map` based on layer sizes and available memory. While the source code in [`obliteratus/models/loader.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/models/loader.py) handles the balancing logic internally around line 630, the architecture supports the underlying Accelerate library's device mapping capabilities. Advanced users can modify the `load_model()` function to pass custom device maps if specific layer placement is required for their hardware topology.