# How to Optimize Power Consumption and Thermal Management with the DS4 `--power` Flag

> Optimize power consumption and thermal management with the DS4 --power flag. Cap GPU utilization between 1-99% to reduce power draw and heat. Learn how to implement.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-08

---

**The `--power` flag enables GPU duty-cycle throttling by inserting calculated sleep intervals after each prefill layer and decode token, allowing you to cap GPU utilization between 1-99% to reduce power draw and thermal output.**

The antirez/ds4 inference engine provides a built-in mechanism to limit GPU utilization through the `--power` command-line option. Available across all DS4 executables—including `ds4`, `ds4-server`, `ds4-agent`, and `ds4-bench`—this flag lets you balance inference throughput against energy consumption and heat generation without modifying model architecture or quantization settings.

## How the `--power` Flag Works Internally

When you pass `--power N` (where N is 1-99), the runtime parses this value into the `engine.power_percent` field. In [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c) (lines 1918-1923) and [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) (lines 12946-12951), the argument handler converts the string input to an unsigned integer stored in the engine configuration structure defined in [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c).

### The Duty-Cycle Throttling Algorithm

The core logic resides in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c). Before applying any throttling, the runtime checks `graph_power_throttle_enabled` to confirm the setting is active:

```c
static bool graph_power_throttle_enabled(const ds4_gpu_graph *g) {
    return g && g->power_percent > 0 && g->power_percent < 100;
}

```

When enabled, the system maintains an exponential moving average of execution times and enforces a specific duty cycle. The target relationship follows this formula:

**Target duty cycle** → `work / (work + sleep) = power / 100`

The sleep duration calculation in `graph_power_sleep` (lines 15645-15653) implements this ratio:

```c
static void graph_power_sleep(double work_sec, uint32_t power_percent) {
    if (power_percent == 0 || power_percent >= 100) return;
    const double sleep = work_sec * (100.0 - (double)power_percent) /
                         (double)power_percent;
    sleep_sec(sleep);
}

```

### Prefill Layer Throttling

During the prefill phase, `graph_power_note_prefill_layer` (lines 15558-15563) records the elapsed time for each layer, updates the moving average for that specific layer index, and invokes the sleep function:

```c
static void graph_power_note_prefill_layer(ds4_gpu_graph *g,
                                           uint32_t il,
                                           double elapsed_sec) {
    if (!graph_power_throttle_enabled(g)) return;
    if (il >= DS4_N_LAYER) return;
    g->prefill_layer_avg_sec[il] =
        graph_power_update_avg(g->prefill_layer_avg_sec[il], elapsed_sec);
    graph_power_sleep(g->prefill_layer_avg_sec[il], g->power_percent);
}

```

### Decode Token Throttling

Similarly, during token generation, `graph_power_note_decode_token` (lines 16565-16570) measures decode latency and applies throttling:

```c
static void graph_power_note_decode_token(ds4_gpu_graph *g, double elapsed_sec) {
    if (!graph_power_throttle_enabled(g)) return;
    g->decode_token_avg_sec =
        graph_power_update_avg(g->decode_token_avg_sec, elapsed_sec);
    graph_power_sleep(g->decode_token_avg_sec, g->power_percent);
}

```

## Optimizing Power Consumption Across DS4 Executables

The `--power` flag works uniformly across the DS4 ecosystem. Because the sleep is inserted **after each measured work chunk**, the impact scales with the actual workload: long-running kernels see proportionally longer sleeps, while short kernels still respect the duty-cycle.

### Running the Server with Reduced Power

For production deployments where thermal constraints matter:

```bash
ds4-server --host 127.0.0.1 --port 8000 --power 60

```

This configuration caps the server to approximately 60% GPU duty-cycle, automatically inserting sleeps to ensure the GPU is active only 60% of the time.

### Configuring Edge Devices with ds4-agent

On resource-constrained edge devices where the GPU is shared with other processes:

```bash
ds4-agent --power 50 --model ./my-model.gguf

```

Setting `--power 50` ensures the GPU spends equal time working and sleeping, roughly halving power consumption compared to full utilization.

### Benchmarking Power vs. Performance

Use `ds4-bench` to measure the throughput impact of different power settings:

```bash
ds4-bench --prompt-file long.txt --power 70

```

This allows you to quantify the latency trade-off when optimizing for thermal management.

### Combining with Memory Optimization

You can stack `--power` with other resource-limiting flags for comprehensive power savings:

```bash
ds4 --power 40 --ssd-streaming --threads 4

```

This reduces GPU duty-cycle while streaming the model from SSD, further lowering overall system power draw.

## Summary

- The `--power` flag sets `engine.power_percent` in DS4, parsed in [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c) (lines 1918-1923) and [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) (lines 12946-12951).
- Throttling applies to both **prefill layers** (`graph_power_note_prefill_layer`) and **decode tokens** (`graph_power_note_decode_token`) via exponential moving averages.
- The algorithm inserts proportional sleeps to achieve the ratio `work/(work+sleep) = power/100`.
- Valid values range from 1-99; 100 (default) disables throttling entirely.
- Lower values significantly reduce thermal output for edge devices, while higher values maximize throughput for latency-sensitive applications.

## Frequently Asked Questions

### What is the default power setting in DS4?

By default, DS4 runs with `--power 100` (or 0, which is treated as disabled), meaning no artificial throttling is applied. The GPU runs at maximum duty-cycle, providing the lowest latency but highest power consumption. You can verify this default behavior in [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c) (lines 1669-1670).

### How does the `--power` flag affect inference latency?

Higher sleep proportions directly increase latency. At `--power 50`, the runtime inserts sleep equal to the work duration, effectively doubling the time per prefill layer and decode token. The exact impact depends on kernel execution time, which the exponential moving average tracks continuously in `graph_power_sleep`.

### Can I use `--power` with multiple GPU contexts?

The `--power` setting applies per `ds4_gpu_graph` instance. If you run multiple DS4 processes or threads, each respects its own `engine.power_percent` value. There is no global GPU lock; throttling is cooperative and process-specific.

### Does power throttling impact model accuracy?

No. The `--power` flag only controls the duty-cycle timing through `sleep_sec()` calls in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c). It does not modify quantization, attention mechanisms, or numerical precision. The model produces identical outputs; only the speed of computation changes.