How to Optimize Power Consumption and Thermal Management with the DS4 `--power` Flag

The --power flag enables GPU duty-cycle throttling by inserting calculated sleep intervals after each prefill layer and decode token, allowing you to cap GPU utilization between 1-99% to reduce power draw and thermal output.

The antirez/ds4 inference engine provides a built-in mechanism to limit GPU utilization through the --power command-line option. Available across all DS4 executables—including ds4, ds4-server, ds4-agent, and ds4-bench—this flag lets you balance inference throughput against energy consumption and heat generation without modifying model architecture or quantization settings.

How the --power Flag Works Internally

When you pass --power N (where N is 1-99), the runtime parses this value into the engine.power_percent field. In ds4_cli.c (lines 1918-1923) and ds4_server.c (lines 12946-12951), the argument handler converts the string input to an unsigned integer stored in the engine configuration structure defined in ds4_gpu_args.c.

The Duty-Cycle Throttling Algorithm

The core logic resides in ds4.c. Before applying any throttling, the runtime checks graph_power_throttle_enabled to confirm the setting is active:

static bool graph_power_throttle_enabled(const ds4_gpu_graph *g) {
    return g && g->power_percent > 0 && g->power_percent < 100;
}

When enabled, the system maintains an exponential moving average of execution times and enforces a specific duty cycle. The target relationship follows this formula:

Target duty cycle → work / (work + sleep) = power / 100

The sleep duration calculation in graph_power_sleep (lines 15645-15653) implements this ratio:

static void graph_power_sleep(double work_sec, uint32_t power_percent) {
    if (power_percent == 0 || power_percent >= 100) return;
    const double sleep = work_sec * (100.0 - (double)power_percent) /
                         (double)power_percent;
    sleep_sec(sleep);
}

Prefill Layer Throttling

During the prefill phase, graph_power_note_prefill_layer (lines 15558-15563) records the elapsed time for each layer, updates the moving average for that specific layer index, and invokes the sleep function:

static void graph_power_note_prefill_layer(ds4_gpu_graph *g,
                                           uint32_t il,
                                           double elapsed_sec) {
    if (!graph_power_throttle_enabled(g)) return;
    if (il >= DS4_N_LAYER) return;
    g->prefill_layer_avg_sec[il] =
        graph_power_update_avg(g->prefill_layer_avg_sec[il], elapsed_sec);
    graph_power_sleep(g->prefill_layer_avg_sec[il], g->power_percent);
}

Decode Token Throttling

Similarly, during token generation, graph_power_note_decode_token (lines 16565-16570) measures decode latency and applies throttling:

static void graph_power_note_decode_token(ds4_gpu_graph *g, double elapsed_sec) {
    if (!graph_power_throttle_enabled(g)) return;
    g->decode_token_avg_sec =
        graph_power_update_avg(g->decode_token_avg_sec, elapsed_sec);
    graph_power_sleep(g->decode_token_avg_sec, g->power_percent);
}

Optimizing Power Consumption Across DS4 Executables

The --power flag works uniformly across the DS4 ecosystem. Because the sleep is inserted after each measured work chunk, the impact scales with the actual workload: long-running kernels see proportionally longer sleeps, while short kernels still respect the duty-cycle.

Running the Server with Reduced Power

For production deployments where thermal constraints matter:

ds4-server --host 127.0.0.1 --port 8000 --power 60

This configuration caps the server to approximately 60% GPU duty-cycle, automatically inserting sleeps to ensure the GPU is active only 60% of the time.

Configuring Edge Devices with ds4-agent

On resource-constrained edge devices where the GPU is shared with other processes:

ds4-agent --power 50 --model ./my-model.gguf

Setting --power 50 ensures the GPU spends equal time working and sleeping, roughly halving power consumption compared to full utilization.

Benchmarking Power vs. Performance

Use ds4-bench to measure the throughput impact of different power settings:

ds4-bench --prompt-file long.txt --power 70

This allows you to quantify the latency trade-off when optimizing for thermal management.

Combining with Memory Optimization

You can stack --power with other resource-limiting flags for comprehensive power savings:

ds4 --power 40 --ssd-streaming --threads 4

This reduces GPU duty-cycle while streaming the model from SSD, further lowering overall system power draw.

Summary

  • The --power flag sets engine.power_percent in DS4, parsed in ds4_cli.c (lines 1918-1923) and ds4_server.c (lines 12946-12951).
  • Throttling applies to both prefill layers (graph_power_note_prefill_layer) and decode tokens (graph_power_note_decode_token) via exponential moving averages.
  • The algorithm inserts proportional sleeps to achieve the ratio work/(work+sleep) = power/100.
  • Valid values range from 1-99; 100 (default) disables throttling entirely.
  • Lower values significantly reduce thermal output for edge devices, while higher values maximize throughput for latency-sensitive applications.

Frequently Asked Questions

What is the default power setting in DS4?

By default, DS4 runs with --power 100 (or 0, which is treated as disabled), meaning no artificial throttling is applied. The GPU runs at maximum duty-cycle, providing the lowest latency but highest power consumption. You can verify this default behavior in ds4_help.c (lines 1669-1670).

How does the --power flag affect inference latency?

Higher sleep proportions directly increase latency. At --power 50, the runtime inserts sleep equal to the work duration, effectively doubling the time per prefill layer and decode token. The exact impact depends on kernel execution time, which the exponential moving average tracks continuously in graph_power_sleep.

Can I use --power with multiple GPU contexts?

The --power setting applies per ds4_gpu_graph instance. If you run multiple DS4 processes or threads, each respects its own engine.power_percent value. There is no global GPU lock; throttling is cooperative and process-specific.

Does power throttling impact model accuracy?

No. The --power flag only controls the duty-cycle timing through sleep_sec() calls in ds4.c. It does not modify quantization, attention mechanisms, or numerical precision. The model produces identical outputs; only the speed of computation changes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →