How to Use the `--power` Flag for GPU Power Throttling in DwarfStar (ds4) to Reduce Heat and Fan Noise

The --power flag lets you cap GPU utilization to a percentage between 1 and 100 by inserting calculated sleeps between work units, directly reducing temperature and fan noise without altering model outputs.

The DwarfStar inference engine (ds4) is an open-source project by antirez designed for efficient local LLM inference. When running large models on consumer GPUs, sustained 100% utilization generates excessive heat and triggers aggressive fan curves. The --power runtime option addresses this by throttling the GPU duty cycle, allowing you to trade peak throughput for cooler, quieter operation.

What Is the --power Flag?

The --power flag is a runtime configuration option accepted by multiple binaries in the ds4 suite, including ds4, ds4-agent, ds4-server, ds4-bench, and ds4-eval. It accepts an integer value N where 1 ≤ N ≤ 100, representing the target GPU duty cycle percentage. When set to values below 100, the engine deliberately idles the GPU for a calculated fraction of each work interval, reducing sustained power draw and thermal output.

According to the help text generated in ds4_help.c (lines 169-174), this flag allows users to "limit GPU power usage to N percent" by slowing down inference to reduce heat and fan noise. The parsing logic is implemented consistently across entry points, as seen in ds4.c (lines 13228-13231), ds4_agent.c (lines 724-727), and ds4_server.c (lines 13228-13231).

How Power Throttling Works Internally

The throttling mechanism operates by measuring execution time and injecting precise delays between computational units. This ensures the GPU remains idle for the requested percentage of each cycle while preserving the logical sequence of model outputs.

Parsing and Configuration

When you invoke any ds4 binary with --power N, the command-line parser records the desired target in the power_target variable. This value persists throughout both the pre-fill phase (layer-by-layer processing) and the generation phase (token-by-token decoding).

Work Interval Measurement

During inference, the engine continuously measures "work intervals"—the time required to execute a single neural network layer or decode one token. These measurements happen in the core runtime within ds4.c, where the system tracks elapsed time using high-resolution timers specific to the Metal, CUDA, or ROCm backends.

Sleep Injection Logic

If power_target is less than 100, the engine calculates a sleep duration such that:


work_time / (work_time + sleep_time) ≈ power_target / 100

As implemented in ds4.c (lines 15652-15655), the system calls usleep (or the platform-specific sleep primitive) for the computed duration before proceeding to the next work unit. Critically, these sleeps occur only between work units, never during active computation, ensuring deterministic output.

Supported Models and Backends

The --power flag currently supports DeepSeek V4 Flash/PRO models across all major GPU backends:

  • Metal (Apple Silicon)
  • CUDA (NVIDIA GPUs)
  • ROCm (AMD GPUs)

Important limitation: GLM 5.2 models only accept --power 100 (full utilization). As noted in ds4.c (lines 48428-48430), the Metal implementation for GLM 5.2 does not yet expose the throttling helpers required for fractional power targets.

Practical Usage Examples

You can apply power throttling across different operational modes of the ds4 ecosystem.

Local Inference with Reduced Power

Run a single prompt at 50% GPU duty cycle to keep temperatures low during extended sessions:

./ds4 --power 50 -p "Summarize the README in a single sentence."

Server Mode with Thermal Constraints

Start the OpenAI-compatible HTTP server with a 40% power budget to maintain quiet operation during API serving:

./ds4-server --power 40 --ctx 100000

Agent Mode for Coding Tasks

When using the coding agent for tool-calling workflows, limit power to prevent thermal throttling during long-running sessions:

./ds4-agent --power 70

Benchmarking with Controlled Throughput

Measure generation speed while explicitly throttling to compare performance-per-watt:

./ds4-bench --power 60 --prompt-file speed-bench/promessi_sposi.txt \
    --ctx-start 32768 --gen-tokens 128

Combining with Other Optimizations

Stack power limiting with SSD streaming and large context windows for memory-constrained but thermally stable deployments:

./ds4 --power 55 --ssd-streaming --ctx 65536

Implementation Details

The power throttling system spans several source files in the antirez/ds4 repository:

  • ds4.c: Contains the core runtime logic (lines 15652-15655) where sleep calculations and GPU backend integration occur. This file also enforces the GLM 5.2 limitation (lines 48428-48430).
  • ds4_help.c: Generates user-facing documentation (lines 169-174) describing the flag's syntax and behavior.
  • ds4_agent.c: Entry point for the agent binary; parses --power at lines 724-727.
  • ds4_server.c: HTTP server implementation; handles flag parsing at lines 13228-13231.
  • ds4_cli.c: Primary CLI entry point for the main ds4 binary.
  • ds4_bench.c: Benchmark utility that respects power limits during throughput testing.

These components work together to parse the percentage target, measure execution intervals, and inject platform-appropriate delays to achieve the requested duty cycle.

Summary

  • The --power flag accepts values from 1 to 100, representing the target GPU duty cycle percentage.
  • Throttling works by measuring work intervals and inserting calculated sleeps between layers and tokens, as implemented in ds4.c.
  • Supported for DeepSeek V4 Flash/PRO on Metal, CUDA, and ROCm; GLM 5.2 requires --power 100.
  • Available across all major binaries: ds4, ds4-server, ds4-agent, and ds4-bench.
  • Sleep injection occurs only between work units, ensuring output quality remains identical to full-power runs.

Frequently Asked Questions

What values can I pass to the --power flag?

You must specify an integer between 1 and 100 inclusive. Values outside this range are rejected by the CLI parser in ds4_cli.c and related entry points. Passing --power 100 effectively disables throttling and allows maximum GPU utilization.

Does power throttling affect the quality or correctness of model outputs?

No. The sleeps are inserted strictly between computational work units (layers or tokens), never during the actual matrix operations or sampling. As implemented in the core runtime (ds4.c, lines 15652-15655), the logical sequence of generated tokens remains unchanged; only the temporal spacing between computations increases.

Why doesn't GLM 5.2 support power throttling below 100%?

The GLM 5.2 Metal implementation lacks the throttling helper functions required to calculate and inject sleep intervals. According to the source in ds4.c (lines 48428-48430), the engine enforces --power 100 for this model architecture until the Metal backend is updated to expose the necessary timing controls.

Can I use --power alongside other performance flags like --ssd-streaming?

Yes. The power limit operates orthogonally to memory management flags. You can combine --power with --ssd-streaming, context size adjustments (--ctx), and other optimizations. The throttling logic applies to GPU compute cycles regardless of whether model weights reside in system RAM, SSD, or GPU VRAM.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →