# How Directional Steering with Power Values Below 100 Is Implemented in ds4

> Discover how ds4 implements directional steering below 100. Learn about the dot-product calculation and power factor for subtle, bounded steering effects in this technical deep-dive.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: internals
- Published: 2026-08-09

---

**Directional steering in ds4 applies a lightweight bias to activation tensors by computing the dot-product between hidden states and a steering vector, scaling the projection by a power factor (≤100), and subtracting the result, where values below 100 produce subtle, bounded steering effects.**

The antirez/ds4 repository implements directional steering as a lightweight inference-time technique to guide model behavior without retraining. When configuring **power values below 100**, users achieve gentle nudges toward desired semantic directions rather than aggressive overrides. This implementation modifies activation tensors in both attention and feed-forward layers through a consistent row-wise projection algorithm available across CPU, CUDA, ROCm, and Metal backends.

## Core Implementation in ds4.c

The steering logic resides primarily in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), where the engine checks whether steering is enabled and then applies the mathematical transformation to each token's hidden-state vector.

### The Enable Check Function

Before processing, the engine validates steering parameters through `cpu_directional_steering_enabled()` at lines 1767–1776. This function returns `false` if the steering vector is `NULL` or if the scale (referred to as *power* in the CLI) equals `0.0f`, effectively bypassing the steering code when disabled.

```c
static bool cpu_directional_steering_enabled(const float *dirs, float scale) {
    return dirs != NULL && scale != 0.0f;
}

```

This early exit prevents unnecessary computation when steering is inactive or when a zero power value effectively disables the feature.

### Row-Wise Projection Algorithm

The core mathematics execute in `cpu_directional_steering_project_rows()` at lines 37031–37052. For each row (representing a token's hidden-state vector), the function computes the dot-product with the steering direction, scales it by the user-provided power, and subtracts the result from the original activation.

```c
static void cpu_directional_steering_project_rows(
        float       *x,            // activation matrix (rows × dim)
        const float *dirs,         // single 4096-wide direction vector
        uint32_t     il,           // layer index (unused here, kept for API symmetry)
        uint32_t     rows,         // number of rows to process
        float        scale) {      // user-provided power (≤ 100)

    if (!cpu_directional_steering_enabled(dirs, scale) || !x || rows == 0) return;

    const uint32_t dim = DS4_N_EMBD;          // 4096 for DeepSeek-V4-Flash
    for (uint32_t r = 0; r < rows; ++r) {
        const float *d = dirs;                // the direction is the same for every row
        float *row = x + (uint64_t)r * dim;   // pointer to the start of the row

        /* Compute dot = Σ_j row[j] * d[j] */
        float dot = 0.0f;
        for (uint32_t j = 0; j < dim; ++j)
            dot += row[j] * d[j];

        /* Apply steering: row[j] -= dot * d[j] * scale */
        const float coeff = dot * scale;
        for (uint32_t j = 0; j < dim; ++j)
            row[j] -= coeff * d[j];
    }
}

```

The algorithm assumes the direction vector is normalized when loaded from the steering file. Consequently, the magnitude of the update is directly proportional to the supplied `scale` parameter. When **power values are below 100**, the subtraction remains modest, creating a gentle bias toward the supplied direction without destabilizing the model's activations.

## GPU Acceleration Across Backends

The ds4 repository provides thin wrappers that port the CPU logic to GPU architectures, ensuring consistent steering behavior regardless of hardware.

### CUDA Implementation

In [`ds4_gpu.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu.c), the `ds4_gpu_directional_steering_project_tensor()` function launches the CUDA kernel with the same mathematical operations:

```c
extern "C" int ds4_gpu_directional_steering_project_tensor(
        ds4_gpu_tensor *x,               // device tensor (rows × dim)
        const ds4_gpu_tensor *dirs,     // direction vector stored as GPU tensor
        uint32_t il,
        uint32_t rows,
        float scale) {

    const uint32_t dim = DS4_N_EMBD;
    const uint32_t nth = 256;                 // threads per block
    directional_steering_project_kernel<<<rows, nth>>>(
        x->data, dirs->data, il, rows, dim, scale);
    return cuda_ok(cudaGetLastError(), "directional steering launch");
}

```

The kernel `directional_steering_project_kernel` mirrors the CPU implementation, performing per-row dot-products and scaled subtractions in parallel across the GPU.

### ROCm and Metal Support

For AMD hardware, `rocm/ds4_rocm_misc_launch.cuh` contains the launch wrapper at lines 58–81, invoking `directional_steering_project_kernel` with ROCm semantics. On Apple Silicon, `metal/dsv4_misc.metal` implements `kernel_dsv4_directional_steering_project_f32` at lines 368–371, ensuring identical **directional steering** behavior across all supported accelerators.

## Working with Sub-100 Power Values

Configuring **power values below 100** requires understanding how the scaling factor interacts with normalized direction vectors.

### Scaling Effects on Activation Tensors

Because the direction vector is normalized to unit length during file loading, the `scale` parameter (power) directly controls the update magnitude. A value of `30.0f` applies 30% of the theoretical maximum steering intensity, while `0.0f` disables steering entirely. This linear relationship ensures predictable, bounded effects when applying low-power directional steering to sensitive layers.

### Practical Implementation Examples

To apply steering during model inference using the C API:

```c
// Load steering directions (one 4096-wide vector per layer)
g->directional_steering_dirs = ds4_gpu_tensor_alloc(n * sizeof(dirs[0]));
ds4_gpu_tensor_write(g->directional_steering_dirs, 0, dirs, n * sizeof(dirs[0]));

// Configure modest steering power (< 100) for subtle guidance
g->directional_steering_attn_scale = 42.0f;   // power for attention layers
g->directional_steering_ffn_scale  = 15.0f;   // power for feed-forward layers

// During inference, the engine automatically applies steering:
if (metal_graph_directional_steering_attn_enabled(g)) {
    metal_graph_apply_directional_steering(g, attn_out, il, 1,
        g->directional_steering_attn_scale);
}

```

For CPU-only testing without GPU dependencies:

```c
float activations[2][4096];   // two tokens
float direction[4096];        // normalized steering vector
float power = 30.0f;          // < 100 for gentle steering

cpu_directional_steering_project_rows(&activations[0][0],
                                      direction,
                                      /*layer*/0,
                                      /*rows*/2,
                                      power);
// activations now contain values shifted 30% toward the direction vector

```

## Summary

- **Directional steering** in ds4 modifies activations by subtracting the scaled projection of hidden states onto a steering vector.
- The implementation spans [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (lines 1767–1776 and 37031–37052), [`ds4_gpu.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu.c), `rocm/ds4_rocm_misc_launch.cuh`, and `metal/dsv4_misc.metal`.
- **Power values below 100** produce gentle, bounded steering effects proportional to the scale factor.
- The algorithm is hardware-agnostic, with identical logic implemented for CPU, CUDA, ROCm, and Metal backends.
- Steering is bypassed entirely when power equals `0.0f` or when direction vectors are `NULL`.

## Frequently Asked Questions

### What happens when the power value is set to 0 in ds4 directional steering?

When the power value (scale) is `0.0f`, the `cpu_directional_steering_enabled()` function returns `false`, causing the engine to skip all steering computations entirely. The activation tensors pass through unchanged, and no GPU kernels are launched for steering operations.

### Why is the maximum power value capped at 100 in ds4?

The value 100 represents the default maximum scale factor in the ds4 CLI and API, corresponding to full-intensity steering where the projection is subtracted at full magnitude. Values below 100 provide fractional intensity, allowing fine-grained control over the strength of the directional bias applied to attention and feed-forward layers.

### How does ds4 ensure consistent steering across different hardware backends?

The repository implements the same mathematical algorithm—computing the dot-product, scaling by power, and subtracting from activations—in functionally identical CPU, CUDA, ROCm, and Metal kernels. The `ds4_gpu_directional_steering_project_tensor()` wrapper and its equivalents handle backend-specific launching while preserving the core row-wise projection logic found in `cpu_directional_steering_project_rows()`.

### Can directional steering be applied selectively to specific layers?

Yes. The implementation accepts a layer index parameter (`il`) in functions like `cpu_directional_steering_project_rows()`, and the global context structure maintains separate scale factors for attention (`directional_steering_attn_scale`) and feed-forward (`directional_steering_ffn_scale`) layers, enabling granular control over which components receive the steering modification.