How Directional Steering with Power Values Below 100 Is Implemented in ds4

Directional steering in ds4 applies a lightweight bias to activation tensors by computing the dot-product between hidden states and a steering vector, scaling the projection by a power factor (≤100), and subtracting the result, where values below 100 produce subtle, bounded steering effects.

The antirez/ds4 repository implements directional steering as a lightweight inference-time technique to guide model behavior without retraining. When configuring power values below 100, users achieve gentle nudges toward desired semantic directions rather than aggressive overrides. This implementation modifies activation tensors in both attention and feed-forward layers through a consistent row-wise projection algorithm available across CPU, CUDA, ROCm, and Metal backends.

Core Implementation in ds4.c

The steering logic resides primarily in ds4.c, where the engine checks whether steering is enabled and then applies the mathematical transformation to each token's hidden-state vector.

The Enable Check Function

Before processing, the engine validates steering parameters through cpu_directional_steering_enabled() at lines 1767–1776. This function returns false if the steering vector is NULL or if the scale (referred to as power in the CLI) equals 0.0f, effectively bypassing the steering code when disabled.

static bool cpu_directional_steering_enabled(const float *dirs, float scale) {
    return dirs != NULL && scale != 0.0f;
}

This early exit prevents unnecessary computation when steering is inactive or when a zero power value effectively disables the feature.

Row-Wise Projection Algorithm

The core mathematics execute in cpu_directional_steering_project_rows() at lines 37031–37052. For each row (representing a token's hidden-state vector), the function computes the dot-product with the steering direction, scales it by the user-provided power, and subtracts the result from the original activation.

static void cpu_directional_steering_project_rows(
        float       *x,            // activation matrix (rows × dim)
        const float *dirs,         // single 4096-wide direction vector
        uint32_t     il,           // layer index (unused here, kept for API symmetry)
        uint32_t     rows,         // number of rows to process
        float        scale) {      // user-provided power (≤ 100)

    if (!cpu_directional_steering_enabled(dirs, scale) || !x || rows == 0) return;

    const uint32_t dim = DS4_N_EMBD;          // 4096 for DeepSeek-V4-Flash
    for (uint32_t r = 0; r < rows; ++r) {
        const float *d = dirs;                // the direction is the same for every row
        float *row = x + (uint64_t)r * dim;   // pointer to the start of the row

        /* Compute dot = Σ_j row[j] * d[j] */
        float dot = 0.0f;
        for (uint32_t j = 0; j < dim; ++j)
            dot += row[j] * d[j];

        /* Apply steering: row[j] -= dot * d[j] * scale */
        const float coeff = dot * scale;
        for (uint32_t j = 0; j < dim; ++j)
            row[j] -= coeff * d[j];
    }
}

The algorithm assumes the direction vector is normalized when loaded from the steering file. Consequently, the magnitude of the update is directly proportional to the supplied scale parameter. When power values are below 100, the subtraction remains modest, creating a gentle bias toward the supplied direction without destabilizing the model's activations.

GPU Acceleration Across Backends

The ds4 repository provides thin wrappers that port the CPU logic to GPU architectures, ensuring consistent steering behavior regardless of hardware.

CUDA Implementation

In ds4_gpu.c, the ds4_gpu_directional_steering_project_tensor() function launches the CUDA kernel with the same mathematical operations:

extern "C" int ds4_gpu_directional_steering_project_tensor(
        ds4_gpu_tensor *x,               // device tensor (rows × dim)
        const ds4_gpu_tensor *dirs,     // direction vector stored as GPU tensor
        uint32_t il,
        uint32_t rows,
        float scale) {

    const uint32_t dim = DS4_N_EMBD;
    const uint32_t nth = 256;                 // threads per block
    directional_steering_project_kernel<<<rows, nth>>>(
        x->data, dirs->data, il, rows, dim, scale);
    return cuda_ok(cudaGetLastError(), "directional steering launch");
}

The kernel directional_steering_project_kernel mirrors the CPU implementation, performing per-row dot-products and scaled subtractions in parallel across the GPU.

ROCm and Metal Support

For AMD hardware, rocm/ds4_rocm_misc_launch.cuh contains the launch wrapper at lines 58–81, invoking directional_steering_project_kernel with ROCm semantics. On Apple Silicon, metal/dsv4_misc.metal implements kernel_dsv4_directional_steering_project_f32 at lines 368–371, ensuring identical directional steering behavior across all supported accelerators.

Working with Sub-100 Power Values

Configuring power values below 100 requires understanding how the scaling factor interacts with normalized direction vectors.

Scaling Effects on Activation Tensors

Because the direction vector is normalized to unit length during file loading, the scale parameter (power) directly controls the update magnitude. A value of 30.0f applies 30% of the theoretical maximum steering intensity, while 0.0f disables steering entirely. This linear relationship ensures predictable, bounded effects when applying low-power directional steering to sensitive layers.

Practical Implementation Examples

To apply steering during model inference using the C API:

// Load steering directions (one 4096-wide vector per layer)
g->directional_steering_dirs = ds4_gpu_tensor_alloc(n * sizeof(dirs[0]));
ds4_gpu_tensor_write(g->directional_steering_dirs, 0, dirs, n * sizeof(dirs[0]));

// Configure modest steering power (< 100) for subtle guidance
g->directional_steering_attn_scale = 42.0f;   // power for attention layers
g->directional_steering_ffn_scale  = 15.0f;   // power for feed-forward layers

// During inference, the engine automatically applies steering:
if (metal_graph_directional_steering_attn_enabled(g)) {
    metal_graph_apply_directional_steering(g, attn_out, il, 1,
        g->directional_steering_attn_scale);
}

For CPU-only testing without GPU dependencies:

float activations[2][4096];   // two tokens
float direction[4096];        // normalized steering vector
float power = 30.0f;          // < 100 for gentle steering

cpu_directional_steering_project_rows(&activations[0][0],
                                      direction,
                                      /*layer*/0,
                                      /*rows*/2,
                                      power);
// activations now contain values shifted 30% toward the direction vector

Summary

  • Directional steering in ds4 modifies activations by subtracting the scaled projection of hidden states onto a steering vector.
  • The implementation spans ds4.c (lines 1767–1776 and 37031–37052), ds4_gpu.c, rocm/ds4_rocm_misc_launch.cuh, and metal/dsv4_misc.metal.
  • Power values below 100 produce gentle, bounded steering effects proportional to the scale factor.
  • The algorithm is hardware-agnostic, with identical logic implemented for CPU, CUDA, ROCm, and Metal backends.
  • Steering is bypassed entirely when power equals 0.0f or when direction vectors are NULL.

Frequently Asked Questions

What happens when the power value is set to 0 in ds4 directional steering?

When the power value (scale) is 0.0f, the cpu_directional_steering_enabled() function returns false, causing the engine to skip all steering computations entirely. The activation tensors pass through unchanged, and no GPU kernels are launched for steering operations.

Why is the maximum power value capped at 100 in ds4?

The value 100 represents the default maximum scale factor in the ds4 CLI and API, corresponding to full-intensity steering where the projection is subtracted at full magnitude. Values below 100 provide fractional intensity, allowing fine-grained control over the strength of the directional bias applied to attention and feed-forward layers.

How does ds4 ensure consistent steering across different hardware backends?

The repository implements the same mathematical algorithm—computing the dot-product, scaling by power, and subtracting from activations—in functionally identical CPU, CUDA, ROCm, and Metal kernels. The ds4_gpu_directional_steering_project_tensor() wrapper and its equivalents handle backend-specific launching while preserving the core row-wise projection logic found in cpu_directional_steering_project_rows().

Can directional steering be applied selectively to specific layers?

Yes. The implementation accepts a layer index parameter (il) in functions like cpu_directional_steering_project_rows(), and the global context structure maintains separate scale factors for attention (directional_steering_attn_scale) and feed-forward (directional_steering_ffn_scale) layers, enabling granular control over which components receive the steering modification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →