Directional Steering in ds4: How to Influence Model Behavior with Hidden-State Biasing

Directional steering is a lightweight inference-time mechanism in the ds4 engine that biases the hidden-state vectors of a DeepSeek-V4 model toward a user-supplied direction vector, allowing you to guide model outputs without retraining.

Directional steering lets you nudge a model's internal representations during inference by projecting layer outputs onto a precomputed direction vector. This technique, implemented in the antirez/ds4 repository, provides fine-grained control over generation style, tone, or reasoning patterns through simple command-line flags or programmatic configuration.

How Directional Steering Works

The ds4 engine implements directional steering as a projection operation applied to either the feed-forward network (FFN) output, the attention block output, or both. When enabled, the engine computes the dot product between the layer output and the steering direction, then adds this projection back to the hidden state scaled by user-defined factors.

The core logic resides in ds4.c, where two primary functions handle the operation:

  • cpu_directional_steering_enabled (line 1767) — validates that steering is properly configured with a non-NULL direction pointer and non-zero scale
  • cpu_directional_steering_project_rows — performs the actual projection of row vectors onto the steering direction

The steering direction itself is a per-layer float32 vector stored in a binary file. The engine expects one floating-point value per transformer layer, which it loads into memory and normalizes before application.

Applying Directional Steering via Command Line

The ds4 CLI exposes three flags for configuring steering behavior. According to ds4_cli.c (line 1924) and documented in ds4_help.c (line 219), you can load a direction file and independently control steering intensity for FFN and attention stages.

Default behavior when only --dir-steering-file is provided:

  • FFN steering scale: 1.0 (full projection applied)
  • Attention steering scale: 0.0 (disabled)

Override these defaults with explicit scale arguments.

Basic FFN Steering

./ds4 -p "Write tersely" \
     --dir-steering-file dir.bin \
     --dir-steering-ffn 0.8

This loads dir.bin and applies 80% of the full projection strength to FFN outputs only.

Combined FFN and Attention Steering

./ds4 -p "Explain quantum entanglement" \
     --dir-steering-file dir.bin \
     --dir-steering-ffn 0.6 \
     --dir-steering-attn 0.3

Here, both transformation stages receive the directional bias, with different intensities. Lower attention scale values are often preferable to preserve positional and contextual information while still biasing semantic content.

Applying Directional Steering Programmatically

For embedded applications or custom inference pipelines, configure steering directly through the engine_t struct defined in ds4.c. The relevant fields (referenced around line 12060) include:

  • steering_dirs — pointer to loaded per-layer direction vector
  • steering_scale — global scaling factor applied to all steering
  • directional_steering_ffn — FFN-specific scale
  • directional_steering_attn — attention-specific scale

C Integration Example

engine_t eng = ds4_engine_create(...);

/* Load binary: one float per transformer layer */
float *dirs = load_dir_file("dir.bin", n_layers);

/* Configure steering parameters */
eng.steering_dirs              = dirs;
eng.steering_scale             = 1.0f;
eng.directional_steering_ffn   = 0.8f;
eng.directional_steering_attn  = 0.3f;

/* Inference automatically applies steering in ds4_forward */
ds4_forward(&eng, ...);

The load_dir_file utility shown here is hypothetical—implement binary file reading to populate a float * sized to your model's layer count. The steering logic triggers automatically within ds4_forward when cpu_directional_steering_enabled returns true.

Steering Integration Points in the Forward Pass

The ds4 engine applies directional steering at two specific locations in the transformer computation graph:

Location Function Call Source Line Purpose
Post-FFN cpu_directional_steering_project_rows 11786 Biases feed-forward transformation outputs
Post-attention cpu_directional_steering_project_rows 13158 Biases self-attention outputs

Both calls use the same projection routine but receive independent scale factors, allowing asymmetric application of the directional bias across the model's two primary transformation pathways.

Creating Direction Vectors

The direction file format is straightforward: raw float32 values, one per layer, little-endian. Extract directions through:

  • Activation analysis: Compute principal components from hidden states across response samples with desired characteristics
  • Contrastive methods: Subtract mean activations of unwanted outputs from desired outputs
  • Fine-tuning residuals: Use weight deltas from adapter training converted to activation space

Store results in a binary file and validate dimensions match your model configuration exactly.

Server Mode Steering Configuration

The ds4 HTTP server (ds4_server.c) mirrors CLI flag parsing, enabling steering via request parameters. This allows dynamic direction loading and scale adjustment per-request in production deployments without process restarts.

Summary

  • Directional steering biases hidden states toward a user-supplied direction without modifying model weights
  • Binary direction files contain one float32 value per layer, loaded via --dir-steering-file
  • Independent FFN and attention scales control projection intensity at --dir-steering-ffn and --dir-steering-attn
  • Defaults favor FFN steering (scale 1.0) with attention disabled (scale 0.0)
  • Implementation spans ds4.c for core logic, ds4_cli.c for command-line parsing, and ds4_help.c for documentation

Frequently Asked Questions

How do I create a direction file for steering?

Generate a binary file containing one 32-bit floating-point value per transformer layer in your model. Common approaches include principal component analysis on desired response activations or contrastive methods comparing activations from different output styles. Ensure values are in little-endian format and match the exact layer count of your DeepSeek-V4 checkpoint.

Can I apply different steering strengths to different layers?

The current ds4 implementation uses a single float32 per layer, loaded from the binary file. Each layer's direction value acts as that layer's steering coefficient. To vary strength across layers, encode those preferences directly into the direction file values rather than using the uniform scale factors.

What is the difference between FFN and attention steering?

FFN steering projects the feed-forward network output, primarily affecting the model's knowledge retrieval and semantic transformations. Attention steering projects the self-attention output, influencing how the model weights token relationships and positional information. FFN steering generally produces more predictable stylistic changes, while attention steering can alter coherence patterns and should be applied with lower scale values.

Does directional steering work with quantized models?

Yes. The steering projection in cpu_directional_steering_project_rows operates on dequantized float32 activations during the forward pass, making it compatible with all quantization schemes supported by ds4. The direction vector remains float32 regardless of weight precision.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →