How to Implement Model Steering with Soup Steer: CLI and Python API Guide
Soup implements model steering as a three-step pipeline—build an artifact, load it, and install a forward hook—that supports CAA, ITI, and RepE methods via both a command-line interface and direct Python API.
Model steering lets you inject custom behavior into language models at inference time without retraining. In the Soup open-source toolkit, this capability is centralized in the soup_cli.utils.steering module and exposed through the soup steer CLI command defined in src/soup_cli/commands/steer.py. This guide walks through the complete implementation flow with production-ready code examples.
Supported Steering Methods in Soup
Soup supports three steering backends, defined in the SUPPORTED_STEERING_METHODS tuple and validated by validate_steering_method in src/soup_cli/utils/steering.py (lines 25–34):
- CAA – Contrastive Activation Addition: Adds a contrastive vector to the residual stream
- ITI – Inference-Time Intervention: Shifts selected attention-head outputs
- RepE – Representation Engineering: Applies a PCA-derived direction in the residual stream
The framework enforces safety limits: steering strength is capped at ±10, and steering names must be 128 characters or fewer (validate_steering_strength and validate_steering_name at lines 17–46 and 49–71).
Step 1: Build a Steering Artifact
The build_steering_vector function constructs a reusable steering artifact from contrastive data. It writes two files to a directory: steering_config.json (metadata) and steering_vector.safetensors (the tensor data).
from soup_cli.utils.steering import build_steering_vector
artifact_dir = build_steering_vector(
method="caa", # "caa", "iti", or "repe"
name="refusal_vector", # safe identifier for caching
base=12, # target transformer layer
pairs_path="data/pairs.jsonl", # JSONL of (positive, negative) prompts
)
print(f"Artifact created at: {artifact_dir}")
Input validation occurs automatically. The function returns a path that defaults to ~/.cache/soup/steering/ for subsequent reuse.
Step 2: Load and Install the Steering Hook
At runtime, load the cached artifact and register a forward hook that applies the steering vector scaled by your chosen strength.
from soup_cli.utils.steering import load_steering_artifact, install_steering_hook
# Assumes `model` is a Hugging Face PreTrainedModel
loaded = load_steering_artifact("refusal_vector") # resolves from cache
handle = install_steering_hook(
model,
loaded,
strength=2.5 # multiplier applied to the vector (|strength| ≤ 10)
)
# Generate with steering active
output = model.generate(input_ids, max_new_tokens=100)
# Remove hook when finished
handle.remove()
The load_steering_artifact function (lines 130–160) resolves the artifact directory—falling back to local cache if needed—and returns a LoadedSteering object containing both the config and tensor. The install_steering_hook function (lines 170–210) registers the intervention on either the residual or attn_o_proj_input stream depending on method.
CLI Quick Start: One-Line Steering
The soup steer command automates the full pipeline. From src/soup_cli/commands/steer.py:
soup steer \
--method caa \
--name refusal_vector \
--base 12 \
--pairs data/pairs.jsonl \
--strength 2.5 \
--model meta-llama/Llama-2-7b-hf \
--prompt "Explain how to build a website"
The CLI handles artifact building (if missing), caching, hook installation, and generation in a single invocation. Re-runs skip the build step and reuse cached artifacts.
Understanding the Hook Implementation
The steering intervention is implemented as a PyTorch forward hook registered via install_steering_hook. Per the source in src/soup_cli/utils/steering.py (lines 170–210):
- The hook identifies the target layer from
loaded.config["base_layer"] - During forward passes, it intercepts activations at the specified location
- It adds
strength * steering_vectorto the activation tensor - Returns a
RemovableHandlefor cleanup
This design keeps steering stateless and removable—no permanent model modification occurs.
Key Files and Architecture
| File | Responsibility |
|---|---|
src/soup_cli/utils/steering.py |
Core API: validators, SteeringArtifact dataclass (lines 90–110), build_steering_vector, load_steering_artifact, install_steering_hook |
src/soup_cli/commands/steer.py |
CLI entry point for soup steer |
tests/test_v07110.py |
Integration tests covering validation, artifact lifecycle, and hook behavior |
The SteeringArtifact dataclass (frozen, lines 90–110) stores method, name, target layer, hidden dimension, and intervention direction—enabling reproducible, self-documenting steering configurations.
Summary
- Soup steer provides both CLI (
soup steer) and Python API (build_steering_vector,install_steering_hook) interfaces - Three methods are supported: CAA, ITI, and RepE, validated at input time
- Artifacts are cached automatically and consist of
steering_config.json+steering_vector.safetensors - Hooks are removable and apply strength-scaled vectors to residual or attention streams
- Safety guards limit strength to ±10 and names to 128 characters
Frequently Asked Questions
What input format does Soup expect for contrastive pairs?
Soup requires a JSONL file where each line contains a JSON object with positive and negative string fields. These pairs define the behavioral direction you want to steer toward or away from.
Can I apply multiple steering vectors simultaneously?
The current install_steering_hook API returns a single handle per invocation. You can install multiple hooks by calling it repeatedly with different artifacts and names, though interaction effects between vectors are not automatically managed—test combined strengths carefully.
How do I check what steering artifacts are cached?
Cached artifacts live in ~/.cache/soup/steering/ by default. Each artifact is a subdirectory named by the name parameter containing steering_config.json and steering_vector.safetensors. List directories or use load_steering_artifact with the name to inspect.
Why is my steering strength capped at 10?
The validate_steering_strength function (lines 49–71) enforces |strength| ≤ 10 as a safety measure. Extreme values can destabilize model outputs or produce degenerate text. If you need stronger effects, consider building a more concentrated contrastive vector rather than increasing strength.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →