How to Balance KL Divergence and Refusal Reduction in Heretic Optimization
Heretic balances model fidelity and safety by treating KL divergence and refusal reduction as a multi-objective optimization problem, using configurable targets and conditional scoring to find Pareto-optimal solutions.
The open-source project p-e-w/heretic implements a sophisticated decensoring pipeline that must preserve the original model's capabilities while eliminating harmful refusals. Understanding how to balance KL divergence and refusal reduction in Heretic optimization requires navigating three core modules that control the trade-off between distribution fidelity and safety alignment.
Understanding the Multi-Objective Trade-Off
Heretic frames the decensoring challenge as a simultaneous minimization of two competing metrics:
- KL Divergence: Measures how much the decensored model's output distribution diverges from the original on "good" (safe) prompts. Lower values indicate better preservation of original capabilities.
- Refusal Rate: The proportion of harmful prompts that the model still refuses after decensoring. Lower values indicate more successful jailbreak or uncensoring.
According to the Heretic source code, these objectives are optimized together using Optuna's multi-objective support, returning a tuple score = (kld_score, refusals_score) with both directions set to StudyDirection.MINIMIZE.
Configuring the KL-Refusal Balance
The primary control surface for balancing these objectives resides in src/heretic/config.py (lines 63-73), which exposes two critical hyperparameters:
Setting the KL Divergence Target
The kl_divergence_target parameter (default 0.01) defines the acceptable threshold for distribution shift. When the current trial's KL falls below this target, the optimizer shifts its focus entirely to refusal reduction.
- Higher values (e.g.,
0.05): Allow more distribution drift, prioritizing aggressive refusal reduction - Lower values (e.g.,
0.005): Force strict fidelity preservation, potentially leaving more refusals intact
Adjusting the KL Divergence Scale
The kl_divergence_scale parameter (default 1.0) rescales the KL term to make it commensurable with the refusal score:
# Increase scale to penalize KL divergence more heavily
heretic --kl-divergence-target 0.005 --kl-divergence-scale 2.0 my-model
# Decrease scale to prioritize refusal reduction over fidelity
heretic --kl-divergence-target 0.02 --kl-divergence-scale 0.5 my-model
You can also persist these settings in config.toml:
kl_divergence_target = 0.02
kl_divergence_scale = 0.8
How Heretic Scores Trials
The scoring logic in src/heretic/evaluator.py (lines 95-119) implements a conditional formulation within the get_score method that dynamically shifts optimization priorities:
- Above target: If
kl_divergence > kl_divergence_target, the score iskl_divergence / kl_divergence_scale - Below target: If KL is already acceptable, the score becomes
refusals_score * kl_divergence_target / kl_divergence_scale
This conditional approach prevents "over-optimizing" KL divergence once it reaches acceptable levels, allowing the optimizer to focus computational budget on cutting refusals instead.
The refusals_score itself is computed as the ratio of current refusals to the baseline count (evaluator.base_refusals), meaning models with high initial refusal rates require more absolute reduction to show improvement.
The Pareto Front Selection Process
In src/heretic/main.py (lines 55-62), each Optuna trial returns the tuple (kld_score, refusals_score), creating a two-dimensional Pareto front of non-dominated solutions. After optimization completes (lines approximately 620-640), Heretic extracts the Pareto-optimal set and presents trials sorted by refusals first, then KL divergence.
This presentation allows users to select the specific trade-off point that matches their risk tolerance and capability requirements from the empirical frontier rather than relying on a single scalarized objective.
Practical Tuning Workflows
Command-Line Optimization
Start with defaults to establish a baseline, then adjust based on the Pareto front results:
# Baseline run
heretic my-model
# If KL is too high but refusals are good, tighten the target
heretic --kl-divergence-target 0.005 my-model
# If you need fewer refusals and can accept more drift
heretic --kl-divergence-target 0.05 --kl-divergence-scale 0.5 my-model
Programmatic Configuration
For custom integrations, modify the Settings object before invoking the run loop:
from heretic.config import Settings
from heretic.main import run
settings = Settings()
settings.kl_divergence_target = 0.02
settings.kl_divergence_scale = 0.8
run(settings=settings)
Analyzing the Pareto Front
Extract and inspect the trade-off curve after optimization completes:
import optuna
from pathlib import Path
checkpoint = Path("checkpoints") / "my-model-heretic.jsonl"
storage = optuna.storages.JournalStorage(
optuna.storages.journal.JournalFileBackend(str(checkpoint))
)
study = optuna.load_study(study_name="heretic", storage=storage)
pareto_front = [
(t.user_attrs["refusals"], t.user_attrs["kl_divergence"])
for t in study.trials
if t.state == optuna.trial.TrialState.COMPLETE
]
pareto_front.sort()
print("Top trade-offs (refusals, KL):", pareto_front[:5])
Summary
- Heretic treats KL divergence and refusal reduction as a multi-objective problem, returning a tuple score to Optuna rather than scalarizing the objectives prematurely.
kl_divergence_targetcontrols the acceptable distribution shift threshold; once KL drops below this value, optimization focuses exclusively on refusals.kl_divergence_scaleadjusts the relative weight of KL compared to refusals, effectively tuning how "expensive" distribution drift appears to the optimizer.- The conditional scoring logic in
src/heretic/evaluator.pydynamically switches between fidelity-focused and refusal-focused optimization based on current performance against the target. - Users select their preferred balance from the empirical Pareto front extracted in
src/heretic/main.py, with results sorted by refusal rate then KL divergence.
Frequently Asked Questions
What happens if I set the KL divergence target too low?
Setting kl_divergence_target to an extremely low value (e.g., 0.001) forces the optimizer to prioritize fidelity preservation above all else. While this maintains model capabilities, it may prevent the optimizer from exploring parameter regions that achieve meaningful refusal reduction, potentially leaving the model heavily censored. The conditional scoring in src/heretic/evaluator.py will continue penalizing KL heavily until the (possibly unattainable) target is met, starving the refusal objective of optimization pressure.
How does the KL divergence scale differ from the target?
The target defines a threshold for "good enough" fidelity, while the scale determines the penalty weight applied to KL violations. A higher scale (e.g., 2.0) makes any KL divergence appear more costly to the optimizer, encouraging conservative exploration even before reaching the target. Conversely, a lower scale (e.g., 0.5) reduces KL's influence in the objective function, effectively prioritizing refusal reduction during the entire optimization trajectory, not just after hitting the target.
Can I optimize for only one objective instead of balancing both?
While Heretic is designed for multi-objective optimization, you can approximate single-objective behavior by adjusting the configuration extremes. To optimize only for KL divergence, set kl_divergence_target very high (effectively disabling the conditional switch) and use a large kl_divergence_scale. To optimize only for refusals, set kl_divergence_scale near zero. However, the src/heretic/main.py implementation always returns a two-element tuple to Optuna, so true single-objective optimization would require modifying the study definition in the source code.
Why does Heretic use a ratio for the refusal score rather than absolute counts?
The refusal score is calculated as the ratio of current refusals to evaluator.base_refusals (the baseline count) because different models start with vastly different refusal rates. A model that refuses 90% of harmful prompts requires different optimization dynamics than one refusing only 10%. Normalizing to the baseline ensures that the optimization pressure scales appropriately to the starting point, preventing the optimizer from converging prematurely on models that simply started with fewer refusals.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →