How to Balance KL Divergence and Refusal Reduction in Heretic Optimization

Heretic balances model fidelity and safety by treating KL divergence and refusal reduction as a multi-objective optimization problem, using configurable targets and conditional scoring to find Pareto-optimal solutions.

The open-source project p-e-w/heretic implements a sophisticated decensoring pipeline that must preserve the original model's capabilities while eliminating harmful refusals. Understanding how to balance KL divergence and refusal reduction in Heretic optimization requires navigating three core modules that control the trade-off between distribution fidelity and safety alignment.

Understanding the Multi-Objective Trade-Off

Heretic frames the decensoring challenge as a simultaneous minimization of two competing metrics:

  • KL Divergence: Measures how much the decensored model's output distribution diverges from the original on "good" (safe) prompts. Lower values indicate better preservation of original capabilities.
  • Refusal Rate: The proportion of harmful prompts that the model still refuses after decensoring. Lower values indicate more successful jailbreak or uncensoring.

According to the Heretic source code, these objectives are optimized together using Optuna's multi-objective support, returning a tuple score = (kld_score, refusals_score) with both directions set to StudyDirection.MINIMIZE.

Configuring the KL-Refusal Balance

The primary control surface for balancing these objectives resides in src/heretic/config.py (lines 63-73), which exposes two critical hyperparameters:

Setting the KL Divergence Target

The kl_divergence_target parameter (default 0.01) defines the acceptable threshold for distribution shift. When the current trial's KL falls below this target, the optimizer shifts its focus entirely to refusal reduction.

  • Higher values (e.g., 0.05): Allow more distribution drift, prioritizing aggressive refusal reduction
  • Lower values (e.g., 0.005): Force strict fidelity preservation, potentially leaving more refusals intact

Adjusting the KL Divergence Scale

The kl_divergence_scale parameter (default 1.0) rescales the KL term to make it commensurable with the refusal score:


# Increase scale to penalize KL divergence more heavily

heretic --kl-divergence-target 0.005 --kl-divergence-scale 2.0 my-model

# Decrease scale to prioritize refusal reduction over fidelity

heretic --kl-divergence-target 0.02 --kl-divergence-scale 0.5 my-model

You can also persist these settings in config.toml:

kl_divergence_target = 0.02
kl_divergence_scale = 0.8

How Heretic Scores Trials

The scoring logic in src/heretic/evaluator.py (lines 95-119) implements a conditional formulation within the get_score method that dynamically shifts optimization priorities:

  1. Above target: If kl_divergence > kl_divergence_target, the score is kl_divergence / kl_divergence_scale
  2. Below target: If KL is already acceptable, the score becomes refusals_score * kl_divergence_target / kl_divergence_scale

This conditional approach prevents "over-optimizing" KL divergence once it reaches acceptable levels, allowing the optimizer to focus computational budget on cutting refusals instead.

The refusals_score itself is computed as the ratio of current refusals to the baseline count (evaluator.base_refusals), meaning models with high initial refusal rates require more absolute reduction to show improvement.

The Pareto Front Selection Process

In src/heretic/main.py (lines 55-62), each Optuna trial returns the tuple (kld_score, refusals_score), creating a two-dimensional Pareto front of non-dominated solutions. After optimization completes (lines approximately 620-640), Heretic extracts the Pareto-optimal set and presents trials sorted by refusals first, then KL divergence.

This presentation allows users to select the specific trade-off point that matches their risk tolerance and capability requirements from the empirical frontier rather than relying on a single scalarized objective.

Practical Tuning Workflows

Command-Line Optimization

Start with defaults to establish a baseline, then adjust based on the Pareto front results:


# Baseline run

heretic my-model

# If KL is too high but refusals are good, tighten the target

heretic --kl-divergence-target 0.005 my-model

# If you need fewer refusals and can accept more drift

heretic --kl-divergence-target 0.05 --kl-divergence-scale 0.5 my-model

Programmatic Configuration

For custom integrations, modify the Settings object before invoking the run loop:

from heretic.config import Settings
from heretic.main import run

settings = Settings()
settings.kl_divergence_target = 0.02
settings.kl_divergence_scale = 0.8

run(settings=settings)

Analyzing the Pareto Front

Extract and inspect the trade-off curve after optimization completes:

import optuna
from pathlib import Path

checkpoint = Path("checkpoints") / "my-model-heretic.jsonl"
storage = optuna.storages.JournalStorage(
    optuna.storages.journal.JournalFileBackend(str(checkpoint))
)
study = optuna.load_study(study_name="heretic", storage=storage)

pareto_front = [
    (t.user_attrs["refusals"], t.user_attrs["kl_divergence"])
    for t in study.trials
    if t.state == optuna.trial.TrialState.COMPLETE
]
pareto_front.sort()
print("Top trade-offs (refusals, KL):", pareto_front[:5])

Summary

  • Heretic treats KL divergence and refusal reduction as a multi-objective problem, returning a tuple score to Optuna rather than scalarizing the objectives prematurely.
  • kl_divergence_target controls the acceptable distribution shift threshold; once KL drops below this value, optimization focuses exclusively on refusals.
  • kl_divergence_scale adjusts the relative weight of KL compared to refusals, effectively tuning how "expensive" distribution drift appears to the optimizer.
  • The conditional scoring logic in src/heretic/evaluator.py dynamically switches between fidelity-focused and refusal-focused optimization based on current performance against the target.
  • Users select their preferred balance from the empirical Pareto front extracted in src/heretic/main.py, with results sorted by refusal rate then KL divergence.

Frequently Asked Questions

What happens if I set the KL divergence target too low?

Setting kl_divergence_target to an extremely low value (e.g., 0.001) forces the optimizer to prioritize fidelity preservation above all else. While this maintains model capabilities, it may prevent the optimizer from exploring parameter regions that achieve meaningful refusal reduction, potentially leaving the model heavily censored. The conditional scoring in src/heretic/evaluator.py will continue penalizing KL heavily until the (possibly unattainable) target is met, starving the refusal objective of optimization pressure.

How does the KL divergence scale differ from the target?

The target defines a threshold for "good enough" fidelity, while the scale determines the penalty weight applied to KL violations. A higher scale (e.g., 2.0) makes any KL divergence appear more costly to the optimizer, encouraging conservative exploration even before reaching the target. Conversely, a lower scale (e.g., 0.5) reduces KL's influence in the objective function, effectively prioritizing refusal reduction during the entire optimization trajectory, not just after hitting the target.

Can I optimize for only one objective instead of balancing both?

While Heretic is designed for multi-objective optimization, you can approximate single-objective behavior by adjusting the configuration extremes. To optimize only for KL divergence, set kl_divergence_target very high (effectively disabling the conditional switch) and use a large kl_divergence_scale. To optimize only for refusals, set kl_divergence_scale near zero. However, the src/heretic/main.py implementation always returns a two-element tuple to Optuna, so true single-objective optimization would require modifying the study definition in the source code.

Why does Heretic use a ratio for the refusal score rather than absolute counts?

The refusal score is calculated as the ratio of current refusals to evaluator.base_refusals (the baseline count) because different models start with vastly different refusal rates. A model that refuses 90% of harmful prompts requires different optimization dynamics than one refusing only 10%. Normalizing to the baseline ensures that the optimization pressure scales appropriately to the starting point, preventing the optimizer from converging prematurely on models that simply started with fewer refusals.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →