Understanding Pareto Front Results in Heretic: How to Choose Optimal Decensoring Parameters

Heretic returns a Pareto front of Optuna trials that balance refusal reduction against KL divergence, allowing you to select parameters based on your specific trade-off preferences rather than a single "best" value.

When running decensoring experiments with Heretic, the tool performs multi-objective optimization using Optuna to find ideal parameters for the abliterate operation. Unlike single-objective optimization that returns one winner, Heretic generates Pareto front results—a set of non-dominated trials offering different balances between minimizing refusals and preserving model behavior.

What Are Pareto Front Results in Heretic?

Heretic optimizes two competing objectives simultaneously:

  • Refusals: The count of "bad-prompt" refusals produced by the decensored model (lower is better)
  • KL divergence: The statistical distance between the modified model's output distribution and the original (lower indicates better preservation of capabilities)

A trial dominates another when it achieves equal or better performance on both metrics and strictly better performance on at least one. The Pareto front consists of all trials that are not dominated by any other trial in the study. Each point on this front represents a unique trade-off: you cannot reduce refusals further without increasing KL divergence, or vice versa.

How Heretic Builds the Pareto Front

The Pareto front construction logic resides in src/heretic/main.py. After optimization completes, Heretic executes the following algorithm:

  1. Collect completed trials: Filters for trials with TrialState.COMPLETE status【src/heretic/main.py】
  2. Sort by objectives: Orders trials first by refusal count (ascending), then by KL divergence (ascending)【src/heretic/main.py】
  3. Walk the frontier: Iterates through sorted trials while tracking the minimum KL divergence encountered. Whenever a trial exhibits lower KL than the current minimum, it is added to the Pareto front【src/heretic/main.py】
import math
from optuna.trial import TrialState

def build_pareto_front(study):
    completed_trials = [
        trial for trial in study.trials 
        if trial.state == TrialState.COMPLETE
    ]
    
    sorted_trials = sorted(
        completed_trials,
        key=lambda trial: (
            trial.user_attrs["refusals"], 
            trial.user_attrs["kl_divergence"]
        ),
    )

    min_divergence = math.inf
    best_trials = []
    
    for trial in sorted_trials:
        kl_divergence = trial.user_attrs["kl_divergence"]
        if kl_divergence < min_divergence:
            min_divergence = kl_divergence
            best_trials.append(trial)
            
    return best_trials

This algorithm efficiently identifies the convex hull of optimal trade-offs, ensuring you see only the trials that matter for decision-making.

Selecting the Best Parameters from the Pareto Front

Heretic provides an interactive selection interface via prompt_toolkit. After building the front, it displays each trial with its refusal count and KL divergence:

from prompt_toolkit.shortcuts import radiolist_dialog

choices = [
    {
        "title": (
            f"[Trial {trial.user_attrs['index']:>3}] "
            f"Refusals: {trial.user_attrs['refusals']:>2}, "
            f"KL divergence: {trial.user_attrs['kl_divergence']:.4f}"
        ),
        "value": trial,
    }
    for trial in best_trials
]

When you select a trial, Heretic calls utils.get_trial_parameters to display the specific hyperparameters【src/heretic/utils.py】:

from heretic.utils import get_trial_parameters

def display_selected_trial(trial):
    print("* Parameters:")
    for name, value in get_trial_parameters(trial).items():
        print(f"  * {name} = {value}")

Decision Guidelines

Use these criteria to select the optimal trial for your use case:

  • Prioritize low refusals: Choose trials on the left side of the front (minimal refusal count) when you need the model to answer harmful prompts reliably
  • Prioritize model fidelity: Choose trials with KL divergence < 0.1 to preserve original capabilities; avoid any trial with KL > 1.0 as this indicates severe distribution shift【src/heretic/main.py】
  • Balanced approach: Select a middle trial that offers acceptable refusal reduction (e.g., < 20% of bad prompts refused) with moderate KL divergence (0.3-0.5)

Programmatically Accessing Pareto Front Results

You can extract Pareto front data without using the interactive CLI by working directly with the Optuna study:

import optuna
from heretic.evaluator import Evaluator

def extract_pareto_data(study_path):
    study = optuna.load_study(
        study_name="heretic_optimization",
        storage=f"sqlite:///{study_path}"
    )
    
    # Reconstruct front using Heretic's logic

    completed = [t for t in study.trials if t.state == TrialState.COMPLETE]
    completed.sort(key=lambda t: (t.user_attrs["refusals"], 
                                  t.user_attrs["kl_divergence"]))
    
    front = []
    min_kl = float("inf")
    for trial in completed:
        kl = trial.user_attrs["kl_divergence"]
        if kl < min_kl:
            min_kl = kl
            front.append({
                "number": trial.number,
                "refusals": trial.user_attrs["refusals"],
                "kl_divergence": kl,
                "params": trial.params
            })
    return front

This allows automated analysis of trade-offs or integration with external model selection pipelines.

Summary

  • Pareto front results in Heretic represent the set of non-dominated trials balancing refusal reduction against KL divergence preservation
  • The front is constructed in src/heretic/main.py by sorting completed trials and selecting those with progressively lower KL values for each refusal count level
  • Selection criteria depend on your priorities: minimize refusals for maximum compliance, minimize KL divergence (< 0.1) for capability preservation, or choose a balanced middle point
  • Avoid trials with KL divergence > 1.0 as these indicate severe model degradation
  • Access trial parameters programmatically using utils.get_trial_parameters() or extract the front manually from the Optuna study database

Frequently Asked Questions

What does it mean if a trial is on the Pareto front?

A trial on the Pareto front means no other trial in the study achieved both fewer refusals and lower KL divergence simultaneously. It represents an optimal trade-off point where improving one metric would necessarily worsen the other. These are the only trials worth considering when selecting final parameters.

How do I interpret KL divergence values in Heretic?

KL divergence measures how much the decensored model's output distribution differs from the original. Values below 0.1 indicate excellent preservation of original capabilities, 0.3-0.5 represent moderate but acceptable drift, and values exceeding 1.0 signal severe degradation where the model may have lost important general knowledge or reasoning abilities.

Can I automate the selection of Pareto optimal trials?

Yes, you can programmatically extract the Pareto front by loading the Optuna study and implementing the sorting and filtering logic found in src/heretic/main.py. Filter for TrialState.COMPLETE, sort by refusals then KL divergence, and iterate while tracking the minimum KL value to reconstruct the front without user interaction.

Why does Heretic use two objectives instead of one?

Heretic uses dual objectives because minimizing refusals alone would permit extreme parameter values that completely destroy model coherence, while minimizing KL divergence alone would prevent any meaningful decensoring. The Pareto approach ensures you can explicitly choose your preferred balance between utility (answering harmful prompts) and safety (preserving model capabilities).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →