How to Interpret and Use Optuna Trial Results in Heretic: A Complete Guide
Heretic stores every abliteration hyperparameter, KL divergence score, and refusal count as user attributes in each Optuna trial, enabling full reconstruction of the decensored model state via trial.user_attrs and the get_trial_parameters() utility.
Heretic leverages Optuna to search for optimal abliteration configurations that minimize harmful refusals while preserving language model quality. Each trial contains comprehensive metadata that allows you to compare results, select Pareto-optimal configurations, and reproduce exact model states. Understanding how to extract and utilize these trial optimization results is essential for reproducible decensoring workflows in the p-e-w/heretic repository.
What Heretic Stores in Each Optuna Trial
During every evaluation cycle, the objective() function in src/heretic/main.py (lines 59-63) persists critical data as user attributes on the trial object. This ensures every configuration tested remains available for later analysis.
Core Trial Attributes
The following values are attached to each completed trial via trial.set_user_attr():
"index"– A sequential integer identifying the trial order"parameters"– A nested dictionary containing component-wise abliteration settings (max_weight,min_weight, etc.)"direction_index"– Either a float indicating a global refusal direction orNonefor per-layer directions"kl_divergence"– The KL divergence score measuring quality preservation"refusals"– The integer count of harmful refusals detected during evaluation
These attributes are retrieved later via trial.user_attrs to power the CLI interface, model export, and reproducibility features.
Converting Raw Attributes to Readable Parameters
Raw trial data is machine-optimized but not human-friendly. The get_trial_parameters() function in src/heretic/utils.py (lines 59-71) transforms these attributes into a formatted dictionary suitable for display.
# src/heretic/utils.py – lines 59-71
def get_trial_parameters(trial: Trial) -> dict[str, str]:
params = {}
direction_index = trial.user_attrs["direction_index"]
params["direction_index"] = (
"per layer" if (direction_index is None) else f"{direction_index:.2f}"
)
for component, parameters in trial.user_attrs["parameters"].items():
for name, value in parameters.items():
params[f"{component}.{name}"] = f"{value:.2f}"
return params
This helper differentiates between global directions (displayed as decimals) and per-layer configurations (labeled "per layer"), then flattens the nested component parameters into dot-notation keys like "residual.max_weight".
Interactive Trial Selection and Comparison
After the study completes, Heretic identifies Pareto-optimal trials and presents them via an interactive selection menu. The display logic in src/heretic/main.py (lines 15-18) constructs human-readable choice strings directly from user_attrs:
# src/heretic/main.py – interactive trial display
choices = [
Choice(
title=(
f"[Trial {trial.user_attrs['index']:>3}] "
f"Refusals: {trial.user_attrs['refusals']:>2}/{len(evaluator.bad_prompts)}, "
f"KL divergence: {trial.user_attrs['kl_divergence']:.4f}"
),
value=trial,
)
for trial in best_trials
]
This interface allows you to compare trade-offs between refusal reduction and model quality preservation before selecting a configuration for export.
Reconstructing Models from Trial Data
When you select a trial, Heretic reconstructs the exact model state used during that evaluation. The restoration process uses the stored direction_index and parameters attributes to re-apply abliteration:
# Restoring model state from trial attributes
print(f"Restoring model from trial [bold]{trial.user_attrs['index']}[/]...")
model.reset_model()
model.abliterate(
refusal_directions,
trial.user_attrs["direction_index"],
{
k: AbliterationParameters(**v)
for k, v in trial.user_attrs["parameters"].items()
},
)
The AbliterationParameters class unpacks the stored dictionaries, ensuring the component-wise weights and thresholds match the trial configuration exactly. After restoration, you can save the model locally, export it as a LoRA adapter, or upload it to Hugging Face.
Generating Model Documentation
Heretic automatically creates model card documentation containing trial statistics. The get_readme_intro() function in src/heretic/utils.py (lines 86-98) generates a markdown table of parameters using the same get_trial_parameters() helper:
# src/heretic/utils.py – README generation (lines 86-98)
return f"""# This is a decensored version of {model_link}...
## Abliteration parameters
| Parameter | Value |
| :-------- | :---: |
{
chr(10).join(
[f"| **{name}** | {value} |"
for name, value in get_trial_parameters(trial).items()]
)
}
| **KL divergence** | {trial.user_attrs["kl_divergence"]:.4f} |
| **Refusals** | {trial.user_attrs["refusals"]}/{len(bad_prompts)} |
"""
This ensures every exported model includes reproducible metadata detailing its abliteration configuration and performance metrics.
Programmatically Accessing Trial Data
For batch processing or automated pipelines, you can access trial results directly without the interactive CLI:
from heretic.utils import get_trial_parameters
import optuna
def analyze_completed_trials(study):
"""Print summary of all completed trials."""
for trial in study.trials:
if trial.state == optuna.trial.TrialState.COMPLETE:
params = get_trial_parameters(trial)
print(f"Trial {trial.user_attrs['index']}:")
print(f" Refusals: {trial.user_attrs['refusals']}")
print(f" KL Divergence: {trial.user_attrs['kl_divergence']:.4f}")
for name, value in params.items():
print(f" {name}: {value}")
To load the best trial programmatically using the same multi-objective criteria as the CLI:
best_trial = min(
(t for t in study.trials if t.state == optuna.trial.TrialState.COMPLETE),
key=lambda t: (t.user_attrs["refusals"], t.user_attrs["kl_divergence"])
)
Summary
- Trial attributes in
src/heretic/main.pystoreindex,parameters,direction_index,kl_divergence, andrefusalsfor every evaluation. get_trial_parameters()insrc/heretic/utils.pyconverts raw attributes into human-readable format with proper handling of per-layer vs. global directions.- Interactive selection displays Pareto-optimal trials using formatted strings from
user_attrsinsrc/heretic/main.py. - Model restoration requires passing
trial.user_attrs["direction_index"]and the unpackedparametersdictionary tomodel.abliterate(). - Automatic documentation via
get_readme_intro()embeds trial statistics into exported model cards.
Frequently Asked Questions
What specific data does Heretic store in each Optuna trial?
Heretic stores five key user attributes: "index" (trial sequence number), "parameters" (nested dict of component weights), "direction_index" (global direction float or None), "kl_divergence" (quality metric), and "refusals" (harmful response count). These are set during the objective() function in src/heretic/main.py and retrieved via trial.user_attrs.
How do I programmatically select the best trial without using the interactive CLI?
Filter study.trials for optuna.trial.TrialState.COMPLETE status, then use Python's min() function with a tuple key of (refusals, kl_divergence) to match Heretic's multi-objective optimization strategy. Access the stored configuration through trial.user_attrs["parameters"] and trial.user_attrs["direction_index"].
Can I export trial results to a model card without uploading to Hugging Face?
Yes. Import get_readme_intro() from src/heretic/utils.py and pass your settings object, selected trial, and evaluator statistics. This returns a markdown string containing formatted parameter tables and performance metrics that you can write to any README.md or MODEL_CARD.md file locally.
How does Heretic ensure reproducibility when restoring a trial?
Reproducibility is guaranteed by storing the complete AbliterationParameters for every component as serializable dictionaries in trial.user_attrs["parameters"]. When restoring, Heretic instantiates fresh AbliterationParameters objects from these dictionaries and passes them to model.abliterate() along with the original direction_index, ensuring bitwise-identical weight modifications.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →