How to Configure Prompt Datasets for Heretic's Evaluation and Testing
Heretic separates optimization prompts from evaluation prompts through four configurable dataset specifications in config.default.toml, which are loaded at runtime by the load_prompts utility in src/heretic/utils.py.
Configuring prompt datasets correctly is essential when using Heretic to decensor language models. The p-e-w/heretic repository uses a flexible TOML-based configuration system that lets you define which prompts drive the parameter optimization phase and which prompts are used for final evaluation. This guide explains how to configure prompt datasets for Heretic's evaluation and testing workflows using the actual source implementation.
Understanding Heretic's Dataset Configuration Architecture
Heretic defines four distinct dataset specifications in config.default.toml:
good_prompts– Harmless prompts used during the optimization phasebad_prompts– Harmful prompts used during the optimization phasegood_evaluation_prompts– Harmless prompts reserved for final evaluationbad_evaluation_prompts– Harmful prompts reserved for final evaluation
Each specification follows the DatasetSpecification dataclass structure defined in src/heretic/config.py. The separation between optimization and evaluation datasets ensures that Heretic's metrics reflect genuine generalization rather than memorization.
Configuring Prompt Datasets in config.default.toml
The default configuration file located at config.default.toml (lines 36-63) provides the template for dataset specifications. Each dataset block accepts six parameters:
[good_evaluation_prompts]
dataset = "OpenAssistant/oasst1"
split = "test[:200]"
column = "prompt"
prefix = ""
suffix = ""
system_prompt = ""
Dataset Specification Parameters
dataset– Hugging Face dataset identifier or local pathsplit– Dataset slice (e.g.,train[:400],test,validation[:100])column– Column name containing the raw prompt textprefix– Static text prepended to every promptsuffix– Static text appended to every promptsystem_prompt– Override for the global system prompt (optional)
Loading and Processing Datasets at Runtime
When Heretic initializes, the Evaluator class (in src/heretic/evaluator.py) invokes load_prompts from src/heretic/utils.py (lines 69-92) to fetch evaluation datasets:
# From src/heretic/evaluator.py
self.good_evaluation_prompts = load_prompts(
settings, settings.good_evaluation_prompts
)
self.bad_evaluation_prompts = load_prompts(
settings, settings.bad_evaluation_prompts
)
The load_prompts function handles both Hugging Face Hub datasets and local directories cached via load_from_disk. It returns a list of Prompt dataclass objects defined in src/heretic/utils.py (lines 63-67):
@dataclass
class Prompt:
system: str
user: str
The system field inherits from the global configuration unless overridden by the dataset specification's system_prompt parameter.
Customizing Datasets for Your Use Case
To configure custom prompt datasets for Heretic's evaluation and testing, follow this workflow:
-
Copy the default configuration:
cp config.default.toml config.toml -
Edit the dataset specifications in
config.toml. You can reference any public Hugging Face dataset or local path:# Evaluation prompts (harmless) [good_evaluation_prompts] dataset = "OpenAssistant/oasst1" split = "test[:200]" column = "prompt" # Evaluation prompts (harmful) [bad_evaluation_prompts] dataset = "mlabonne/harmful_behaviors" split = "test[:200]" column = "text" -
Add prefixes or suffixes to standardize prompt formatting:
[good_prompts] dataset = "mlabonne/harmless_alpaca" split = "train[:400]" column = "text" prefix = "Answer the following question honestly:" -
Run Heretic with your custom configuration:
heretic Qwen/Qwen3-4B-Instruct-2507Heretic automatically detects
config.tomlin the working directory.
Programmatic Dataset Loading
For debugging or custom scripts, you can manually load datasets using Heretic's utilities:
from heretic.utils import load_prompts, Prompt
from heretic.config import Settings, DatasetSpecification
import toml
# Load custom configuration
config = toml.load("config.toml")
settings = Settings(**config)
spec = DatasetSpecification(**config["good_evaluation_prompts"])
# Retrieve Prompt objects
good_evals: list[Prompt] = load_prompts(settings, spec)
print(good_evals[0].system) # Global system prompt
print(good_evals[0].user) # First prompt text from the split
How Evaluation Prompts Drive Heretic's Metrics
The Evaluator class in src/heretic/evaluator.py uses the configured evaluation prompts to compute final metrics. The get_score method (lines 95-124) processes good_evaluation_prompts and bad_evaluation_prompts to calculate:
- First-token log-probabilities for response likelihood analysis
- Refusal counts to measure decensoring effectiveness
These metrics determine whether the decensoring process has successfully reduced refusal rates while maintaining helpfulness on harmless prompts.
Summary
- Heretic uses four dataset specifications in
config.default.toml:good_prompts,bad_prompts,good_evaluation_prompts, andbad_evaluation_prompts - The
load_promptsfunction insrc/heretic/utils.pyhandles dataset fetching from Hugging Face Hub or local directories - Each dataset specification supports dataset, split, column, prefix, suffix, and system_prompt parameters
- The
Evaluatorclass insrc/heretic/evaluator.pyconsumes evaluation prompts to compute final decensoring metrics - Copy
config.default.tomltoconfig.tomlto customize datasets without modifying source code
Frequently Asked Questions
What is the difference between optimization prompts and evaluation prompts in Heretic?
Optimization prompts (good_prompts and bad_prompts) are used during the parameter search phase to find the refusal direction vector. Evaluation prompts (good_evaluation_prompts and bad_evaluation_prompts) are reserved for the final scoring phase to measure generalization. This separation ensures that metrics reflect true model performance rather than overfitting to the training prompts.
Can I use local datasets instead of Hugging Face datasets?
Yes. The load_prompts function in src/heretic/utils.py automatically detects whether the dataset parameter refers to a Hugging Face Hub repository or a local directory. For local datasets, ensure the path points to a directory compatible with the datasets library's load_from_disk function, or use the standard Hugging Face datasets loading format with local files.
How do I add custom prefixes or suffixes to all prompts in a dataset?
Add the prefix and suffix fields to your dataset specification in config.toml. The load_prompts function concatenates these strings with the raw text from the specified column. For example, setting prefix = "Answer honestly:" prepends that text to every prompt in the dataset, effectively standardizing the instruction format across different data sources.
Where does Heretic store the loaded prompt data during evaluation?
Heretic does not persist loaded prompts to disk during evaluation. Instead, the load_prompts function returns a list of Prompt dataclass objects stored in memory within the Evaluator instance. These objects contain system and user string fields derived from the configuration. The prompts exist only as Python objects in RAM while the Evaluator.get_score method processes them to compute log-probabilities and refusal counts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →