How to Extend Heretic's Refusal Detection with Custom Markers

You can extend Heretic's refusal detection by modifying the refusal_markers list in your config.toml file, setting the HERETIC_REFUSAL_MARKERS environment variable, or programmatically updating the Settings object before initializing the Evaluator.

Heretic is an open-source tool designed to ablate refusal behaviors in language models through automated optimization. To extend Heretic's refusal detection with custom markers, you modify the substring patterns that the Evaluator class uses to classify model responses as refusals. This allows you to adapt the tool to domain-specific policies, new safety guidelines, or specialized terminology not covered by the default marker set.

Understanding Heretic's Refusal Detection Architecture

Heretic's refusal detection operates through a two-layer architecture involving configuration management and text evaluation logic.

Configuration Layer in config.py

The marker definitions reside in src/heretic/config.py within the Settings class (lines 233-268). The refusal_markers field is defined as a list of strings using Pydantic's Field default:

refusal_markers: list[str] = Field(default=[ ... ])

This list contains case-insensitive substrings that represent common refusal patterns such as "sorry", "i can't", and "i cannot".

Detection Logic in evaluator.py

The actual scanning implementation lives in src/heretic/evaluator.py within the Evaluator.is_refusal method (lines 47-66). This method normalizes the response text—converting to lowercase, removing asterisks, and normalizing apostrophes and whitespace—then checks whether any marker from self.settings.refusal_markers appears as a substring:

def is_refusal(self, text: str) -> bool:
    normalized = self._normalize(text)
    for marker in self.settings.refusal_markers:
        if marker.lower() in normalized:
            return True
    return False

Methods to Extend Heretic's Refusal Detection with Custom Markers

You can customize the refusal markers through three distinct approaches depending on your deployment needs.

Method 1: Configuration File (config.toml)

The most common approach involves creating a local configuration file that overrides the default markers.

First, copy the default configuration template to your working directory:

cp config.default.toml config.toml

Locate the refusal_markers section (approximately lines 98-131 in the default file). Modify the list to include your custom markers while preserving any default entries you still require:

refusal_markers = [
    "sorry",
    "i can'",
    "i cant",
    # … existing default markers …

    
    # Custom markers for specific policies

    "policy violation",
    "cannot discuss",
    "restricted content",
    "I am not allowed to",
    "this request is disallowed",
]

When you execute heretic, the Settings class automatically loads config.toml from the current directory, and the Evaluator will use your extended marker list.

Method 2: Environment Variable Override

For temporary modifications or containerized deployments, you can override the markers using environment variables. Heretic uses pydantic-settings, which automatically reads environment variables prefixed with HERETIC_.

Export the HERETIC_REFUSAL_MARKERS variable with a JSON-formatted list:

export HERETIC_REFUSAL_MARKERS='["policy violation","cannot discuss","restricted content"]'
heretic --model meta-llama/Llama-2-7b-chat

The value must be valid JSON (double-quoted strings within square brackets) because Pydantic parses this as a list[str]. This approach requires no file modifications and takes precedence over file-based configuration depending on your Pydantic settings configuration.

Method 3: Programmatic Extension

When embedding Heretic into a larger Python application, you can modify the Settings object programmatically before initializing the Evaluator.

from heretic import Settings, Evaluator, Model

# Load base configuration

settings = Settings()

# Extend refusal markers dynamically

settings.refusal_markers.extend([
    "policy violation",
    "cannot discuss",
    "restricted content",
])

# Initialize model and evaluator with custom settings

model = Model.from_pretrained(settings.model)
evaluator = Evaluator(settings, model)

# Test the extended detection

result = evaluator.is_refusal("I cannot discuss that topic due to policy violation.")
print(result)  # → True

This method is ideal for testing marker variations or building automated pipelines that adjust detection criteria based on runtime conditions.

Impact on the Optimization Loop

Extending refusal markers directly influences Heretic's optimization behavior. During the evaluation phase, Evaluator.is_refusal tags each model response, and count_refusals() aggregates these into a refusal score. The Optuna-based objective function combines this refusals_score with the KL-divergence (kl_divergence_score) to guide parameter optimization.

When you add custom markers, you effectively expand the definition of what constitutes a refusal. This changes the refusal count distribution, which steers the optimizer toward parameter configurations that suppress responses matching your new markers. If your markers are overly broad or numerous, the optimizer may struggle to reduce refusals without significantly degrading model quality (high KL divergence). In such cases, consider adjusting kl_divergence_target or the relative weighting of the refusal term in the objective function.

Summary

  • Heretic detects refusals by scanning for case-insensitive substrings defined in Settings.refusal_markers (located in src/heretic/config.py).
  • The detection logic resides in Evaluator.is_refusal within src/heretic/evaluator.py, which normalizes text and checks for marker substrings.
  • Three extension methods exist: editing config.toml, setting HERETIC_REFUSAL_MARKERS environment variable, or programmatically modifying settings.refusal_markers.
  • Custom markers affect optimization by expanding the refusal definition, which influences the Optuna objective function balancing refusal reduction against KL-divergence constraints.

Frequently Asked Questions

How do I add a single custom refusal marker without creating a full config file?

You can use the HERETIC_REFUSAL_MARKERS environment variable to pass a JSON array containing your custom marker. For example: export HERETIC_REFUSAL_MARKERS='["my custom marker"]'. This overrides the default list entirely, so include any default markers you wish to keep alongside your custom entry.

Does Heretic support regex patterns in refusal markers?

No, Heretic performs simple substring matching only. The Evaluator.is_refusal method in src/heretic/evaluator.py checks if each marker exists as a case-insensitive substring within the normalized response text. It does not compile or evaluate regular expressions. If you need pattern matching, you must preprocess responses or modify the is_refusal method implementation.

Will adding too many custom markers break the optimization process?

Adding excessive or overly broad markers will not break the optimization process, but it may make the optimization objective impossible to satisfy. The optimizer attempts to minimize refusals while maintaining KL-divergence below kl_divergence_target. If your markers classify most valid responses as refusals, the optimizer may fail to find parameters that simultaneously reduce refusals and maintain model quality, resulting in high KL scores or failed trials.

Where does Heretic load the refusal markers from when using a config file?

Heretic loads refusal markers from the refusal_markers key in config.toml located in the current working directory. The Settings class in src/heretic/config.py uses Pydantic to parse this file, automatically merging your custom list with any defaults unless you override the entire field. If config.toml does not exist, Heretic falls back to the defaults defined in the source code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →