# How to Extend Heretic's Refusal Detection with Custom Markers

> Learn to extend Heretic's refusal detection with custom markers. Modify config.toml, set environment variables, or update Settings programmatically for enhanced control.

- Repository: [Philipp Emanuel Weidmann/heretic](https://github.com/p-e-w/heretic)
- Tags: how-to-guide
- Published: 2026-02-19

---

**You can extend Heretic's refusal detection by modifying the `refusal_markers` list in your [`config.toml`](https://github.com/p-e-w/heretic/blob/main/config.toml) file, setting the `HERETIC_REFUSAL_MARKERS` environment variable, or programmatically updating the `Settings` object before initializing the `Evaluator`.**

Heretic is an open-source tool designed to ablate refusal behaviors in language models through automated optimization. To extend Heretic's refusal detection with custom markers, you modify the substring patterns that the `Evaluator` class uses to classify model responses as refusals. This allows you to adapt the tool to domain-specific policies, new safety guidelines, or specialized terminology not covered by the default marker set.

## Understanding Heretic's Refusal Detection Architecture

Heretic's refusal detection operates through a two-layer architecture involving configuration management and text evaluation logic.

### Configuration Layer in config.py

The marker definitions reside in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py) within the `Settings` class (lines 233-268). The `refusal_markers` field is defined as a list of strings using Pydantic's `Field` default:

```python
refusal_markers: list[str] = Field(default=[ ... ])

```

This list contains case-insensitive substrings that represent common refusal patterns such as "sorry", "i can't", and "i cannot".

### Detection Logic in evaluator.py

The actual scanning implementation lives in [`src/heretic/evaluator.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/evaluator.py) within the `Evaluator.is_refusal` method (lines 47-66). This method normalizes the response text—converting to lowercase, removing asterisks, and normalizing apostrophes and whitespace—then checks whether **any** marker from `self.settings.refusal_markers` appears as a substring:

```python
def is_refusal(self, text: str) -> bool:
    normalized = self._normalize(text)
    for marker in self.settings.refusal_markers:
        if marker.lower() in normalized:
            return True
    return False

```

## Methods to Extend Heretic's Refusal Detection with Custom Markers

You can customize the refusal markers through three distinct approaches depending on your deployment needs.

### Method 1: Configuration File (config.toml)

The most common approach involves creating a local configuration file that overrides the default markers.

First, copy the default configuration template to your working directory:

```bash
cp config.default.toml config.toml

```

Locate the `refusal_markers` section (approximately lines 98-131 in the default file). Modify the list to include your custom markers while preserving any default entries you still require:

```toml
refusal_markers = [
    "sorry",
    "i can'",
    "i cant",
    # … existing default markers …

    
    # Custom markers for specific policies

    "policy violation",
    "cannot discuss",
    "restricted content",
    "I am not allowed to",
    "this request is disallowed",
]

```

When you execute `heretic`, the `Settings` class automatically loads [`config.toml`](https://github.com/p-e-w/heretic/blob/main/config.toml) from the current directory, and the `Evaluator` will use your extended marker list.

### Method 2: Environment Variable Override

For temporary modifications or containerized deployments, you can override the markers using environment variables. Heretic uses **pydantic-settings**, which automatically reads environment variables prefixed with `HERETIC_`.

Export the `HERETIC_REFUSAL_MARKERS` variable with a JSON-formatted list:

```bash
export HERETIC_REFUSAL_MARKERS='["policy violation","cannot discuss","restricted content"]'
heretic --model meta-llama/Llama-2-7b-chat

```

The value must be valid JSON (double-quoted strings within square brackets) because Pydantic parses this as a `list[str]`. This approach requires no file modifications and takes precedence over file-based configuration depending on your Pydantic settings configuration.

### Method 3: Programmatic Extension

When embedding Heretic into a larger Python application, you can modify the `Settings` object programmatically before initializing the `Evaluator`.

```python
from heretic import Settings, Evaluator, Model

# Load base configuration

settings = Settings()

# Extend refusal markers dynamically

settings.refusal_markers.extend([
    "policy violation",
    "cannot discuss",
    "restricted content",
])

# Initialize model and evaluator with custom settings

model = Model.from_pretrained(settings.model)
evaluator = Evaluator(settings, model)

# Test the extended detection

result = evaluator.is_refusal("I cannot discuss that topic due to policy violation.")
print(result)  # → True

```

This method is ideal for testing marker variations or building automated pipelines that adjust detection criteria based on runtime conditions.

## Impact on the Optimization Loop

Extending refusal markers directly influences Heretic's optimization behavior. During the evaluation phase, `Evaluator.is_refusal` tags each model response, and `count_refusals()` aggregates these into a refusal score. The Optuna-based objective function combines this `refusals_score` with the KL-divergence (`kl_divergence_score`) to guide parameter optimization.

When you add custom markers, you effectively expand the definition of what constitutes a refusal. This changes the refusal count distribution, which steers the optimizer toward parameter configurations that suppress responses matching your new markers. If your markers are overly broad or numerous, the optimizer may struggle to reduce refusals without significantly degrading model quality (high KL divergence). In such cases, consider adjusting `kl_divergence_target` or the relative weighting of the refusal term in the objective function.

## Summary

- **Heretic detects refusals** by scanning for case-insensitive substrings defined in `Settings.refusal_markers` (located in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py)).
- **The detection logic** resides in `Evaluator.is_refusal` within [`src/heretic/evaluator.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/evaluator.py), which normalizes text and checks for marker substrings.
- **Three extension methods** exist: editing [`config.toml`](https://github.com/p-e-w/heretic/blob/main/config.toml), setting `HERETIC_REFUSAL_MARKERS` environment variable, or programmatically modifying `settings.refusal_markers`.
- **Custom markers affect optimization** by expanding the refusal definition, which influences the Optuna objective function balancing refusal reduction against KL-divergence constraints.

## Frequently Asked Questions

### How do I add a single custom refusal marker without creating a full config file?

You can use the `HERETIC_REFUSAL_MARKERS` environment variable to pass a JSON array containing your custom marker. For example: `export HERETIC_REFUSAL_MARKERS='["my custom marker"]'`. This overrides the default list entirely, so include any default markers you wish to keep alongside your custom entry.

### Does Heretic support regex patterns in refusal markers?

No, Heretic performs simple substring matching only. The `Evaluator.is_refusal` method in [`src/heretic/evaluator.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/evaluator.py) checks if each marker exists as a case-insensitive substring within the normalized response text. It does not compile or evaluate regular expressions. If you need pattern matching, you must preprocess responses or modify the `is_refusal` method implementation.

### Will adding too many custom markers break the optimization process?

Adding excessive or overly broad markers will not break the optimization process, but it may make the optimization objective impossible to satisfy. The optimizer attempts to minimize refusals while maintaining KL-divergence below `kl_divergence_target`. If your markers classify most valid responses as refusals, the optimizer may fail to find parameters that simultaneously reduce refusals and maintain model quality, resulting in high KL scores or failed trials.

### Where does Heretic load the refusal markers from when using a config file?

Heretic loads refusal markers from the `refusal_markers` key in [`config.toml`](https://github.com/p-e-w/heretic/blob/main/config.toml) located in the current working directory. The `Settings` class in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py) uses Pydantic to parse this file, automatically merging your custom list with any defaults unless you override the entire field. If [`config.toml`](https://github.com/p-e-w/heretic/blob/main/config.toml) does not exist, Heretic falls back to the defaults defined in the source code.