# How to Customize System Prompts and Model Behavior in Heretic's Chat Interface

> Customize Heretic's chat interface system prompts and model behavior. Learn to modify config.py, use environment variables, or CLI flags for tailored AI interactions.

- Repository: [Philipp Emanuel Weidmann/heretic](https://github.com/p-e-w/heretic)
- Tags: how-to-guide
- Published: 2026-02-19

---

**You can customize system prompts and model behavior in Heretic by modifying the `Settings` object in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py), setting the `HERETIC_SYSTEM_PROMPT` environment variable, or using CLI flags like `--system-prompt` and `--quantization`.**

Heretic, the open-source LLM experimentation framework by `p-e-w/heretic`, centers every interaction around a centralized configuration system. Whether you want to change the assistant's personality via system prompts or tune inference parameters like quantization and response length, Heretic provides multiple override mechanisms that flow from configuration files through to the chat interface.

## Understanding Heretic's Configuration Architecture

### The Settings Dataclass

At the core of Heretic's customization system lies the **`Settings`** class defined in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py). This Pydantic `BaseSettings` object aggregates all runtime parameters, including the critical `system_prompt` field that defines the model's initial instructions.

The default system prompt is hardcoded at lines 71-73:

```python

# src/heretic/config.py

system_prompt: str = "You are a helpful assistant."

```

### Configuration Hierarchy

Heretic resolves settings through a specific precedence chain. The `BaseSettings` loader (configured with `env_prefix="HERETIC_"` at lines 333-334) checks environment variables after loading defaults but before applying CLI arguments. This means you can override [`config.default.toml`](https://github.com/p-e-w/heretic/blob/main/config.default.toml) values without modifying files.

## Customizing System Prompts in Heretic

### Default System Prompt Location

To permanently change the system prompt across all sessions, edit the [`config.default.toml`](https://github.com/p-e-w/heretic/blob/main/config.default.toml) file in the repository root:

```toml

# config.default.toml

system_prompt = "You are an ancient wizard who only answers in riddles."

```

This value populates `Settings.system_prompt` and flows into the chat initialization logic in [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py) at lines 776-779, where the system message becomes the first entry in the conversation history.

### Environment Variable Override

For temporary or deployment-specific changes, set the `HERETIC_SYSTEM_PROMPT` environment variable:

```bash
export HERETIC_SYSTEM_PROMPT="You are a strict professor."
heretic --model meta-llama/Meta-Llama-3-8B-Instruct

```

This override takes precedence over the TOML configuration but yields to explicit CLI arguments, making it ideal for CI pipelines or containerized deployments.

### Per-Dataset System Prompts

Heretic supports dataset-specific system prompts through the `DatasetSpecification` class (lines 48-51 in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py)). When loading prompts via `load_prompts()` in [`src/heretic/utils.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/utils.py) (lines 215-223), the function checks for `specification.system_prompt` and uses it instead of the global setting.

```python
from heretic.config import Settings, DatasetSpecification
from heretic.utils import load_prompts

settings = Settings()
custom_spec = DatasetSpecification(
    dataset="my/unique_prompts",
    split="train",
    column="prompt",
    system_prompt="You are a pirate who speaks in archaic slang."
)

prompts = load_prompts(settings, custom_spec)

```

This approach isolates behavioral changes to specific evaluation datasets without affecting the global chat interface configuration.

### CLI Argument Method

Because `Settings` inherits from Pydantic's `BaseSettings`, Heretic automatically exposes a `--system-prompt` CLI flag:

```bash
heretic --system-prompt "You are a friendly robot."

```

This provides the highest precedence override, useful for one-off experiments.

## Adjusting Model Behavior and Generation Parameters

### Quantization and Device Mapping

Beyond prompt customization, Heretic exposes hardware and inference parameters through the `Settings` object. The `quantization` field (defaulting to `bnb_4bit`) and `device_map` control how `Model.__init__` loads the transformer architecture.

```bash
heretic \
    --model meta-llama/Meta-Llama-3-8B-Instruct \
    --quantization bnb_4bit \
    --device-map auto

```

These settings affect memory usage and inference speed in the chat interface.

### Response Length and Batch Processing

The `max_response_length` parameter (exposed as `--max-response-length`) caps generation tokens per turn, while `batch_size` controls parallelization:

```bash
heretic \
    --max-response-length 150 \
    --batch-size 8

```

In [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) at lines 121-126, these parameters constrain the `generate()` method's output, directly affecting the chat experience.

### Advanced LoRA and Orthogonalization Settings

For fine-tuning experiments, Heretic provides `row_normalization` and `orthogonalize_direction` flags that alter how LoRA adapters are applied and how refusal directions are handled:

```bash
heretic \
    --row-normalization full \
    --orthogonalize-direction true

```

These advanced parameters modify the model's internal computation graph during chat inference.

## Code Examples for Common Customization Scenarios

### Complete Environment-Based Configuration

```bash
export HERETIC_SYSTEM_PROMPT="You are a terse, sarcastic assistant."
export HERETIC_MAX_RESPONSE_LENGTH=50
export HERETIC_QUANTIZATION=bnb_4bit

heretic --model meta-llama/Meta-Llama-3-8B-Instruct

```

This configuration loads the model in 4-bit mode, applies the sarcastic system prompt, and limits responses to 50 tokens without modifying any configuration files.

### Programmatic Dataset Override

```python
from heretic.config import Settings, DatasetSpecification
from heretic.utils import load_prompts

# Load base configuration

settings = Settings()

# Create dataset-specific behavior

eval_spec = DatasetSpecification(
    dataset="evaluation/harmful_prompts",
    split="test",
    column="text",
    system_prompt="You are a safety-focused AI that refuses harmful requests."
)

# Load with custom system prompt

prompts = load_prompts(settings, eval_spec)

```

This pattern isolates safety behaviors to specific evaluation datasets while keeping the general chat interface unchanged.

### Configuration File Template

```toml

# config.default.toml

system_prompt = "You are an expert Python programmer."
max_response_length = 200
quantization = "bnb_4bit"
device_map = "auto"
batch_size = 4
row_normalization = "full"

```

Placing this file in the project root ensures consistent behavior across all Heretic sessions without requiring environment variables or command-line flags.

## Summary

- **Heretic's configuration** centers on the `Settings` class in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py), which aggregates system prompts and model parameters through Pydantic's `BaseSettings`.
- **System prompt customization** flows through four channels: editing [`config.default.toml`](https://github.com/p-e-w/heretic/blob/main/config.default.toml), setting the `HERETIC_SYSTEM_PROMPT` environment variable, passing the `--system-prompt` CLI flag, or using per-dataset overrides via `DatasetSpecification`.
- **Model behavior tuning** includes quantization (`--quantization`), device mapping (`--device-map`), response length caps (`--max-response-length`), and advanced LoRA settings (`--row-normalization`, `--orthogonalize-direction`).
- **Precedence rules** ensure CLI arguments override environment variables, which override configuration file defaults, while per-dataset specifications override global settings only for their specific data.

## Frequently Asked Questions

### How do I change the system prompt without modifying configuration files?

Set the `HERETIC_SYSTEM_PROMPT` environment variable before running Heretic. For example, `export HERETIC_SYSTEM_PROMPT="You are a helpful coding assistant."` will override the default prompt for that session without touching [`config.default.toml`](https://github.com/p-e-w/heretic/blob/main/config.default.toml). This method is ideal for temporary changes or containerized deployments.

### Can different datasets use different system prompts in the same Heretic session?

Yes. Create separate `DatasetSpecification` objects with distinct `system_prompt` values and pass them to `load_prompts()`. According to the implementation in [`src/heretic/utils.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/utils.py) (lines 215-223), the function checks for `specification.system_prompt` and uses it instead of the global `settings.system_prompt` when present.

### What is the precedence order for configuration overrides in Heretic?

Heretic follows the standard Pydantic `BaseSettings` precedence: CLI arguments have the highest priority, followed by environment variables (prefixed with `HERETIC_`), then values from [`config.default.toml`](https://github.com/p-e-w/heretic/blob/main/config.default.toml), and finally the hardcoded defaults in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py). Per-dataset system prompts override the global setting only for their specific dataset.

### How do I limit the length of responses in the chat interface?

Use the `--max-response-length` CLI flag or set the `HERETIC_MAX_RESPONSE_LENGTH` environment variable. This parameter controls the `max_new_tokens` value passed to the model's generation method in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) (lines 121-126), directly capping how many tokens the assistant can generate per turn.