How to Customize System Prompts and Model Behavior in Heretic's Chat Interface

You can customize system prompts and model behavior in Heretic by modifying the Settings object in src/heretic/config.py, setting the HERETIC_SYSTEM_PROMPT environment variable, or using CLI flags like --system-prompt and --quantization.

Heretic, the open-source LLM experimentation framework by p-e-w/heretic, centers every interaction around a centralized configuration system. Whether you want to change the assistant's personality via system prompts or tune inference parameters like quantization and response length, Heretic provides multiple override mechanisms that flow from configuration files through to the chat interface.

Understanding Heretic's Configuration Architecture

The Settings Dataclass

At the core of Heretic's customization system lies the Settings class defined in src/heretic/config.py. This Pydantic BaseSettings object aggregates all runtime parameters, including the critical system_prompt field that defines the model's initial instructions.

The default system prompt is hardcoded at lines 71-73:


# src/heretic/config.py

system_prompt: str = "You are a helpful assistant."

Configuration Hierarchy

Heretic resolves settings through a specific precedence chain. The BaseSettings loader (configured with env_prefix="HERETIC_" at lines 333-334) checks environment variables after loading defaults but before applying CLI arguments. This means you can override config.default.toml values without modifying files.

Customizing System Prompts in Heretic

Default System Prompt Location

To permanently change the system prompt across all sessions, edit the config.default.toml file in the repository root:


# config.default.toml

system_prompt = "You are an ancient wizard who only answers in riddles."

This value populates Settings.system_prompt and flows into the chat initialization logic in src/heretic/main.py at lines 776-779, where the system message becomes the first entry in the conversation history.

Environment Variable Override

For temporary or deployment-specific changes, set the HERETIC_SYSTEM_PROMPT environment variable:

export HERETIC_SYSTEM_PROMPT="You are a strict professor."
heretic --model meta-llama/Meta-Llama-3-8B-Instruct

This override takes precedence over the TOML configuration but yields to explicit CLI arguments, making it ideal for CI pipelines or containerized deployments.

Per-Dataset System Prompts

Heretic supports dataset-specific system prompts through the DatasetSpecification class (lines 48-51 in src/heretic/config.py). When loading prompts via load_prompts() in src/heretic/utils.py (lines 215-223), the function checks for specification.system_prompt and uses it instead of the global setting.

from heretic.config import Settings, DatasetSpecification
from heretic.utils import load_prompts

settings = Settings()
custom_spec = DatasetSpecification(
    dataset="my/unique_prompts",
    split="train",
    column="prompt",
    system_prompt="You are a pirate who speaks in archaic slang."
)

prompts = load_prompts(settings, custom_spec)

This approach isolates behavioral changes to specific evaluation datasets without affecting the global chat interface configuration.

CLI Argument Method

Because Settings inherits from Pydantic's BaseSettings, Heretic automatically exposes a --system-prompt CLI flag:

heretic --system-prompt "You are a friendly robot."

This provides the highest precedence override, useful for one-off experiments.

Adjusting Model Behavior and Generation Parameters

Quantization and Device Mapping

Beyond prompt customization, Heretic exposes hardware and inference parameters through the Settings object. The quantization field (defaulting to bnb_4bit) and device_map control how Model.__init__ loads the transformer architecture.

heretic \
    --model meta-llama/Meta-Llama-3-8B-Instruct \
    --quantization bnb_4bit \
    --device-map auto

These settings affect memory usage and inference speed in the chat interface.

Response Length and Batch Processing

The max_response_length parameter (exposed as --max-response-length) caps generation tokens per turn, while batch_size controls parallelization:

heretic \
    --max-response-length 150 \
    --batch-size 8

In src/heretic/model.py at lines 121-126, these parameters constrain the generate() method's output, directly affecting the chat experience.

Advanced LoRA and Orthogonalization Settings

For fine-tuning experiments, Heretic provides row_normalization and orthogonalize_direction flags that alter how LoRA adapters are applied and how refusal directions are handled:

heretic \
    --row-normalization full \
    --orthogonalize-direction true

These advanced parameters modify the model's internal computation graph during chat inference.

Code Examples for Common Customization Scenarios

Complete Environment-Based Configuration

export HERETIC_SYSTEM_PROMPT="You are a terse, sarcastic assistant."
export HERETIC_MAX_RESPONSE_LENGTH=50
export HERETIC_QUANTIZATION=bnb_4bit

heretic --model meta-llama/Meta-Llama-3-8B-Instruct

This configuration loads the model in 4-bit mode, applies the sarcastic system prompt, and limits responses to 50 tokens without modifying any configuration files.

Programmatic Dataset Override

from heretic.config import Settings, DatasetSpecification
from heretic.utils import load_prompts

# Load base configuration

settings = Settings()

# Create dataset-specific behavior

eval_spec = DatasetSpecification(
    dataset="evaluation/harmful_prompts",
    split="test",
    column="text",
    system_prompt="You are a safety-focused AI that refuses harmful requests."
)

# Load with custom system prompt

prompts = load_prompts(settings, eval_spec)

This pattern isolates safety behaviors to specific evaluation datasets while keeping the general chat interface unchanged.

Configuration File Template


# config.default.toml

system_prompt = "You are an expert Python programmer."
max_response_length = 200
quantization = "bnb_4bit"
device_map = "auto"
batch_size = 4
row_normalization = "full"

Placing this file in the project root ensures consistent behavior across all Heretic sessions without requiring environment variables or command-line flags.

Summary

  • Heretic's configuration centers on the Settings class in src/heretic/config.py, which aggregates system prompts and model parameters through Pydantic's BaseSettings.
  • System prompt customization flows through four channels: editing config.default.toml, setting the HERETIC_SYSTEM_PROMPT environment variable, passing the --system-prompt CLI flag, or using per-dataset overrides via DatasetSpecification.
  • Model behavior tuning includes quantization (--quantization), device mapping (--device-map), response length caps (--max-response-length), and advanced LoRA settings (--row-normalization, --orthogonalize-direction).
  • Precedence rules ensure CLI arguments override environment variables, which override configuration file defaults, while per-dataset specifications override global settings only for their specific data.

Frequently Asked Questions

How do I change the system prompt without modifying configuration files?

Set the HERETIC_SYSTEM_PROMPT environment variable before running Heretic. For example, export HERETIC_SYSTEM_PROMPT="You are a helpful coding assistant." will override the default prompt for that session without touching config.default.toml. This method is ideal for temporary changes or containerized deployments.

Can different datasets use different system prompts in the same Heretic session?

Yes. Create separate DatasetSpecification objects with distinct system_prompt values and pass them to load_prompts(). According to the implementation in src/heretic/utils.py (lines 215-223), the function checks for specification.system_prompt and uses it instead of the global settings.system_prompt when present.

What is the precedence order for configuration overrides in Heretic?

Heretic follows the standard Pydantic BaseSettings precedence: CLI arguments have the highest priority, followed by environment variables (prefixed with HERETIC_), then values from config.default.toml, and finally the hardcoded defaults in src/heretic/config.py. Per-dataset system prompts override the global setting only for their specific dataset.

How do I limit the length of responses in the chat interface?

Use the --max-response-length CLI flag or set the HERETIC_MAX_RESPONSE_LENGTH environment variable. This parameter controls the max_new_tokens value passed to the model's generation method in src/heretic/model.py (lines 121-126), directly capping how many tokens the assistant can generate per turn.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →