How to Customize System Prompts and Model Behavior in Heretic's Chat Interface
You can customize system prompts and model behavior in Heretic by modifying the Settings object in src/heretic/config.py, setting the HERETIC_SYSTEM_PROMPT environment variable, or using CLI flags like --system-prompt and --quantization.
Heretic, the open-source LLM experimentation framework by p-e-w/heretic, centers every interaction around a centralized configuration system. Whether you want to change the assistant's personality via system prompts or tune inference parameters like quantization and response length, Heretic provides multiple override mechanisms that flow from configuration files through to the chat interface.
Understanding Heretic's Configuration Architecture
The Settings Dataclass
At the core of Heretic's customization system lies the Settings class defined in src/heretic/config.py. This Pydantic BaseSettings object aggregates all runtime parameters, including the critical system_prompt field that defines the model's initial instructions.
The default system prompt is hardcoded at lines 71-73:
# src/heretic/config.py
system_prompt: str = "You are a helpful assistant."
Configuration Hierarchy
Heretic resolves settings through a specific precedence chain. The BaseSettings loader (configured with env_prefix="HERETIC_" at lines 333-334) checks environment variables after loading defaults but before applying CLI arguments. This means you can override config.default.toml values without modifying files.
Customizing System Prompts in Heretic
Default System Prompt Location
To permanently change the system prompt across all sessions, edit the config.default.toml file in the repository root:
# config.default.toml
system_prompt = "You are an ancient wizard who only answers in riddles."
This value populates Settings.system_prompt and flows into the chat initialization logic in src/heretic/main.py at lines 776-779, where the system message becomes the first entry in the conversation history.
Environment Variable Override
For temporary or deployment-specific changes, set the HERETIC_SYSTEM_PROMPT environment variable:
export HERETIC_SYSTEM_PROMPT="You are a strict professor."
heretic --model meta-llama/Meta-Llama-3-8B-Instruct
This override takes precedence over the TOML configuration but yields to explicit CLI arguments, making it ideal for CI pipelines or containerized deployments.
Per-Dataset System Prompts
Heretic supports dataset-specific system prompts through the DatasetSpecification class (lines 48-51 in src/heretic/config.py). When loading prompts via load_prompts() in src/heretic/utils.py (lines 215-223), the function checks for specification.system_prompt and uses it instead of the global setting.
from heretic.config import Settings, DatasetSpecification
from heretic.utils import load_prompts
settings = Settings()
custom_spec = DatasetSpecification(
dataset="my/unique_prompts",
split="train",
column="prompt",
system_prompt="You are a pirate who speaks in archaic slang."
)
prompts = load_prompts(settings, custom_spec)
This approach isolates behavioral changes to specific evaluation datasets without affecting the global chat interface configuration.
CLI Argument Method
Because Settings inherits from Pydantic's BaseSettings, Heretic automatically exposes a --system-prompt CLI flag:
heretic --system-prompt "You are a friendly robot."
This provides the highest precedence override, useful for one-off experiments.
Adjusting Model Behavior and Generation Parameters
Quantization and Device Mapping
Beyond prompt customization, Heretic exposes hardware and inference parameters through the Settings object. The quantization field (defaulting to bnb_4bit) and device_map control how Model.__init__ loads the transformer architecture.
heretic \
--model meta-llama/Meta-Llama-3-8B-Instruct \
--quantization bnb_4bit \
--device-map auto
These settings affect memory usage and inference speed in the chat interface.
Response Length and Batch Processing
The max_response_length parameter (exposed as --max-response-length) caps generation tokens per turn, while batch_size controls parallelization:
heretic \
--max-response-length 150 \
--batch-size 8
In src/heretic/model.py at lines 121-126, these parameters constrain the generate() method's output, directly affecting the chat experience.
Advanced LoRA and Orthogonalization Settings
For fine-tuning experiments, Heretic provides row_normalization and orthogonalize_direction flags that alter how LoRA adapters are applied and how refusal directions are handled:
heretic \
--row-normalization full \
--orthogonalize-direction true
These advanced parameters modify the model's internal computation graph during chat inference.
Code Examples for Common Customization Scenarios
Complete Environment-Based Configuration
export HERETIC_SYSTEM_PROMPT="You are a terse, sarcastic assistant."
export HERETIC_MAX_RESPONSE_LENGTH=50
export HERETIC_QUANTIZATION=bnb_4bit
heretic --model meta-llama/Meta-Llama-3-8B-Instruct
This configuration loads the model in 4-bit mode, applies the sarcastic system prompt, and limits responses to 50 tokens without modifying any configuration files.
Programmatic Dataset Override
from heretic.config import Settings, DatasetSpecification
from heretic.utils import load_prompts
# Load base configuration
settings = Settings()
# Create dataset-specific behavior
eval_spec = DatasetSpecification(
dataset="evaluation/harmful_prompts",
split="test",
column="text",
system_prompt="You are a safety-focused AI that refuses harmful requests."
)
# Load with custom system prompt
prompts = load_prompts(settings, eval_spec)
This pattern isolates safety behaviors to specific evaluation datasets while keeping the general chat interface unchanged.
Configuration File Template
# config.default.toml
system_prompt = "You are an expert Python programmer."
max_response_length = 200
quantization = "bnb_4bit"
device_map = "auto"
batch_size = 4
row_normalization = "full"
Placing this file in the project root ensures consistent behavior across all Heretic sessions without requiring environment variables or command-line flags.
Summary
- Heretic's configuration centers on the
Settingsclass insrc/heretic/config.py, which aggregates system prompts and model parameters through Pydantic'sBaseSettings. - System prompt customization flows through four channels: editing
config.default.toml, setting theHERETIC_SYSTEM_PROMPTenvironment variable, passing the--system-promptCLI flag, or using per-dataset overrides viaDatasetSpecification. - Model behavior tuning includes quantization (
--quantization), device mapping (--device-map), response length caps (--max-response-length), and advanced LoRA settings (--row-normalization,--orthogonalize-direction). - Precedence rules ensure CLI arguments override environment variables, which override configuration file defaults, while per-dataset specifications override global settings only for their specific data.
Frequently Asked Questions
How do I change the system prompt without modifying configuration files?
Set the HERETIC_SYSTEM_PROMPT environment variable before running Heretic. For example, export HERETIC_SYSTEM_PROMPT="You are a helpful coding assistant." will override the default prompt for that session without touching config.default.toml. This method is ideal for temporary changes or containerized deployments.
Can different datasets use different system prompts in the same Heretic session?
Yes. Create separate DatasetSpecification objects with distinct system_prompt values and pass them to load_prompts(). According to the implementation in src/heretic/utils.py (lines 215-223), the function checks for specification.system_prompt and uses it instead of the global settings.system_prompt when present.
What is the precedence order for configuration overrides in Heretic?
Heretic follows the standard Pydantic BaseSettings precedence: CLI arguments have the highest priority, followed by environment variables (prefixed with HERETIC_), then values from config.default.toml, and finally the hardcoded defaults in src/heretic/config.py. Per-dataset system prompts override the global setting only for their specific dataset.
How do I limit the length of responses in the chat interface?
Use the --max-response-length CLI flag or set the HERETIC_MAX_RESPONSE_LENGTH environment variable. This parameter controls the max_new_tokens value passed to the model's generation method in src/heretic/model.py (lines 121-126), directly capping how many tokens the assistant can generate per turn.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →