heretic
Fully automatic censorship removal for language models
Customize Heretic's chat interface system prompts and model behavior. Learn to modify config.py, use environment variables, or CLI flags for tailored AI interactions.
Understanding Pareto Front Results in Heretic: How to Choose Optimal Decensoring ParametersLearn how Heretic presents Pareto front results balancing refusal reduction and KL divergence. Choose optimal decensoring parameters based on your specific trade-offs for better model control.
How to Interpret and Use Optuna Trial Results in Heretic: A Complete GuideLearn to interpret and use Optuna trial results in Heretic. Understand hyperparameters, KL divergence, and refusal counts to reconstruct your model state with this complete guide.
How to Add Support for New Model Architectures in Heretic: A Complete GuideLearn how to add support for new model architectures in Heretic by modifying three key Python functions. This guide simplifies integrating custom models into the Heretic framework.
How to Debug Issues with Loading Prompt Datasets in HereticSolve Heretic prompt dataset loading problems. Debug `DatasetSpecification`, split, and column issues using `load_prompts`. Run loading logic interactively for faster issue resolution.
How to Upload Decensored Models to Hugging Face Using HereticUpload decensored models to Hugging Face securely with Heretic. This guide shows you how to use the interactive CLI to push adapters or merged models without saving credentials.
How to Configure Multi-GPU and Device Mapping for Heretic: A Complete GuideConfigure multi-GPU and device mapping for Heretic using device_map and max_memory in TOML, CLI, or env variables. Optimize your performance with this complete guide.
Row Normalization Strategies in Heretic: A Complete Guide to LoRA AbliterationExplore Heretic's three row normalization strategies NONE PRE and FULL Understand how to use them to balance computational speed and weight magnitude preservation for LoRA adapters
How Heretic's Winsorization Handles Massive Activations in Language ModelsLearn how Heretic's winsorization effectively handles massive activations in language models by calculating per-layer quantile thresholds and capping extreme values for numerical stability.
Benefits of Using orthogonalize_direction in Heretic for Targeted Model AbliterationLearn how orthogonalize_direction in Heretic removes harmful refusal components while preserving model helpfulness for targeted abliteration. Improve model safety today.
Understanding direction_index and Layer Selection in Heretic's Ablation ScopeMaster Heretic's ablation scope with precise control over direction_index and layer selection. Learn how these parameters fine-tune your model's refusal directions for optimal results.
How to Tune Batch Size and Optimize Performance in HereticLearn how to tune batch size and optimize performance in Heretic. automatically benchmark for optimal throughput or set a max batch size to prevent memory errors. Boost your computational efficiency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →