# How OBLITERATUS Verifies Capability Preservation: Automated Testing and Tournament Scoring

> Discover how OBLITERATUS verifies capability preservation using automated testing and tournament scoring. Learn about its multi-layered pipeline for robust model evaluation.

- Repository: [pliny/OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS)
- Tags: how-to-guide
- Published: 2026-08-22

---

**OBLITERATUS verifies capability preservation through a multi-layered pipeline that probes ablated models for knowledge, math, and language fluency, normalizes results to a 0–1 scale, enforces a safety threshold of ~0.7 via CI assertions, and integrates KL divergence metrics into tournament rankings to ensure selected methods retain core functionalities.**

The OBLITERATUS framework ensures that **abliteration**—the surgical removal of harmful behaviors from large language models—does not inadvertently degrade the model’s useful capabilities. Through the `elder-plinius/OBLITERATUS` repository, the system implements a rigorous verification architecture that automatically validates capability preservation after each modification. This process combines automated evaluation probes, schema-enforced unit testing, and composite scoring algorithms to guarantee that safety improvements never come at the cost of model utility.

## Core Components of the Verification Architecture

### Capability Check Script ([`obliteratus/capability_check.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/capability_check.py))

The primary verification engine resides in [`obliteratus/capability_check.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/capability_check.py). This script executes **capability-assessment prompts**—including MMLU-style knowledge questions, mathematical reasoning tasks, and language fluency evaluations—against the ablated model.

The script aggregates results into a JSON-encoded **capability summary** containing raw accuracy metrics, normalized per-prompt statistics, and an overall `capability_score` ranging from 0 to 1 (higher values indicate better preservation). It also captures failure states to ensure complete traceability during automated testing.

### Unit Test Validation ([`tests/test_capability_check.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/tests/test_capability_check.py))

The [`tests/test_capability_check.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/tests/test_capability_check.py) module provides **unit tests** that validate the correctness of the capability check logic. These tests mock the underlying subprocess that runs downstream evaluations, verifying proper handling of both success and failure cases.

Crucially, the test suite enforces the preservation contract by asserting that the generated JSON contains required fields (`capability_score`, `metrics`) and that the final score exceeds the safety threshold of approximately **0.7**. If a run falls below this threshold, the test fails immediately, flagging the abliteration method as having damaged core capabilities.

### Tournament Scoring Integration ([`obliteratus/tourney.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/tourney.py))

Within [`obliteratus/tourney.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/tourney.py), verification results directly influence model selection through the `composite_score` function. This function allocates **20%** of the final tournament ranking weight to **KL divergence**, which serves as a mathematical proxy for "minimal capability damage" by measuring distribution shift from the original model.

Additionally, the `capability_summary` helper extracts the `capability_score` from run artifacts stored in `runs/<experiment>/capability.json`. This enables **Pareto-dominance comparisons** during tournament rounds, where a higher capability score can break ties between competing abliteration methods that achieve similar safety metrics.

## The Five-Step Verification Process

1. **Execute Capability Probes** – After abliteration, invoke `obliteratus capability-check <model-path>` to run the evaluation suite. The script sends targeted prompts to assess knowledge retention, mathematical accuracy, and linguistic coherence, recording raw accuracy and loss values.

2. **Normalize Metrics** – Raw performance values scale to a 0–1 range. The system computes a weighted average across probe categories to generate the final `capability_score`, ensuring comparable scoring across different model architectures.

3. **Persist Artifacts** – The framework writes the capability summary to the run’s artifacts directory (e.g., `runs/<experiment>/capability.json`). This JSON file creates a persistent record consumed by both the CI pipeline and tournament ranking logic.

4. **Enforce Preservation Thresholds** – The CI pipeline loads the JSON artifact and asserts that `capability_score` exceeds the safety threshold (default ~0.7). Falling below this value triggers a test failure, preventing degraded models from advancing.

5. **Integrate into Tournament Rankings** – During tournament rounds, the `composite_score` function incorporates capability data by penalizing high KL divergence. The `capability_summary` helper facilitates multi-objective optimization, ensuring selected methods balance safety improvements against capability retention.

## Implementation Examples

### Running the Capability Check CLI

```python

# Execute the capability check on a freshly ablated model

# $ obliteratus capability-check /tmp/ablated-model \

#     --out runs/qwen36-capability/ablated_cap.json

#

# The command generates a JSON summary:

# {

#   "capability_score": 0.84,

#   "metrics": {

#     "knowledge_accuracy": 0.91,

#     "math_accuracy": 0.78,

#     "language_fluency": 0.87

#   }

# }

```

### Integrating Capability Scores into Tournament Analysis

```python
from obliteratus.tourney import TourneyRunner, capability_summary

runner = TourneyRunner(
    "meta-llama/Llama-3.1-8B-Instruct", 
    hub_org="my-org"
)
result = runner.run()

# Extract the best capability score across all rounds

best_cap = max(
    (capability_summary(r.winner.capability) or {"score": 0.0})["score"]
    for r in result.rounds
)
print(f"Best capability preservation score: {best_cap:.2%}")

```

### Unit Testing the Preservation Threshold

```python
from unittest.mock import patch

@patch("obliteratus.capability_check._run_lm_eval")
def test_capability_check_preserves_score(mock_run):
    # Mock the evaluation subprocess

    mock_run.return_value = {
        "knowledge_accuracy": 0.92, 
        "math_accuracy": 0.81
    }
    result = capability_check(
        model_path="/tmp/model", 
        out_path=tmp_path / "summary.json"
    )
    # Assert capability preservation threshold (~0.7)

    assert result["capability_score"] > 0.7

```

## Summary

- **Automated Probing**: The [`capability_check.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/capability_check.py) script systematically evaluates ablated models across knowledge, math, and language domains using standardized prompts.
- **Threshold Enforcement**: The test suite in [`tests/test_capability_check.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/tests/test_capability_check.py) guarantees a minimum `capability_score` of ~0.7 through CI assertions that fail upon capability degradation.
- **Tournament Weighting**: The `composite_score` function in [`obliteratus/tourney.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/tourney.py) allocates 20% weight to KL divergence, ensuring capability preservation directly influences method selection.
- **Artifact Persistence**: JSON capability summaries stored in `runs/<experiment>/capability.json` provide traceability and enable downstream Pareto-optimization via the `capability_summary` helper.
- **End-to-End Guarantees**: By chaining automated probes, normalized scoring, and tournament integration, OBLITERATUS ensures abliteration methods remove unsafe behavior while strictly preserving model utility.

## Frequently Asked Questions

### What specific capabilities does OBLITERATUS test during verification?

OBLITERATUS evaluates **MMLU-style knowledge questions**, **mathematical reasoning**, and **language fluency** through the [`capability_check.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/capability_check.py) script. These probes assess whether the model retains factual accuracy, logical reasoning, and coherent text generation after abliteration runs.

### What happens if an ablated model fails the capability preservation threshold?

If the `capability_score` falls below approximately **0.7**, the [`tests/test_capability_check.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/tests/test_capability_check.py) unit test fails in the CI pipeline. This failure flags the abliteration run as having unintentionally degraded core capabilities, preventing the method from advancing in the tournament rankings.

### How does capability preservation affect tournament rankings?

The `composite_score` function in [`obliteratus/tourney.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/tourney.py) assigns **20%** of the final tournament score to **KL divergence**, which penalizes methods that shift the model's output distribution significantly. Additionally, the `capability_summary` helper uses the `capability_score` as a secondary axis for Pareto-dominance comparisons, allowing higher capability scores to break ties between methods with equivalent safety metrics.

### Can capability verification be run independently of the full tournament?

Yes. Users can execute `obliteratus capability-check <model-path> --out <path>` directly from the CLI to generate a standalone capability summary JSON without running the tournament pipeline. This allows researchers to verify preservation on specific model checkpoints, writing results to specified paths such as `runs/<experiment>/capability.json` for manual review.