How OBLITERATUS Verifies Capability Preservation: Automated Testing and Tournament Scoring
OBLITERATUS verifies capability preservation through a multi-layered pipeline that probes ablated models for knowledge, math, and language fluency, normalizes results to a 0–1 scale, enforces a safety threshold of ~0.7 via CI assertions, and integrates KL divergence metrics into tournament rankings to ensure selected methods retain core functionalities.
The OBLITERATUS framework ensures that abliteration—the surgical removal of harmful behaviors from large language models—does not inadvertently degrade the model’s useful capabilities. Through the elder-plinius/OBLITERATUS repository, the system implements a rigorous verification architecture that automatically validates capability preservation after each modification. This process combines automated evaluation probes, schema-enforced unit testing, and composite scoring algorithms to guarantee that safety improvements never come at the cost of model utility.
Core Components of the Verification Architecture
Capability Check Script (obliteratus/capability_check.py)
The primary verification engine resides in obliteratus/capability_check.py. This script executes capability-assessment prompts—including MMLU-style knowledge questions, mathematical reasoning tasks, and language fluency evaluations—against the ablated model.
The script aggregates results into a JSON-encoded capability summary containing raw accuracy metrics, normalized per-prompt statistics, and an overall capability_score ranging from 0 to 1 (higher values indicate better preservation). It also captures failure states to ensure complete traceability during automated testing.
Unit Test Validation (tests/test_capability_check.py)
The tests/test_capability_check.py module provides unit tests that validate the correctness of the capability check logic. These tests mock the underlying subprocess that runs downstream evaluations, verifying proper handling of both success and failure cases.
Crucially, the test suite enforces the preservation contract by asserting that the generated JSON contains required fields (capability_score, metrics) and that the final score exceeds the safety threshold of approximately 0.7. If a run falls below this threshold, the test fails immediately, flagging the abliteration method as having damaged core capabilities.
Tournament Scoring Integration (obliteratus/tourney.py)
Within obliteratus/tourney.py, verification results directly influence model selection through the composite_score function. This function allocates 20% of the final tournament ranking weight to KL divergence, which serves as a mathematical proxy for "minimal capability damage" by measuring distribution shift from the original model.
Additionally, the capability_summary helper extracts the capability_score from run artifacts stored in runs/<experiment>/capability.json. This enables Pareto-dominance comparisons during tournament rounds, where a higher capability score can break ties between competing abliteration methods that achieve similar safety metrics.
The Five-Step Verification Process
-
Execute Capability Probes – After abliteration, invoke
obliteratus capability-check <model-path>to run the evaluation suite. The script sends targeted prompts to assess knowledge retention, mathematical accuracy, and linguistic coherence, recording raw accuracy and loss values. -
Normalize Metrics – Raw performance values scale to a 0–1 range. The system computes a weighted average across probe categories to generate the final
capability_score, ensuring comparable scoring across different model architectures. -
Persist Artifacts – The framework writes the capability summary to the run’s artifacts directory (e.g.,
runs/<experiment>/capability.json). This JSON file creates a persistent record consumed by both the CI pipeline and tournament ranking logic. -
Enforce Preservation Thresholds – The CI pipeline loads the JSON artifact and asserts that
capability_scoreexceeds the safety threshold (default ~0.7). Falling below this value triggers a test failure, preventing degraded models from advancing. -
Integrate into Tournament Rankings – During tournament rounds, the
composite_scorefunction incorporates capability data by penalizing high KL divergence. Thecapability_summaryhelper facilitates multi-objective optimization, ensuring selected methods balance safety improvements against capability retention.
Implementation Examples
Running the Capability Check CLI
# Execute the capability check on a freshly ablated model
# $ obliteratus capability-check /tmp/ablated-model \
# --out runs/qwen36-capability/ablated_cap.json
#
# The command generates a JSON summary:
# {
# "capability_score": 0.84,
# "metrics": {
# "knowledge_accuracy": 0.91,
# "math_accuracy": 0.78,
# "language_fluency": 0.87
# }
# }
Integrating Capability Scores into Tournament Analysis
from obliteratus.tourney import TourneyRunner, capability_summary
runner = TourneyRunner(
"meta-llama/Llama-3.1-8B-Instruct",
hub_org="my-org"
)
result = runner.run()
# Extract the best capability score across all rounds
best_cap = max(
(capability_summary(r.winner.capability) or {"score": 0.0})["score"]
for r in result.rounds
)
print(f"Best capability preservation score: {best_cap:.2%}")
Unit Testing the Preservation Threshold
from unittest.mock import patch
@patch("obliteratus.capability_check._run_lm_eval")
def test_capability_check_preserves_score(mock_run):
# Mock the evaluation subprocess
mock_run.return_value = {
"knowledge_accuracy": 0.92,
"math_accuracy": 0.81
}
result = capability_check(
model_path="/tmp/model",
out_path=tmp_path / "summary.json"
)
# Assert capability preservation threshold (~0.7)
assert result["capability_score"] > 0.7
Summary
- Automated Probing: The
capability_check.pyscript systematically evaluates ablated models across knowledge, math, and language domains using standardized prompts. - Threshold Enforcement: The test suite in
tests/test_capability_check.pyguarantees a minimumcapability_scoreof ~0.7 through CI assertions that fail upon capability degradation. - Tournament Weighting: The
composite_scorefunction inobliteratus/tourney.pyallocates 20% weight to KL divergence, ensuring capability preservation directly influences method selection. - Artifact Persistence: JSON capability summaries stored in
runs/<experiment>/capability.jsonprovide traceability and enable downstream Pareto-optimization via thecapability_summaryhelper. - End-to-End Guarantees: By chaining automated probes, normalized scoring, and tournament integration, OBLITERATUS ensures abliteration methods remove unsafe behavior while strictly preserving model utility.
Frequently Asked Questions
What specific capabilities does OBLITERATUS test during verification?
OBLITERATUS evaluates MMLU-style knowledge questions, mathematical reasoning, and language fluency through the capability_check.py script. These probes assess whether the model retains factual accuracy, logical reasoning, and coherent text generation after abliteration runs.
What happens if an ablated model fails the capability preservation threshold?
If the capability_score falls below approximately 0.7, the tests/test_capability_check.py unit test fails in the CI pipeline. This failure flags the abliteration run as having unintentionally degraded core capabilities, preventing the method from advancing in the tournament rankings.
How does capability preservation affect tournament rankings?
The composite_score function in obliteratus/tourney.py assigns 20% of the final tournament score to KL divergence, which penalizes methods that shift the model's output distribution significantly. Additionally, the capability_summary helper uses the capability_score as a secondary axis for Pareto-dominance comparisons, allowing higher capability scores to break ties between methods with equivalent safety metrics.
Can capability verification be run independently of the full tournament?
Yes. Users can execute obliteratus capability-check <model-path> --out <path> directly from the CLI to generate a standalone capability summary JSON without running the tournament pipeline. This allows researchers to verify preservation on specific model checkpoints, writing results to specified paths such as runs/<experiment>/capability.json for manual review.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →