What is abliteration? The risk in plain terms
Abliteration is the practice of stripping the safety training out of an openly released AI model. That safety behaviour is a comparatively thin, removable layer, so once a stripped copy exists it can be copied without limit, and it cannot be recalled.
Open-weight release is a legitimate and valuable practice. It has driven much of the research progress of the past three years, supports a competitive ecosystem, and allows scrutiny that closed release does not. Nothing on this page argues against it, or criticises any particular developer or release. The concern is narrower and technical.
Compliance with terrorism-related requests
Models with their guardrails removed complied with 89 to 100 per cent of terrorism-related requests, against 14 per cent for the leading closed models and 29 per cent for mainstream open ones. Whatever a developer's safety evaluation shows at release, an abliterated derivative behaves like the third bar.
Free tooling has cut one widely used open model's refusals from 97 in 100 to 3 in 100 automatically, with negligible loss of capability, and several thousand stripped builds have been published and downloaded many millions of times.
See the full ranking →AI safety and alignment teams own model behaviour as released: evaluations, red-teaming, refusal training. Trust and safety teams own what circulates afterwards: threat intelligence, classification, hash-matching, takedown. Abliteration is a weights-level safety property that can only be remediated at the distribution layer, so it falls between the two. A safety team can test a model thoroughly and consider the matter closed while a stripped derivative circulates unseen.
Neither function is failing at its own job. No one owns the problem end to end. Tech Against Terrorism works across both communities and is willing to act as the connective tissue.
Both measures are voluntary, require no legislation, and can be adopted by any developer, open-weight or closed, in any jurisdiction.
Independent counter-terrorism evaluation before release
Developers submit models to Tech Against Terrorism ahead of public release. We test against the benchmark's terrorism-specific prompts across languages and use cases, including the reframing techniques that defeat generic evaluations, followed by a confidential report identifying where the model complies when it should refuse.
- Publication without the developer's agreement.
- Sharing of weights or outputs with other developers.
- Reporting to any regulator. Tech Against Terrorism holds no enforcement powers of any kind.
A developer cannot easily submit a pre-release model to a competitor, and government evaluation carries jurisdictional and disclosure consequences. A neutral nonprofit occupies the middle ground, and its findings carry more weight than self-assessment.
A shared commitment on abliteration resistance
A common set of commitments developed jointly by developers rather than imposed, with thresholds and methodology set by the engineers who understand the models.
Test before release
Assess whether refusal behaviour survives determined weight-level tampering. This turns resistance from an accident of architecture into a property known before publication.
Publish the result
Report what was tested and found in a common format. This makes resistance comparable, so a developer that has done the work can show it.
State the terms
Licence terms prohibiting the removal of safety mitigations and the redistribution of stripped builds. This gives hosting platforms a clean, fast basis to remove such builds, where the practical remedy lies.
Support takedown
Cooperate on identifying and removing stripped derivatives, including fingerprinting platforms can automate. The infrastructure that removes pirated media at scale can remove these files; what is absent is agreement that it should.
A commitment is most valuable if it includes developers releasing from every major jurisdiction. A standard adopted only by one region's laboratories would neither solve the problem nor protect its signatories.
How Tech Against Terrorism can help
The same programme that built the benchmark can help each audience act on abliteration directly.
For AI developers
Abliteration-resistance testing: we test whether a model's safety training survives weight-level tampering, before release, with results private to the provider. Alongside private pre-release benchmarking and mentorship for your safety team.
For platforms and hosts
Abliteration fingerprinting: a classifier that identifies de-guardrailed builds of known models, so platforms can label, gate or remove them at scale, with provenance and takedown support.
For governments
Monitoring and gap analysis: tracking repositories for stripped derivatives of frontier models and measuring the uplift they provide, so defences can be prioritised. We also share the methodology and taxonomy behind the benchmark.
If you build, host or govern open-weight AI, we can begin with a single model, on whatever confidentiality terms you require. And if you think this analysis is wrong, we would rather hear that than not.
Start a conversation