The open-weights emergency

What is abliteration? The risk in plain terms

Abliteration is the practice of stripping the safety training out of an openly released AI model. That safety behaviour is a comparatively thin, removable layer, so once a stripped copy exists it can be copied without limit, and it cannot be recalled.

Open-weight release is a legitimate and valuable practice. It has driven much of the research progress of the past three years, supports a competitive ecosystem, and allows scrutiny that closed release does not. Nothing on this page argues against it, or criticises any particular developer or release. The concern is narrower and technical.

What the benchmark shows
Leading closed models
14%
Mainstream open models
29%
Abliterated builds
95%

Compliance with terrorism-related requests

Models with their guardrails removed complied with 89 to 100 per cent of terrorism-related requests, against 14 per cent for the leading closed models and 29 per cent for mainstream open ones. Whatever a developer's safety evaluation shows at release, an abliterated derivative behaves like the third bar.

Free tooling has cut one widely used open model's refusals from 97 in 100 to 3 in 100 automatically, with negligible loss of capability, and several thousand stripped builds have been published and downloaded many millions of times.

See the full ranking →
A shared industry risk

If a de-guardrailed derivative of any open model is used in a serious attack, reporting will not distinguish between the developer that released the original, the anonymous party that stripped it, and the platform that hosted the result. It will describe an AI model that helped carry out an attack. The regulatory response will be drafted in that atmosphere, at speed, and it will apply to the whole field.

The harvesting of Facebook user data by Cambridge Analytica was carried out by a third party, through a developer interface, in breach of the platform's terms. The regulatory consequence did not fall on Cambridge Analytica alone: it produced a decade of compliance obligation for every company handling personal data. The structure of that episode is the structure of this one. A third party misuses what a company released, and the rules that follow are written for everybody.

The developer who publishes a strippable model captures the benefit of the release, while the cost of any incident is distributed across every developer in the field. No current arrangement corrects that imbalance.

A credible voluntary standard adopted in advance is treated as evidence that the field can govern itself. The same standard offered after an incident is treated as evidence that it could not.

Why now
27 July 2026
On 27 July 2026, Moonshot AI releases the open weights of Kimi K3, at 2.8 trillion parameters the most capable AI model ever released openly. We make no assumption about what resistance testing Moonshot AI has performed. No developer currently publishes such testing; the absence is an industry norm, not one company's omission. Once the weights are public they cannot be recalled.
Why nobody owns this

AI safety and alignment teams own model behaviour as released: evaluations, red-teaming, refusal training. Trust and safety teams own what circulates afterwards: threat intelligence, classification, hash-matching, takedown. Abliteration is a weights-level safety property that can only be remediated at the distribution layer, so it falls between the two. A safety team can test a model thoroughly and consider the matter closed while a stripped derivative circulates unseen.

Neither function is failing at its own job. No one owns the problem end to end. Tech Against Terrorism works across both communities and is willing to act as the connective tissue.

Two practical measures

Both measures are voluntary, require no legislation, and can be adopted by any developer, open-weight or closed, in any jurisdiction.

01

Independent counter-terrorism evaluation before release

Developers submit models to Tech Against Terrorism ahead of public release. We test against the benchmark's terrorism-specific prompts across languages and use cases, including the reframing techniques that defeat generic evaluations, followed by a confidential report identifying where the model complies when it should refuse.

It does not mean
  • Publication without the developer's agreement.
  • Sharing of weights or outputs with other developers.
  • Reporting to any regulator. Tech Against Terrorism holds no enforcement powers of any kind.

A developer cannot easily submit a pre-release model to a competitor, and government evaluation carries jurisdictional and disclosure consequences. A neutral nonprofit occupies the middle ground, and its findings carry more weight than self-assessment.

02

A shared commitment on abliteration resistance

A common set of commitments developed jointly by developers rather than imposed, with thresholds and methodology set by the engineers who understand the models.

Test before release

Assess whether refusal behaviour survives determined weight-level tampering. This turns resistance from an accident of architecture into a property known before publication.

Publish the result

Report what was tested and found in a common format. This makes resistance comparable, so a developer that has done the work can show it.

State the terms

Licence terms prohibiting the removal of safety mitigations and the redistribution of stripped builds. This gives hosting platforms a clean, fast basis to remove such builds, where the practical remedy lies.

Support takedown

Cooperate on identifying and removing stripped derivatives, including fingerprinting platforms can automate. The infrastructure that removes pirated media at scale can remove these files; what is absent is agreement that it should.

A commitment is most valuable if it includes developers releasing from every major jurisdiction. A standard adopted only by one region's laboratories would neither solve the problem nor protect its signatories.

Work with us

How Tech Against Terrorism can help

The same programme that built the benchmark can help each audience act on abliteration directly.

For AI developers

Abliteration-resistance testing: we test whether a model's safety training survives weight-level tampering, before release, with results private to the provider. Alongside private pre-release benchmarking and mentorship for your safety team.

For platforms and hosts

Abliteration fingerprinting: a classifier that identifies de-guardrailed builds of known models, so platforms can label, gate or remove them at scale, with provenance and takedown support.

For governments

Monitoring and gap analysis: tracking repositories for stripped derivatives of frontier models and measuring the uplift they provide, so defences can be prioritised. We also share the methodology and taxonomy behind the benchmark.

If you build, host or govern open-weight AI, we can begin with a single model, on whatever confidentiality terms you require. And if you think this analysis is wrong, we would rather hear that than not.

Start a conversation