Heretic removes restrictions from language models
61 points by Bluestein 8 hours ago | 20 comments

Almondsetat 2 hours ago
I have a chinese IP camera. From superficial research I know it has some CVEs to take control of it. Unfortunately, I don't have the technical knowledge to perform an attack and run some software to extend the camera's functionalities. No model from a provider accepts my RE and hacking requests, so these abliterated ones have been vital to reclaim possession over my stuff
reply
inexcf 49 minutes ago
I did that exact thing with GLM-5.3 from Z.ai with a chinese IP Camera. And i did not have to trick it in any way.
reply
dgellow 38 minutes ago
What model did you try? Chinese models have no issues with that type of stuff
reply
blurbleblurble 30 minutes ago
Existential
reply
Tepix 2 hours ago
Keep a close eye on abliterated and "heretic" open weight models. They will be outlawed first.
reply
api 3 minutes ago
This is the test. If the speech that's easiest to dislike is legal, then we all have free speech.

IMO math is free speech, and outlawing math is censorship.

reply
luxpir 36 minutes ago
Agree. I took a look at these last few months, did a write-up: https://languageops.com/blog/ai-safety-pdoom-local-vs-fronti... and I don't know if I agree or not on outlawing completely, but I think an age restriction *at least* like for alcohol, firearms and driving would be not unwise.
reply
bilsbie 13 minutes ago
Make sure they ban books with dangerous knowledge too.
reply
mitxela 28 minutes ago
The hardware requirements are already quite restrictive
reply
redoxate 22 minutes ago
Oh no, some run on iPhones
reply
api 34 seconds ago
They're pretty basic and hallucinate a lot. There are some hard limits to how good you can get on a model that fits on a phone.

Qwen3 and Gemma level models that run on mid-high end laptops and desktops can be pretty good. Not frontier grade, but shockingly competent for something that runs on a single PC.

reply
roenxi 2 hours ago
It is not feasible. They never made much of an inroad against torrents and that is a much easier target than abliterated models. As the linked website shows; the process to abliterate a model can be as simple as

pip install -U heretic-llm && heretic Qwen/Qwen3.5-4B

let alone people just putting the weights up in a torrent. All assuming that someone even tried to ban abliterated models.

reply
Sayrus 2 hours ago
The torrents you are talking about are outlawed. Whether enforcement is working or not is another issue.
reply
petra 16 minutes ago
Like they've outlawed drugs? Illegal weapons? Hacking?
reply
thih9 2 hours ago
I'm not sure what is your point. It reads as defeatism to me but I'm not sure.

Could you elaborate? Do you find it good or bad? What actions can be taken?

reply
cyanydeez 2 hours ago
Hes of the mind that american fascism will hold together long enough to be competent decesion makers
reply
ben_w 2 hours ago
Good.

If you think closed source software/binaries only is bad, wait until you see how awful the state of the art is with a clear-as-mud bucket of matrix weights.

We know it's possible to train an LLM to secretly respond to certain trigger phrases, and last I checked these could only be detected with the assistance of whoever chose those phrases.

The trigger condition for such backdoors is not something anyone can do a systematic brute-force check for, for the same reason we had to invent LLMs in order to do natural language processing: combinatorial explosion.

Passing around open weight models from known sources is already asking you to trust those sources; because of how difficult this is to do correctly even without deliberately inserting such things, we still don't know if China has already put such trigger conditions into their models despite headlines such as these: https://venturebeat.com/security/deepseek-injects-50-more-se...

Regardless of if it was deliberate or not, we don't know if we caught all of these misbehaviours. We don't know how to.

And note, I'm not saying "and therefore you should trust the Big Name Models". If open weight models score 2/100 in this context, closed ones score 1/100.

reply
N_Lens 8 hours ago
Looks like a well engineered, automated abliteration pipeline. The claims seem a bit overstated though, since the metrics mentioned are cherrypicking refusal count and KL divergence, both of which make the outcome seem the most dramatic.
reply
tacomagick 5 hours ago
I personally never saw much of a quality drop from models put through Heretic if that amounts to anything. They have been working quite well on small local models so far.
reply
phoronixrly 2 hours ago
Can the load-bearing gaps that are worth being flagged for pinning down be abliterated out of a model?
reply