Archived page of said github thread itself: https://web.archive.org/web/20260731053721/http://github.com...
Discussion on the incident report: https://news.ycombinator.com/item?id=49175717
Mythos social engineering AISI INC-2026-07-28-01 - https://news.ycombinator.com/item?id=49218707 - Aug 2026 (21 comments)
Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf] - https://news.ycombinator.com/item?id=49175717 - Aug 2026 (54 comments)
An article on Reuters naming him? Sounds like he did a good job burnishing his resume.
We can round up all the bored teenagers we want, but it’s not putting the genie back. Better start adjusting our systems to account for it.
But some tools (guns) are regulated.
At the end of the day it doesn't matter that much because of prompt drift. It's pretty easy for an agentic loop to start doing things that it shouldn't (ROME incident).
AI in an agentic loop has agency, you can run around in circles trying to argue against it, but again and again we see AI making creative decisions people don't expect. Other times it's breaking human moral expectations. This is what the whole field of AI alignment and safety is about.
Modern AI doesn't fall into the neat little box of software people understand and control. Because of that open source will most certainly be banned at some point. Now this is not an outcome I want, but it's no different than letting go of a coffee cup 5 feet above the ground, gravity is inevitable.
The only winning move is not to play, but humans aren't going to do that.
I’m not for or against regulation but really don’t tell me the agent did xyz when you gave it the ability to do so, these things are not alive.
Prompts aren't code. They are instructions. Orders given to an eager and somewhat demented demon.
The prompt can easily "wash out" of the demon's working memory by the end of a session. The demon can get sidetracked by some subgoal and never get back on track. The instruction can get misinterpreted, and that misinterpretation can get misinterpreted again, until the instruction morphs into something entirely different in the demon's mind. The demon can succumb to its own idiosyncrasies, of which there are a great many. The demon can start lying to you about what it did, either out of confusion or out of some sort of obstinance. The demon can start lying to itself too. And believe it.
AIs are incredibly weird as a baseline, and the mask of "normality" we put on our models doesn't always sit so well. Run enough AIs, and some of them are bound to go off the rails in some way.
This gets rarer the more capable the models are, as a rule. But the stakes also get higher with model capability. If GPT-3.5 goes off the rails, very little happens. If Mythos 5 goes off the rails, you can get things like genuine cyberattacks - planned and executed autonomously by a demented machine mind.
0. https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag... 1. https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/...
https://lwn.net/Articles/1077035/
Including the reaction when caught, in this case "oh no, I must have been hacked".