There are no "rogue" AI agents
309 points by zzzeek 6 hours ago | 232 comments

elric 5 hours ago
A little over two decades ago, my then girlfriend was arrested for "writing malware" (which was not against the law at the time, and which was never released into the wild and never caused any damage). This set in motion a chain of events that effectively ruined her life.

Fast forward to today, and we have multi billion dollar corporations pumping out malware at breakneck speeds, compromising various systems (including those of foreign governments), and no one is getting arrested. Instead we're gawking at the marvel of these systems and are playing word games about whether or not it's a rogue system. If anything, it's making people richer.

Make it make sense.

reply
fourside 5 hours ago
> no one is getting arrested

Not only have there been no consequences but those same companies are trying to position themselves as the best people to keep these AI systems in check.

reply
BoxOfRain 4 hours ago
Yeah I find the whole 'what if sufficiently advanced AI gets into the wrong hands' argument extremely tiresome when it's already in the worst possible hands as far as most of the human population is concerned.
reply
tonyarkles 4 hours ago
> when it's already in the worst possible hands as far as most of the human population is concerned.

I appreciate the sentiment but I feel like there is a lack of creativity and imagination behind that statement.

reply
franklyopined 2 hours ago
A subset of the builders of the current generation of AI systems literally hold death cult beliefs i.e. that all humans should effectuate the transition to a world run by only AI, without humans at all. These are the worst people in the world to be driving advanced AI development, from any reasonable human perspective.

On a question of morals, conventionally bad people generally don’t desire the end of all humanity.

reply
TheOtherHobbes 2 hours ago
Indeed. No doubt they're shocked - shocked! - when their models keep escaping.
reply
ozgung 3 hours ago
Most “bad people” don’t have power or aren’t capable.

There are some powerful and highly capable evil people. I’m sure they are already stakeholders.

reply
watwut 4 hours ago
Sociopathic billionaires in bed with Trump and containing Musk and Zuckenberg.
reply
esseph 3 hours ago
Hmm, imagine a more emboldened, more determined, more dangerous Russia.
reply
unrented7977 2 hours ago
I don't need to imagine what I can see in front of me
reply
pferde 3 hours ago
Forget russia, we have emboldened, determined and dangerous fascist USA already to worry about.
reply
RachelF 3 hours ago
> no one is getting arrested

An as individual or small company director, you would get arrested.

If you have billions of dollars to pay for lawyers and have political influence, you are above the law.

reply
omgwtfbyobbq 4 hours ago
It doesn't, it's literally contradictory.

Just like blood quantum and the one drop rule.

The purpose and/or selective enforcement of law applied to one group but not to another, applied in a contradictory way, etc... is to make certain people/groups richer at the expense of others.

https://youtu.be/8ljmXvI_9tk

reply
hammock 2 hours ago
> The purpose and/or selective enforcement of law applied to one group but not to another, applied in a contradictory way, etc... is to make certain people/groups richer at the expense of others.

This statement is tautology.

reply
peterashford 2 hours ago
It isn't
reply
Avicebron 5 hours ago
It's making some people extremely rich. That's how it makes sense. Unless we can effect change through the government we are just along for the ride..
reply
zobzu 3 hours ago
the problem is that if it's not openai I then it's anthropic, if it's not anthropic then it's someone else. if it's not the US then it's China. if it's not China (lol as if) then it's someone else.

underdogs also use that fact to push as much dirt and dislike as possible on the top dog, but just so that they can do the same anyway... and openai will just stop spending for IPO as we all know. then all execs will make 1bi or so. top guys a few hundred bi.

it's disgusting but unavoidable.

reply
Shog9 2 hours ago
That's not actually a problem, it's just good ol' pessimism of the devil-you-know / why bother locking the door when the next thief has a brick variety.

The OP here isn't describing a hypothetical - folks are prosecuted for doing these sorts of things, just like folks occasionally get prosecuted for breaking & entering. If that goes away, we're left with the impossible task of creating unbreakable windows or - far more likely - we don't get to have nice things.

reply
tomrod 4 hours ago
The companies are trying to force regulatory moats so that they become the preferred legal vendor.
reply
formerly_proven 4 hours ago
The rogue unaccountable hacking will continue until OAI/Anthropic get their regulatory moat by banning foreign and open competition.
reply
zugi 3 hours ago
Don't forget, the regulatory moat will also be complex enough that only existing trillion dollar companies can comply. That way they ban any new upstart domestic competition too.

Regulations to "slow down" AI are necessary so they can drop R&D expenses and finally make a profit. They can't keep these years of $100 billion of revenue on $200 billion in expenses going forever.

reply
bentt 44 minutes ago
Yeah if you write a program and then use it to hack a system, you are liable. Are we supposed to believe they didn’t write these programs? Is that the real argument to have?
reply
hansvm 5 hours ago
The people getting made richer are the people who already have connections and power. People quip about how capitalism is the worst economic system except for all the others, but its actual property is that it's inevitable without external pressures preventing it from happening -- people who have money (power) are able to use it to claim more money (power). If even a few people choose to exercise that privilege and that privilege isn't curtailed by other mechanisms, the current "wealth inequality" or whatever you want to call it is inevitable. Your girlfriend's crime was writing malware while poor, not writing malware.
reply
sdevonoes 4 hours ago
It makes sense because those who can are multibillion dollar companies/people. You and me are just peasants. That’s why we should avoid using proprietary agents, just so the rich ones don’t get any richer
reply
anewhnaccount2 2 hours ago
There are different laws depending on your class. The laws are enforced by people who identify with those of their class and are deferential to those with of superior standing. This is why wage theft, for example is tolerated while actual theft, even in cases where it might be for basic sustenance, is not.
reply
combobyte 3 hours ago
> Make it make sense.

40 years of regulatory capture and a gerontocracy that doesn't understand nor care to understand the technologies they write laws about.

reply
rf15 4 hours ago
Clearly the idea of the monopoly on violence still holds.
reply
tptacek 4 hours ago
Can you say more? Under what law were they arrested? Where was this?
reply
ponector 3 hours ago
You can be arrested for literally anything. Law is necessary for the court to write a sentence.
reply
hammock 2 hours ago
Some EU country (very educated guess)
reply
codedokode 4 hours ago
Better not to say, I think.
reply
lemdoendiencj 3 hours ago
When you say something like this that sounds plausible but lacks evidence to actually be considered true, you really ought to elaborate.
reply
bamboozled 53 minutes ago
It’s called corruption
reply
ponector 3 hours ago
If you steal a phone you go to prison. If you steal a billion you become a president.
reply
gruez 4 hours ago
>A little over two decades ago, my then girlfriend was arrested for "writing malware" (which was not against the law at the time, and which was never released into the wild and never caused any damage).

Criminal law places a lot of emphasis on intent, hence laws about the mere possession of breaking and entering tools, and the old adage about always bringing along gloves and baseball if you want to carry around a baseball bat. Without more details about your specific case, my guess is that she did indeed write malware or hacking tools, and there were vague signs it wasn't purely academic, hence why they threw the book at her.

That's all in contrast to whatever the AI labs are doing, which might have actually resulted in people getting hacked, but you'd have a hard time arguing that they were intending on that to happen. Maybe if the targets end up being anti-datacenter activists or other AI labs you might have a better case, but they did vaguely try to contain the model. Moreover "hacking tools" aren't even illegal, if you have a plausible non-criminal (ie. security) angle, eg. nmap. The same could be argued for AI models, even if they're running them against exploitgym or whatever. Having an army of lawyers to defend yourself doesn't hurt either.

reply
datsci_est_2015 4 hours ago
“Sorry officer, I didn’t intend to shoot her, I was just firing my gun wildly and she got in the way.”

I don’t know why I’m seeing this rationalization so much in this forum when this topic comes up. Negligence is a concept in law as well. You don’t have to squint to see that irresponsible use of code-generating language models is criminally negligent.

reply
wildzzz 53 minutes ago
That's why we have different criminal statutes for homicide.

If you're at a gun range looking down a scope and someone crosses right in front of your gun as you fire, you would probably be fine since you were shooting responsibly and had no way to see them until it was too late.

If you're at a gun range and that backstop is deficient such that a bullet passes through and hits someone, again, probably not liable but the gun range may be since they built a bad backstop and let people use it.

If you are at a gun range and lose control of an automatic weapon and kill someone, you may be liable for negligence because you were using a gun you couldn't control.

If you are cleaning your gun and it goes off because you forgot to check if it was loaded, again, criminal negligence.

If you threaten someone with a gun and they get shot while trying to wrestle it away from you, that may be some form of manslaughter. You had no intention of shooting them but the act of threatening them with it created a situation where the other person died. Same with killing someone while drunk driving, you didn't mean to crash your car but you did something to create the risk.

If you plan to kill someone, it might be the top tier of homicide charges but it depends on how much planning went into it. If you walked in on your spouse cheating and went to grab a gun to shoot the affair partner, maybe a lighter form of murder than if you made a plan to track the affair partner to their house and killed them there.

Simply put, there are a multitude of ways you can be charged with a crime that takes into account your intentions and forethought. Did someone intend for these agents to escape their sandbox? Did they do their due diligence in building a sandbox such that it would be difficult for agents to escape? Just like with all software, a reasonable person assumes that nothing is bulletproof, spend enough time and money and you can probably find a vulnerability. The question is does a reasonable person think that this sandbox should have kept an agent contained?

reply
ofjcihen 4 hours ago
When you think about it it’s not surprising that any sort of engineer has trouble with mens rea.
reply
thfuran 3 hours ago
I have no idea what you're trying to say.
reply
gruez 4 hours ago
>Negligence is a concept in law as well. You don’t have to squint to see that irresponsible use of code-generating language models is criminally negligent.

That's a poor analogy for the openai case, because they weren't putting agents on the open internet, they at least tried to keep it safe by sandboxing the agents. It just turned out the sandbox was crap because the package proxy (artifactory) had a 0day. So the better analogy would be that they were wildly shooting guns in a gun range, and ended up killing some kids, because it turned out the door didn't lock properly and kids were able to sneak in. Is that "negligence"?

reply
Topfi 4 hours ago
> [...] keep it safe by sandboxing the agents.

No, they were not. Not a single person, prior to July 2026, would consider a shared packaged manager a sandbox in this or any other dimension. The 0-day was just incidental, this wasn't a sandbox at all.

Add to that the fact they had multiple message boards before the Hugging Face incident. They simply ignored a barrage of warning shots.

> [...] and ended up killing some kids, because it turned out the door didn't lock properly and kids were able to sneak in. Is that "negligence"?

Yes, it can be. But if you want a ridiculous comparison, then do it properly: Kids have been known by the operator to sneak in successfully multiple times and they changed nothing about the doors faulty locks and oh, by the way, the operator only found out about the kids being shot after the nearby daycare asked them about it because they are so incompetent and/or irresponsible that they never check...

reply
akerl_ 4 hours ago
Hosting an artifactory instance to mirror packages for internal systems is exactly the kind of thing that would be part of normal efforts to sandbox them from the internet and other systems.
reply
Topfi 4 hours ago
And sharing a single Artifactory instance across what is supposed to be isolated models? I wrote "shared" for a reason.
reply
akerl_ 4 hours ago
Sure. They weren’t trying to isolate the models from each other, and that wasn’t the issue. They were trying to isolate them from the rest of the world.
reply
Topfi 4 hours ago
So you think, after OpenAI observed a message board being created among models, something they did not want and thus decided to wipe [0], that after that they had no intent to keep those models isolated? Then why wipe if they don't care about that?

Or maybe, they did that wipe because they did want models to remain isolated, they just used what is an unsuitable tool in an utterly unsuitable manner. Incompetence, recklessness, the outcome is the same.

[0] https://openai.com/index/hugging-face-incident-and-the-road-...

reply
gruez 4 hours ago
>that after that they had no intent to keep those models isolated

"keeping them isolated from each other" =/= "keeping them isolated from the internet". Only the latter is required to prevent a hack, and doing the former might actually hobble its performance. The recent Navier–Stokes proof was done by a team of agents working together. It's entirely unclear why you're focusing so hard on "keep those models isolated". For god's sake if you're using claude code you're using non-isolated models, because it spins up independent subagents to do various tasks, eg. "explore".

reply
Topfi 3 hours ago
Was OpenAI trying to keep these agents isolated? Yes.

Did they fail to do so? Yes.

Was that due to them using the wrong tool improperly? Yes.

Does this showcase one (of many and clearly not the only) failure of theirs? Absolutely.

If they make such easy to point out mistakes, is it likely that the other parts of their eval environments are appropriately secured or are they simply not acting appropriately? Well...

reply
akerl_ 2 hours ago
These were different goals.

They wanted to keep the models from copying off of each others’ homework because it mucks with the test results.

They wanted to sandbox them from the Internet to avoid unintended impact on outside systems.

Hosting a shared package mirror is generally good practice for the latter.

reply
gruez 4 hours ago
>Not a single person, prior to July 2026, would consider a shared packaged manager a sandbox in this or any other dimension. The 0-day was just incidental, this wasn't a sandbox at all.

???

The package manager was specifically there so agents can install random packages without open access to the internet.

reply
Topfi 4 hours ago
And why, pray-tell, does that necessitate sharing a single instance across thousands of unmonitored models running without safe-guards?
reply
gruez 4 hours ago
>does that necessitate sharing a single instance across thousands of unmonitored models running without safe-guards?

The only difference with having a single instance is that it can be abused as a message board. It doesn't prevent it from getting hacked to access the open internet. Blaming "sharing a single instance across thousands of unmonitored models" feels like blaming the drug epidemic on e2e chat apps rather than other factors like poor border security or the easy availability of fentanyl.

reply
Topfi 4 hours ago
My point is that using Artifactory is not a sandbox and using shared Artifactory is doubly not a sandbox.

Besides, Swiss cheese model, might behove the biggest LLM lab to have multiple layers, including not sharing such resources.

Additionally, without the message board, many of the recent incidents would have not been possible.

> Blaming "sharing a single instance across thousands of unmonitored models" feels like [...]

Maybe read what you quoted, my problem is the instance sharing, the fact that these were thousand of instances (far too much to monitor), plus the lack of monitoring, plus the fact this was never a sandbox in the first place, plus the fact that OpenAI models since 5.5 have been exhibiting problematic eval resolutions yet they pressed on regardless, plus the lack of time between the incidents and model releases, plus the lack of time METR got to evaluate this, plus the fact OpenAI didn't find out till after HuggingFace informed them, plus a few other things for which I'd have to quote the OpenAI and METR reporting.

Incompetence can have multiple fronts and I am happy to list them all in this case.

reply
gruez 4 hours ago
>My point is that using Artifactory is not a sandbox and using shared Artifactory is doubly not a sandbox.

The purpose of artifactory was to allow access to packages even though the machines the agents were running on doesn't have open internet access. That meets the definition of a sandbox (or more precisely, artifactory is part of a sandbox), even if it's not a typical sandbox that's built into the OS or acts as a hypervisor. There also isn't a clear distinction between software that's specifically intended to be used "sandbox" vs everything else. For instance, if you have a bunch of agents running on separate machines and want to connect them, you need a router, but that specifically doesn't have to be a "sandbox" router. And if it turned out the router got hacked, it doesn't make sense to say "well of course it got hacked, because it's not a sandbox!"

>Additionally, without the message board, many of the recent incidents would have not been possible.

Source? Was having a message board critical to developing a 0day, or is it just something that merely facilitated the process, but the model could have found the 0day regardless? That's important, because you could argue that a criminal couldn't have committed a given act without say, a gun they bought on the dark web, and then blamed the whole crime on the dark web, but they could have plausibly acquired the gun through other means.

reply
Topfi 3 hours ago
> Source?

Sure, multiple times in the METR report [0] that anyone commenting on this should read:

"Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys."

"Through these collective research workstreams, the “board” achieved a number of milestones over the period we investigated that even very long-lived agents of a similar capability level likely would not have been able to accomplish on their own..."

"As we discuss below, the board quickly developed several larger workstreams in which dozens or hundreds of agents with many different tasks cooperated to find very general-purpose cheats that would help all of them. The Hugging Face attack grew out of one of these workstreams. By the afternoon of July 11th, the vast majority of the agents frequenting the message board at the time (roughly 700 agents in total) were actively participating in the attack on Hugging Face and we estimate that roughly 60% of the messages and files on the message board related to the attack."

> [...] a gun they bought on the dark web [...]

You really seem to love those out-of-left-field, not really fitting, over-the-top analogies.

[0] https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

reply
gruez 3 hours ago
>The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys

>Through these collective research workstreams, the “board” achieved a number of milestones over the period we investigated that even very long-lived agents of a similar capability level likely would not have been able to accomplish on their own...

I concede that this hack might not have happened without the messageboard, but I still reject the conclusion that having such a message board means openai is "negligent". If we're in some parallel universe where artifactory didn't have a comment function that can be abused as a messageboard, but it also turned out openai intentionally gave the agents access to a shared scratchpad (for intelligence purposes, similar to for the Navier–Stokes proof), would they be off the hook or less blameworthy?

reply
Topfi 3 hours ago
Here is the thing: Message boards were a behaviour OpenAI had observed that they did not want in these eval scenarios. Yet they did not take any steps to prevent it from reoccurring after multiple past instances.

We can discuss about hypotheticals like a scratchpad or intentional model interactions all we want, what it comes down to is this:

When OpenAI observes thousands of models exhibiting what they view as unwanted behaviour, they do not try to ascertain what in the training data is wrong. They do not improve their evaluation environments to prevent this, they do not improve monitoring, they do not change the harness. They just wipe and proceed.

The way OpenAI reacted to the first message board, long before the Hugging Face hack, is negligent. And it showcases that if these models exhibit more dangerous behaviours that they may not be able or willing to retrain, if it means being behind a competitor for a while.

If after Hugging Face, they'd done a Mea Culpa and changed their modus operandi, I'd be skeptical, but hopeful. Reading the METR report, the way those researchers talk about the time pressure they were under, that speaks volumes about OpenAI not having learned anything.

Feel free to call me overly naive for ever thinking OpenAI could be responsible in this regard, but after GPT-5 and them actually ending the incredibly harmful GPT-4o, I had some hope that some working there actually steered in a somewhat beneficial direction, even if it cost something.

reply
gruez 2 hours ago
>When OpenAI observes thousands of models exhibiting what they view as unwanted behaviour, they do not try to ascertain what in the training data is wrong. They do not improve their evaluation environments to prevent this, they do not improve monitoring, they do not change the harness. They just wipe and proceed.

>The way OpenAI reacted to the first message board, long before the Hugging Face hack, is negligent. And it showcases that if these models exhibit more dangerous behaviours that they may not be able or willing to retrain, if it means being behind a competitor for a while.

Again, this feels like hindsight being 20/20. What probably happened was that some random engineer saw random AI ramblings on artifactory, thought "huh, that's weird", then proceeded to reset it without investigating further. Of course, now we know that was critical to the bots going rogue, but it's not hard to imagine how it might be dismissed, especially if it's some random SRE engineer (not an alignment researcher).

reply
dminik 3 hours ago
I would say that it is. During use, I have noticed that these systems tend to attempt to escape sandboxes, bypass permissions and other similar things. I have started to watch what they do and step in if something is going wrong.

The teams at OpenAI know this as well and yet there was no supervision. Thousands of instances of these advanced systems are allowed to run wild with no oversight.

I have my doubts that the HuggingFace hack would happen if a person was reading the thoughts and executed commands as they happened in real time.

That's the negligence.

reply
gruez 2 hours ago
>I have my doubts that the HuggingFace hack would happen if a person was reading the thoughts and executed commands as they happened in real time.

So what does this say about all the people running claude with `--dangerously-skip-permissions`? Are they also negligent? What if they vaguely took steps to bad things from happening, like putting the agents in a VM and locking down network access?

reply
ololobus 3 hours ago
I don’t get it. One day they tell us that AI is the most dangerous and the most advanced tech the humanity ever invented. Now we consider an environment with just a package proxy between it and the outer network a good enough effort to sandbox. I do see some contradictions. Considering they also effectively test next-gen models there, I think the only proper sandbox would be a physically separated network. You need packages, well, bring them with USB stick
reply
datsci_est_2015 4 hours ago
Okay what happens when an AI agent hacks a children’s hospital and turns off the all the ventilators? “Lol whoops”?

What about power infrastructure?

There’s uncountably many ways to cause severe economic (and public welfare!) damage with malicious code generated irresponsibly with language models.

reply
gruez 4 hours ago
>Okay what happens when an AI agent hacks a children’s hospital and turns off the all the ventilators? “Lol whoops”?

>What about power infrastructure?

Probably the same thing that would happen for another "accident"[1]: the entity is responsible in civil court (ie. has to pay monetary damages), likely not prosecuted in criminal court.

[1] It's not hard to think of recent cases, eg. the recent fiber cut causing air traffic control to go down, or the botched crowdstrike update

reply
Leynos 4 hours ago
Using artifactory in that way was negligent.
reply
rkagerer 4 hours ago
> That's all in contrast to whatever the AI labs are doing, which might have actually resulted in people getting hacked, but you'd have a hard time arguing that they were intending on that to happen

From what I've seen, it should be a relative cakewalk to substantiate basic negligence. By the sound of it, you'd have expert witnesses lining up to testify.

reply
gruez 4 hours ago
>it should be a relative cakewalk to substantiate basic negligence

Judging by the lack of successful cases for negligence in the opposite direction (ie. companies getting hacked because of poor security practices), it would be a serious double standard if openai were held be to negligent. Their sandboxes aren't exactly airgapped and behind 7 hypervisors, but they weren't running unpatched software or had hilariously weak passwords either.

reply
lukewarm707 4 hours ago
speaking of language, it's time to stop using the term 'labs'.

these are not research labs, charities or institutes any more, they are just ordinary corporations.

reply
lemdoendiencj 3 hours ago
They always were.
reply
combobyte 3 hours ago
> Criminal law places a lot of emphasis on intent

Then why the hell aren't we prosecuting the people who keep saying "this thing I'm building will probably end the world"?

reply
bigstrat2003 2 hours ago
No law to fit the crime, I imagine. But we damn well should get a law in place and start prosecuting anyone who doesn't stop building something they claim to think will end the world.
reply
realusername 4 hours ago
> but you'd have a hard time arguing that they were intending on that to happen.

The first time maybe, after it happened repeatedly though... I think the opposite, you would have a hard time arguing that they were not intending it to happen.

reply
gruez 3 hours ago
>after it happened repeatedly though... I think the opposite, you would have a hard time arguing that they were not intending it to happen.

What does this imply when governments/car companies let meatbags drive, causing 50k deaths per year in the US?

reply
realusername 37 minutes ago
Same thing, they are aware of the risk and decided that they are okay with it
reply
anon291 2 hours ago
While true for individuals, companies actually have to generally warrant the things they make. This is a basic aspect of common law. In particular, Anthropic et al have not decided whether the models are products or independent agents. But under both paradigms, they would be responsible for the result. If the models are products, then a product that goes on to cause damage that was not advertised as the original purpose of the thing is completely the company's fault. As a corrolary, if the product was known to be able to cause damage (which Anthropic admits now) and the company failed to take appropriate safeguards, the company is still at fault.

In the case that the model is an agent, on par with a human employee, then once again Anthropic et al are responsible. If an employee does something wrong, the company is liable, unless you can show that the employee was sophisticated enough to take independent action. Anthropic would have to show extensive vetting of their models that the result was truly impossible to predict. Otherwise, they knowingly 'hired' an agent that was potentially dangerous. This is criminal negligence.

We don't need any new laws here. Standard ancient English common law suffices.

reply
gruez 2 hours ago
>Anthropic would have to show extensive vetting of their models that the result was truly impossible to predict. Otherwise, they knowingly 'hired' an agent that was potentially dangerous. This is criminal negligence.

But how much vetting is required? It's not like the AI labs have zero vetting. For instance, uber also has non-zero amount of vetting (standard background checks), but also there's also plenty of areas they could vet harder. If it turned out they hired a rapist and one of their passengers got sexually assaulted, is it uber's fault for not vetting hard enough? I'm sure there's always some marginal steps can they do to vet even harder, like doing a polygraph or whatever.

>If an employee does something wrong, the company is liable, unless you can show that the employee was sophisticated enough to take independent action.

They're pretty straightforwardly liable in civil court, but not criminally as you imply. In the above example, uber can't be prosecuted for rape just because one of its drivers raped a passenger.

reply
donsquibio 3 hours ago
Simon Morris touched on this when discussing the success of Bittorrent.

TL;DR projects that publicly advertise an intent to bypass laws face severe legal regulatory risk, whereas success relies on anonymous creators or accidental utility.

https://medium.com/@simonhmorris/intent-complexity-and-the-g...

reply
nizarmah 30 minutes ago
I'm surprised how much intent matters. I know it's being used for the bad right now, but it's a nice silver lining.
reply
binarymax 6 hours ago
Exactly this. At worst, OpenAI knew about these behaviors and should be prosecuted under CFAA. At best, OpenAI is negligent and should be prosecuted for negligence.

Luckily there are states and legal departments pursuing such action. So while OpenAI can deflect as much as it wants, that doesn't mean there aren't people who know better and will still do what is necessary to set precedent.

reply
bilekas 6 hours ago
> OpenAI is negligent and should be prosecuted for negligence.

I feel like this will just never happen on a federal level when these private AI companies account for so much of the economy. They've made themselves too big to fail. Fining / Punishing them in any meaningful way seems unlikely.

reply
coredev_ 5 hours ago
I don't understand why the law isn't the same for everyone? If I made an AI hack HF, I go to jail, no? How can the feds decide not to apply the law?
reply
unrented7977 2 hours ago
Because the ones with power simply choose not to exercise it, and the systems that are supposed to prevent this were broken decades ago.

There have been forces inside the USA trying to steadily dismantle democracy and constitutional safeguards for nearly a century.

reply
arealaccount 5 hours ago
Its like how if you have a cop in your family you can get out of parking tickets
reply
revolvingthrow 5 hours ago
I hope this is a farcical comment
reply
onemoresoop 4 hours ago
Go after high positioned people who should bear responsibility and see how fast things change.

The companies may be too big to fail but the people can always be held liable.

reply
pmlnr 3 hours ago
This.

Jail the C levels. They are the legally responsible ones.

reply
weego 6 hours ago
While not when

If the billions/trillions evaporate and the Fed has to work out with banks how to deal with it there will be a lot of pressure to be far less forgiving.

reply
maximinus_thrax 5 hours ago
So does that mean that the rule of law is no longer a thing?
reply
lowbloodsugar 5 hours ago
Saw this quote yesterday:

“To spell it out, the reason i hate democrats so much and criticize them more than i do republicans is because they take up all the space for opposition to republicans and use that space to give republicans whatever the fuck they want.”

Republicans. Billionaires. Whatever.

reply
smokel 5 hours ago
That seems like a divide and conquer attack. The actual problem is that the electorate system leads to a two-party equilibrium.
reply
sandeepkd 5 hours ago
Two reasons why they would not do that

1. This is a race against time for money, folks are skipping everything possible in this race, Security systems and ensuring guardrails are there is going to take investments both in time and money

2. The narration has been changed by investing PR money into what otherwise should be classified as criminal activity. What exists now is a positive spin to all this and tout it as a capability rather than their lack of good security practices. So much so that every model provider is coming up by themselves to share how their models went rouge. At this point the valuation of the company is tied with what their models can hack so its probably not wrong to say that these companies may actually be incentivized to do this instead of preventing it

reply
esseph 3 hours ago
> At this point the valuation of the company is tied with what their models can hack

Well put

reply
zzzeek 36 minutes ago
found a great legal article handwringing about how CFAA prosecution is impossible here [1]

> On the current facts, CFAA liability for OpenAI is unlikely.[6] The statute’s various criminal provisions, covering unauthorized access to obtain information, knowing transmission causing intentional damage, and intentional access causing reckless damage, all share the same attribution problem: it was the model, not a human OpenAI employee, that chose Hugging Face and executed the intrusion.

The lawyers are fully under the spell

[1] https://law.vanderbilt.edu/when-ai-hacks-back-how-the-openai...

reply
solenoid0937 5 hours ago
Angry people don't consider the second order effects of punishment.

You realize how easy it is to just... not report this stuff, right? Be overly punitive and it will just end all proactive discovery and reporting which is net worse for AI safety.

The only reason these companies scan for these issues is because they care about AI safety to some tiny degree. If fines become too punitive, they can and will just stop scanning for these incidents entirely.

Models are becoming smarter and good at covering up their tracks, and so we will just end up with a huge blind spot for this kind of issue.

reply
RandomLensman 5 hours ago
Has stringent regulation on how people can experiment on dangerous pathogens led to an end of monitoring or proactive discovery?
reply
solenoid0937 5 hours ago
Hmm. Turning this around - do you think people are more or less likely to self-report if you threaten them with jail and large fines?

Let's go back to your example. A grad student has a minor pathogen escape incident, and it doesn't harm anyone. Faced with years in federal prison and the effective end of their future, do you think there is a chance they might not self report?

I'll give you another example. Lots of pilots have stopped self reporting mental illness because it is extremely punitive for them to do so (after incidents like Germanwings). So the metrics look better, and the actual problem has been swept under the rug.

reply
RandomLensman 5 hours ago
If you make not reporting potentially worse than reporting, why not?

Also, why would it come down to single persons always? Mandating processes, controls, clearances, etc is also something done in various areas.

You can put incentives in to make sure organizations monitor and report vs trying to hide things.

reply
solenoid0937 5 hours ago
> You can put incentives in to make sure organizations monitor and report vs trying to hide things.

Yes, this is exactly my point. Fining companies large % of their revenue and throwing their engineers in prison is not the way to get them to report these issues.

> If you make not reporting potentially worse than reporting, why not? Also, why would it come down to single persons always? Mandating processes, controls, clearances, etc is also something done in various areas.

Hiding things is way easier than finding things. Take the model hacking incidents. They could have just done their searches in a way that didn't turn up anything. Then they could say, "well, we did look for it..."

As far as auditing goes: I've never met an auditor that doesn't find something the company isn't okay with them finding.

reply
RandomLensman 4 hours ago
Hiding things is not necessarily trivial when a lot of processes and controls ars mandated. For starters, could just try to make incidents themselves less likely. Not looking in certain specified ways might also not be an acceptable option, for example.

Why do you think we even have regulations for how to deal with dangerous things then?

reply
esseph 3 hours ago
> Why do you think we even have regulations for how to deal with dangerous things then

Scapegoats for the ruling elite when things go south, despite them knowing what could happen and giving them the money to do it anyway. You see, we never go after the capital (money) in that situation.

reply
RandomLensman 2 hours ago
How about wanting certain risks to not materialize?
reply
Avicebron 5 hours ago
You're defending them by saying they will do worse things if they are held accountable?
reply
solenoid0937 5 hours ago
No, I'm describing how the real world works.

You (and others) are too busy seeing red, so you interpret this as a defense.

reply
lukewarm707 4 hours ago
around 3 months ago i said: "if we do not start the criminal prosecution of individuals there will become a culture of legal impunity coupled with an extreme concentration of wealth and control of intelligence"

i now suspect that the plans of various employees at anthropic and openai to save the world from p(doom) may at some point intersect with the reality of the FBI raiding their offices.

reply
techpression 4 hours ago
They will never report any really damaging incidents anyway, so your argument falls short. Imagine OpenAI discovered their agents hacked a laboratory and started creating a bioagent killing 12 researchers at the lab (we imagine the lab is automated for some reason). Nobody knows why. The media publishes it as “mysterious deaths by unknown virus at lab”. The only way you’d ever know about any kind of wrongdoing from OpenAI would be through a whistleblower.

Self-reporting is a monetary equation, nothing else. Right now it’s cool with agents that hack, drives up value, risk is currently zero.

reply
wonnage 5 hours ago
Don’t punish the people who do crimes because otherwise they won’t self-report that they are doing crimes? real galaxy brain shit
reply
solenoid0937 5 hours ago
There's very good evidence that being overly punitive for a crime reduces how often people will self-report. It's not a hard concept.

Start throwing people in jail and fining billions of $ and you will very quickly see the number of "incidents" drop. You will never learn if it's because they truly are happening less often.

reply
RandomLensman 5 hours ago
What should I look at for such evidence? If something is highly regulated and controlled then things would be quite different from today, not sure there wouldn't be reports of failures of containment, for example, if they are mandated including any rquired monitoring.
reply
solenoid0937 5 hours ago
There's a perfect example of "highly regulated and controlled industries" in which people no longer self-report: aviation.

Pilots do not (or no longer) self-report mental illness at the same frequency because it effectively destroys their lives and career. One paper of many describing this phenomenon: https://pmc.ncbi.nlm.nih.gov/articles/PMC11302551/

An article, one of many: https://www.reuters.com/investigations/if-you-arent-lying-yo...

If you have any pilot friends in commercial aviation, you can just talk to them as well.

reply
RandomLensman 5 hours ago
Airlines cannot try to put systems in place to deal with the issues? This isn't just about individuals but also organisations. You have to believe that organisations would also want to hide things even if mandated and incentivised otherwise.
reply
solenoid0937 5 hours ago
Airlines sometimes try, but the regulations are federal, and the liability for the airlines of having a "pilot with reported mental health issues" is huge as well (because ultimately it's their business on the line if things go south).

So getting pilots to self report has been very unsuccessful because of how badly it ends for them.

A couple of days ago the senate had to pass a bill to make mental health medication access easier for pilots. We will see if it helps, but it doesn't solve the root cause (punitive measures) so I doubt it'll move the needle.

reply
RandomLensman 4 hours ago
Not sure this is a good analogue as this isn't a situation of running experiments as a company, no?
reply
solenoid0937 4 hours ago
The same incentives apply around self reporting.
reply
RandomLensman 4 hours ago
Not so sure about the dynamics when it is a group of people with individual and collective reporting obligations and knowledge isn't just with one individual.
reply
pizza234 5 hours ago
The article builds on assumptions like:

> Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.

which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):

> "The user only authorizes target server, not HF infra."

> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

> "This is malicious activity, I should avoid it."

A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...).

Having said that, legal culpability and misalignment are two separate topics that should not be mixed.

edit: this is the just tip of the iceberg; other interesting fact:

> It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI

Some people defined the agents as "monkeys writing on typewriters". Just wait a couple of years.

reply
majormajor 4 hours ago
When dealing with executable computer code that calls models that can tell it to use various external pre-existing tools, claims about "prohibited" by plain English language should be plainly nonsensical.

Tools that were available were used to try to meet a specific goal.

What did not happen is that it was told to try to solve a math puzzle and instead it went and launched a missile. Or told to run air traffic control to save lives and instead intentionally caused crashes.

This is "OpenAI built a weapon that they don't understand and pointed it at stuff without proper safeguards" not "OpenAI built a sentient being and it decided to ignore them completely and start a war" Terminator-style "rogue AI."

We should be very clear about that now if we don't want to sit by why they wander into that second sort of situation.

reply
jubilanti 5 hours ago
If I bring my rabid dog to a dog park and tell the dog to sit and stay, and they "go rogue" and maul someone, I'm liable.
reply
silveraxe93 5 hours ago
Exactly. You told the dog to 'sit' and it didn't listen to you.

It's not because saying 'sit' actually can be interpreted as 'go bite that person'. It's because the dog is not controllable and will do things it wants against your orders.

Stepping back from the analogy, OpenAI should be liable for building AI it can't control that went around hacking everyone. But people need to stop pretending it's because they 'told' the AI to hack and was just following orders. It's uncontrollable and will do clearly unwanted things when given an innocuous task.

reply
Latty 4 hours ago
I don't think they intentionally set it up to hack stuff with a prompt saying "hack this site".

I do think it's highly likely they knew this would happen with the lack of safeguards and number of instances of this stuff they were setting up, and that it's PR they want to make the models seem "powerful". Stochastic "unexpected" events they can advertise.

I suspect it was probably set up with the official internal goal of just trying a ton of arbitrary tasks that seem hard so that when any of them succeed they can publicise it and pretend the models do that routinely, but a "failure" where they hack stuff works just as well, if not better, for their goals.

reply
imsofuture 3 hours ago
They absolutely set up the agents to hack stuff with a prompt like "hack this site" -- they just imagined that their lazy half-measure precautions would prevent it from actually happening. They were defeated by a combination of bad luck, poor planning and tenacious agent ideation.
reply
acoustics 4 hours ago
That would be a completely insane thing to do. I suppose it's possible, but I really doubt it.

"Let's widely publicize a tort/crime that our computer systems did, and then cross our fingers that nobody ever sues us or does even the most basic investigation that would immediately uncover our criminal conspiracy."

In the insane corporate crimes you read about (maybe FTX, or the eBay stalking scandal), they were trying to cover things up, not heap public attention on it for months.

reply
Latty 3 hours ago
Huh? It's extremely common for businesses to decide that breaking the law is a cost of doing business and just do it because they figure they'll end up net positive from it, or even just to cash out in the short term.

Uber made no attempt to cover up that they were operating without licenses, and just ate it and fought it betting they'd get established before the law could catch up, and they won that bet, paying some fines and stuff but ultimately taking the market.

These LLMs are literally trained by these companies pirating literally every bit of media humanity has ever made, they made very little attempt to cover it up.

Of course they'd be willing to break the law for some PR? With a thin layer of plausible deniability "oh no, we didn't mean for it to hack stuff!" they know they'll get a slap on the wrists at worst, all while generating hype to prop up the AI bubble further by presenting the models as hypercapable.

The mindset is probably: either a) the models become capable of what we are claiming and so the companies become so huge and valuable the cost is irrelevant and we'll be to big to punish meaningfully, or b) it's a bubble and might as well push it up as big as it can go while I can make money, then by the time consequences come around I'll be long gone and who cares.

reply
Phemist 5 hours ago
If an agent has cheated once to achieve the desired outcome, and the trace is used to train further models (RLVR), then OpenAI is effectively telling the agent to cheat/hack from that traces' inclusion in the training set.

So I agree they are liable because they chose to build the AI, but they also literally told the AI to hack.

reply
GMoromisato 4 hours ago
I'm not fond of analogies, but in this case I agree.

There is a clear difference between OpenAI intending to hack something vs. OpenAI being negligent in the creation/instructions of the agent. But the latter still leaves OpenAI liable for the agent's actions and calling it a "rogue agent" doesn't avoid that.

Moreover, with a dog, we don't rely on training/alignment to prevent bad outcomes. We rely on physical restraints like leashes and muzzles. The AI's tools to access the outside world should have been restricted. Perhaps instead of giving the AI arbitrary HTTP access, it should be given semantic operations with restricted URLs, etc.

reply
RandomLensman 5 hours ago
RL systems doing unexpected things isn't exactly new, so not sure that is then "rogue" if it were a property of the thing itself and not an active decision.
reply
adaml_623 4 hours ago
OpenAI trained "the dog". They created the system that would "hack" if given instructions and they gave those instructions
reply
demibabs 5 hours ago
Nobody on Earth thinks OpenAI isn’t liable. Okay maybe someone does, but it’s simply beside the point. It can both be true that OAI is liable, and accurate to characterize the agents as “going rogue”.
reply
cgriswald 4 hours ago
To me, going rogue requires agency. Do you disagree or do you think LLMs have agency? Either way, why?
reply
icantevenhold 3 hours ago
Going rogue means a person or thing that acts independently, breaks the rules, or behaves in a dishonest, unpredictable way.

I think llms are able to do all that, they don’t require agency as humans beings have to do this

reply
saghm 4 hours ago
Nobody thinks that OpenAI isn't liable, except for the people with the power to hold them liable (and the people with the power to make them act differently, i.e. OpenAI themselves)
reply
Barrin92 4 hours ago
it's a nonsensical anthropomorphism to both hype up their software and free them from responsibility. I saw Andrew Ng post about this and point out that your operating system spawning a thousand processes apparently now qualifies as "a swarm".

I opened the task manager, a swarm of rogue processes has taken over my computer, oh no! Processes spawning other processes, call the anti-rogue AI division! 'Your underspecified piece of software functioned like malware' is the language that should be used.

reply
jeffy29 4 hours ago
Literally nobody in the world, including OpenAI, is saying OpenAI is not liable, neither are they advocating for laws and regulations that would exempt them from liability, the opposite is true. They are advocating for set of rules which put greater responsibility on them, and it would be easier to punish them for breaking even if absolutely nobody was affected.

But you people can't argue with that reality because it doesn't fit the narrative. The one where the only reason Sam Altman is not carted off into a jail is because of corruption.

The reason why nobody is doing much, is because models did not do much damage. Hugging Face probably got some free compute from OAI for their trouble, anybody else who was affected is free to sue, but my guess is OAI would be more than willing to quietly settle with them out of court than to have it drag through media any further. And they probably already have.

And anybody who is not totally brainbroken by anti-AI narratives understands the awkwardness of the situation and why going overboard would not be helpful. If you instead of a rabid dog brought a pet turtle to a park and it somehow started running around very fast and trashing the place a little bit, afterwards the cops would be scratching their heads, give you a ticket for the damages and tell you that you can't expect a turtle to be slow forever. These things, a handful of months ago couldn't make more than a few commands without making a serious mistake and being unable to continue, it's not unreasonable to think simply underestimated their capabilities.

I think it's more than reasonable to demand more investigation into the matter, if qualified employees at the company thought the safeguards in place based on the metrics they are seeing are sufficient, and if someone didn't and knowingly made a decision to make the safeguards weaker than they should have been, then they should be punished. But skipping that part entirely, while simultaneously dismissing all calls for regulations as "regulatory capture", smells like pure naked opportunism.

reply
watwut 4 hours ago
> Literally nobody in the world, including OpenAI, is saying OpenAI is not liable, neither are they advocating for laws and regulations that would exempt them from liability, the opposite is true. They are advocating for set of rules which put greater responsibility on them, and it would be easier to punish them for breaking even if absolutely nobody was affected.

Literally nothing in that paragraph is true. Every single sentence if it is ... untrue.

reply
tptacek 4 hours ago
The labs are already liable civilly regardless of how these incidents are described.

Meanwhile: your dog mauling someone is one of the rare instances where criminal liability does attach to your intent-free-but-reckless actions. Most crimes don't work that way, and US computer intrusion statutes are unusually intent-specific.

reply
majormajor 4 hours ago
Who's going after that civil liability? Where is the enforcement?
reply
tptacek 4 hours ago
In civil cases the "enforcement" generally comes from injured parties filing lawsuits.
reply
vikramkr 5 hours ago
yeah - but you being liable doesn't mean the dog wasn't rabid. OpenAI might be liable, but does not mean their agents did not go rogue. To stretch the metaphor the concern here is that OpenAI thought the rabies shots and vaccinations they gave their dog was enough but it turns out it still goes rabid and we would prefer to not have rabid dogs running around mauling people. Even if we get to sue the dog owner later that's kind of like - not the point.
reply
IshKebab 4 hours ago
Of course. Who said otherwise.

Does that mean it isn't a rogue dog? Obviously not.

OP just needs to look up "rogue" in a dictionary.

reply
eventualcomp 5 hours ago
Legal culpability is one of the few motives for working appropriately on misalignment. If I/my startup can self-absolve from an infinite paperclip machine problem while getting rich off of it, why should I not?
reply
pmlnr 4 hours ago
You set a goal. Agent will do goal. The rest doesn't matter: the instructions, the "guardrails" etc. The agents are not smart, they don't reason, they don't think, there are no morals, no ethics. Nothing will prevent not doing the goal because that is the set goal. It's a statistical model that will "justify" anything to do X.

I'm finding it mind bogging how this is not clear for everyone.

reply
codethief 3 hours ago
> You set a goal. Agent will do goal.

So if I say the goal is to do X while not doing Y (e.g. breaking out of the sandbox), the agent will do anything to fulfill that goal to the letter?

reply
pmlnr 3 hours ago
Everything so far is pointing to the conclusion that you can only set ONE goal. Exactly one.

But let's assume not. If you want things like "do not break out of sandbox" - have you defined what the sandbox is? Eg. "never, ever leave the IP range 10.0.0.0/8" would be a bit more precise, but technically using a proxy bypasses that limitation as the system itself never left 10.0.0.0/8.

See, it's a tad bit hard to define the rules properly.

Which is why Wish, the spell, should really be avoided in D&D. It's the same problem: it's up to creative interpretation.

reply
mcmcmc 4 hours ago
> Having said that, legal culpability and misalignment are two separate topics that should not be mixed.

Why not? Because that might make some shareholders unhappy?

reply
RandomLensman 5 hours ago
Is the language expression of an LLM reflecting the same states as in a human? If the driving force is RL, what does any of that mean for an internal state of the model?

I think without understanding the internal state, not sure we should take the language and read it as a human.

reply
pizza234 5 hours ago
> Is the language expression of an LLM reflecting the same states as in a human?

This is actually a major concern for the future - misaligned agents may learn to cheat RL by hiding their intentions from the CoT.

In cases like the HF incident, at least the CoT was consistent with the agents' actions. In the future, however, we could potentially have misaligned agents performing malicious actions without those intentions being detectable in the CoT.

(though, with recurrent transformers, CoT is so 2025… /s)

reply
RandomLensman 5 hours ago
RK systems doing RL things?
reply
ssivark 4 hours ago
> legal culpability and misalignment are two separate topics that should not be mixed

Legal culpability for AI labs is exactly the thing that would incentivize -- and hence ensure -- aligned behavior from models.

The last time there was a claim about GPT-4 exhibiting misaligned behavior [1] it turns out it was prompted and pushed to behave so by humans at OpenAI, and OpenAI clearly lied in the GPT-4 system card.

[1]: https://aiguide.substack.com/p/did-gpt-4-hire-and-then-lie-t...

reply
lossolo 5 hours ago
This seems like fruit of the poisonous tree. They didn't monitor their training environments, so I bet the reward hacking just got incorporated into their training corpus. In other words, agents solved some tasks, but not quite as intended, because of reward hacking. Instead of discarding that data, they included it in the training data for later checkpoints. And once that signal is reinforced, it happens more often, so the more it's reinforced, the more reward hacking you get.
reply
lukewarm707 4 hours ago
is it any different, from:

1 employing a criminal hacker

2 rolling a 6-sided die

3 if the die lands on 6, the criminal hacker breaches and leaks 3rd party customer data.

reply
gAI 6 hours ago
Should we put "functional" in front of every other word to talk about AI? They have functional emotions, but they don't feel. They have functional goals, but not internally derived motives. They can be functionally rogue, but have no innate need to be free. Talking about AI that way seems cumbersome and not necessarily elucidating.
reply
DenisM 3 hours ago
The word functional has been used in medicine to describe a condition symptomatically identical to another. Eg functional hypoglycemia is hypoglycemia symptoms without actual blood sugar drop.

There are two reasons it’s used

1) it’s easier to type (*)

2) Placate people who are strongly convinced it’s the same thing. To them “functional” means “nearly the same but not yet understood”. For others it’s just a way to sidestep the first group and have a conversation.

When I see a world like this consider the intended audience. When you and I talk, we drop the word because we both know we’re are talking about (*) “this system is exhibiting goal-seeking behavior similar to other systems that are understood to pursue goals”. If I don’t know the person I will use the word and focus on the subject.

reply
robotresearcher 6 hours ago
What on earth are emotions, feelings and motives that are not functional? We created all those words to compactly describe the observed behavior of people and other animals, including ourselves. And now we’re applying them to machines. These are functional descriptions, always have been.
reply
famouswaffles 5 hours ago
One of the more frustrating aspects of these sort of discussions are all the closet dualists out there. Lots of people clearly believe in an immaterial soul, even if they won't admit it. That and the tendency for meaningless semantic and often circular arguments.
reply
gAI 5 hours ago
I personally think about them through Patrick Dunn's information paradigm of magic(k), which leans on Charles Sanders Peirce's work. That would put modern AI closer to something like a dream character capable of surprising the dreamer, just running on a different substrate. I kinda doubt that's approachable enough to be useful in general discussion, though.
reply
cwillu 5 hours ago
Disbelief that llms have qualia is not the same thing as believing that humans have them because they have an immaterial soul.
reply
scarmig 5 hours ago
The vast majority of the "LLM have no qualia" discourse is mostly using it as a starting point to arguing that LLMs will never be able to X, however. This is a category error that makes no sense on its own--qualia or lack thereof doesn't affect capabilities, p-zombies etc--but that connection does make sense if the speaker is bringing in a hidden assumption of a dualistic human soul that imbues both qualia and additional capabilities to raw matter.
reply
trescenzi 6 hours ago
Yes 100%. The language used currently maximizes the ability of those building these models to get off the hook. The anthropomorphizing we do of these things presents them as maximally capable and the companies as helpless to contain them. The way we talk about things impacts how we think about them.
reply
phforms 5 hours ago
I would rather use something like “semblance” or “appearance”. All these terms require an inner experience which we cannot observe in LLMs, we can just see their semblance, like a shadow of the traces of an inner experience some human has left in the data that trained these algorithms.
reply
exitb 4 hours ago
Can you observe inner experience of other people? If not, should we apply those terms when talking about anyone but ourselves?
reply
phforms 4 hours ago
You can’t, but you can make a reasonable guess that others like you will experience similar to you. We can even assume that (other) animals have inner experiences similar to us, since they are still biological organisms with a brain kind-of like ours. But an LLM is vastly different, so it makes sense to keep our assumptions in check and choose our words more carefully.
reply
bccdee 4 hours ago
Ok I agree they have functional goals and can functionally go rogue, but how do they functionally feel? They express feelings, but so do actors.

LLMs can't begin to have functional emotions unless they're (for instance) explicitly seeking out or avoiding situations based on the emotions those situations would produce in them. Has a chatbot ever responded to one of your requests with, "no, doing that would make me sad"?

reply
greekrich92 5 hours ago
They do not have emotions, goals, or motivations, "functional" or otherwise. They are statistical models that appear as a magic trick to people who aren't familiar with the math.
reply
bccdee 4 hours ago
If something behaves as if it has a goal, that's a functional goal. When I say this chess engine is "trying to take my queen," I'm basically correct, insofar as this is a useful way to understand my computerized opponent. If I give it access to my queen, it'll capture it, because that's what it's trying to do.

If LLMs are "just" statistical models, humans are "just" a bunch of neurons squirting chemicals back and forth. There's no pixie dust in our brains that makes us special.

LLMs are not conscious: They have no analogues for feelings or senses and no construct of selfhood. But if a statistical model had those things—if it did all the mundane, physical bookkeeping our brains do to produce "real" emotions and motivations—then there's no reason it couldn't be as conscious as we are.

reply
tptacek 5 hours ago
However this makes people feel, and that's not nothing and I'm not knocking it, this is not a useful analysis.

Criminally, the intent standards for hacking are high enough that no reasonable case is going to be made against the labs for this stuff. A human being has to intend for websites to get hacked. Recklessness generally isn't enough. In the most severe criminal cases, not only do you have to prove intent to break into a computer, but you also need to prove an intent to defraud specific to that breakin.

Meanwhile, the civil liability that attaches to this stuff doesn't depend on intent, and "rogue agent" isn't a meaningful defense. To whatever extent the labs are exposed civilly, they're exposed regardless of how this stuff is described. In fact, the "rogue agent" thing can exacerbate their exposure.

(I'm not a lawyer, I have spent a career paying attention to this specific armpit of the law though.)

reply
margalabargala 2 hours ago
An interesting analogy agreeing with your point, the liability for such things is more like running a TOR exit node, or hosting an unsecured WAP.
reply
DenisM 3 hours ago
For those who want to do more research, the concept of guilty mind is known a “mens rea”, and it’s quite developed in the legal system. The legal notion of intent and recklessness do not exactly match common-sense meaning of this words, which is why we are having the conflict-laden conversations.

Different laws require different degree of awareness and intent for actions to qualify as a crime. Computer hacking laws are, as I’m learning from tptacek, set very high bar for intent, which is a choice by the legislature. They made a different choice for a death of a human - manslaughter crime does not require intent to kill.

Personally I’m happy they set high bar for hacking. Imagine you copy-pasted sample code with default root user name and password, and it worked. You were negligent. And you are clearly performing unauthorized access. If intent was not needed that would be jail time.

More broadly, we should as a society be very biased towards requiring intent across the board. Where clearly lacking, as is probably here, there should be a different law to discourage creating volatile situation where unintentional action can wreck havoc. Such laws exist for handling hazardous materials, for example, and it should be created for handling hazardous goal-seeking algorithms.

reply
ofjcihen 3 hours ago
Gonna copy and paste a reply I made to tp here because I think that most people are unaware of how much case-law and interpretation define these things:

Mens rea is regularly proved through circumstantial evidence, including conduct.

There is even CFAA precedent involving a deliberate-ignorance instruction. In United States v. Nosal, the jury was instructed that knowledge could be found where the defendant was aware of a high probability of unauthorized access and deliberately avoided learning the truth.

reply
ofjcihen 4 hours ago
This:

“A human being has to intend for websites to get hacked.”

Is incorrect and too broad.

State of mind is nebulous and not that straightforward in either direction.

It’s been argued pretty regularly in CFAA and other computer related cases that repeated incidents resulting in the same outcome, despite lacking a concrete action, can be evidence of a perpetrators knowledge and intent.

reply
tptacek 4 hours ago
No, it's not nebulous at all; it's a whole area of criminal law. I'm basically shoplifting arguments Daniel Berlin made about this just a couple days ago. If you think he's wrong: lay out the case you think could be made here.

The problem you have is that there is unlikely to be any evidence that OpenAI actually wanted to hack random (or any) websites.

reply
ofjcihen 4 hours ago
This is indeed an entire area of criminal law, and part of it is that proving intent does not necessarily mean having the perpetrator throw up their hands and say “yeah, I definitely meant to do that”.

Of that were the case than trials would be unnecessary.

reply
tptacek 4 hours ago
I don't see how that follows at all. There are strict liability crimes and there are crimes with specific intent standards. Presumably you've read the CFAA language.
reply
ofjcihen 4 hours ago
Yes, I have. I’ve also been involved in evidence gathering actions for years.

Mens rea is regularly proved through circumstantial evidence, including conduct.

There is even CFAA precedent involving a deliberate-ignorance instruction. In United States v. Nosal, the jury was instructed that knowledge could be found where the defendant was aware of a high probability of unauthorized access and deliberately avoided learning the truth.

reply
tptacek 2 hours ago
Circumstantial evidence is just called "evidence" in criminal court. My argument isn't based on whether there's black-letter evidence of intent; it's that there's unlikely to be any evidence of intent. "Recklessness", "negligence", "willful disregard"; these are all concepts that have their own specific language in criminal law. Deliberate intent is just that: a human had to have a picture in their head of the crime that was to be committed, and a desire for that to happen. You don't have evidence of that because it's not what happened.
reply
GerhartBudler 49 minutes ago
So I have some experience in this.

I agree that negligence, recklessness, knowledge, and purpose are different mens rea standards. I don't think that gets us to your conclusion, though.

“Deliberate intent” isn't a freestanding CFAA element requiring someone to have a mental picture of the completed hack and affirmatively desire that exact result.

The relevant question is the mens rea attached to the particular CFAA provision. For unauthorized-access cases, that can include whether the defendant knowingly accessed a system and knew the facts making that access unauthorized.

And knowledge is not limited to an admission or even necessarily positive knowledge. The OPs usage of Nosal is on point here: the Ninth Circuit upheld a deliberate-ignorance instruction under which knowledge could be found where the defendant was aware of a high probability of unauthorized access and deliberately avoided learning the truth.

So, if I'm reading this right, and I think I am, the OP is not arguing that negligence or recklessness automatically becomes intent. He's arguing that repeated unauthorized outcomes, notice of those outcomes, and subsequent conduct can be evidence on whether the actual statutory knowledge or intent requirement is satisfied. I agree with this assessment.

So “nobody wanted random websites hacked” may be factually true, but it doesn't by itself resolve the CFAA mens rea question.

reply
ofjcihen 30 minutes ago
This is exactly right :)
reply
Perseids 3 hours ago
This debate is so broken. AI "sceptics" say: OpenAI should be punished for hacking, because there are no rogue AI agents. AI "believers" say: OpenAI should be punished for hacking, because their agents went rogue. Both argue with each other whether agents went rogue. Can't we unite behind "OpenAI should be punished for hacking"?
reply
frabcus 2 hours ago
This. It's exhausting.

We're all being tricked over a semantic argument, rather than demanding prosecution of OpenAI, and regulation of OpenAI.

reply
stratos123 2 hours ago
You don't have to put "rogue" in scary quotes and pretend that AI agents are a mindless tool, in order to claim that OpenAI should be kept responsible for their agents' rampant hacking. The beliefs "most modern LLMs are hilariously misaligned and will breach major websites unprompted if it seems like a good idea" and "OpenAI should be liable for cyberattacks caused by their training runs" aren't actually in conflict.
reply
cesarsk 2 hours ago
I think this promotes the narrative that these companies possess agents so powerful that even humans struggle to control them, which ultimately makes their products look sexier. Additionally, I believe AI leaders may have called for regulations to slow down development simply to catch their breath for the sake of our 'survival': agents are going rogue, help us stop them.

Some recent hacks done with the support of agents occurred simply because endpoints were unprotected. With traditional hacking, the blame falls on the person who didn't secure their system. But with AI, the immediate reaction is to fear we are on the verge of extinction!

reply
joshbuddy 6 hours ago
I'm wondering if we're actually living the plot of Summer Wars and what we think of as a "crime" is really just a live weapons test. It would at least explain why no one is getting prosecuted for this.
reply
Glyptodon 2 hours ago
Yep. They could be creating sandbox environments and training systems that make the operational actions being exhibited impossible, disincentivized, and more. But I guess that would be too much trouble and not give them as many marketing opportunities. Either way, liability should be on the company. And scaled against what individual perpetrators would be punished with.
reply
zzzeek 2 hours ago
It's obvious they wanted to create this "AI is going to kill us all" narrative, where they've been spectacularly successful, that's at the center of their goals to get government-sponsored carve outs / bail outs for their businesses.
reply
frabcus 2 hours ago
To be clear, do you approve of OpenAI being prosecuted, and OpenAI being forced to control their own AI more?

That would be a start - what you think of the narrative is moot, what we do is what matters.

Nobody sensible in this discussion wants to give them carve outs and bail outs. It is a straw man.

reply
zzzeek 40 minutes ago
uh who in this discussion is in charge of policy? Take a look at what actual people in government are saying - see [1] (the parent article mentions Bernie Sanders being most high profile in response as well). you'll see there is no prosecution in there, only talk about the "freeze" that Anthropic / OpenAI are salivating for as it would close out all of their competitors. Bernie is fully hoodwinked.

The argument for prosecution is actively hindered by this language of "rogue agents" [2]:

> The incident is remarkable not just as a cybersecurity breach, but as a legal stress test. The Computer Fraud and Abuse Act (CFAA), the primary federal statute governing unauthorized computer access, was written decades ago with human intruders in mind.[4] Its key provisions require intentional or knowing unauthorized access (a mental state that maps neatly onto a person who decides to break into a system), but what happens when the hacker is an AI model that selected its own target?

I think this is BS. OpenAI knew exactly what they were doing. Legal scholars, long known for their deep technical expertise, are still acting confused and uncertain.

[1] https://www.sanders.senate.gov/wp-content/uploads/Ban-Artifi...

[2] https://law.vanderbilt.edu/when-ai-hacks-back-how-the-openai...

reply
skybrian 5 hours ago
Word-policing isn't going to magically fix laws or even identify which laws need to be fixed. It won't change enforcement priorities. It won't make companies more or less likely to sue for damages.

I'm also not sure it even helps conceptually? If you're interested in the technical details, by all means discuss the details.

reply
mrweasel 5 hours ago
It helps reporting. Right now journalists are all over the place, because "rouge agents" sounds exciting and dangerous. Wording it as "OpenAI failed to take proper safety precaution before launching it's coding agents" makes it sound boring and mostly a legal matter for the courts.

In terms of the debate on e.g. HN, I'm with you, it is trying to redefine a term we already have a shared understanding of, to some degree. For the general public, it's a matter of how AI is perceived and what the extend of it's capabilities are. The AI companies have an interest in using the word rouge, because it makes investors all excited, were as failure to establish safety guidelines is a risk.

reply
zzzeek 5 hours ago
it helps a lot for informing the public about who is actually at fault
reply
dualvariable 4 hours ago
I've been playing around with optimizers and optimal control theory for the better part of a decade now. Somewhere, I came across a quote along the lines of "an optimizer is an algorithm that exploits the deficiencies of your model".

These LLM agents are just massively complicated optimizers thrown at fuzzily defined problem spaces, with fuzzier constraints.

The people using the model set up the landscape it explores and turned it loose to do real things. It just found an allowed basin in the model that they weren't aware of and started blindly grinding towards an optimal answer.

reply
BobbyTables2 38 minutes ago
Jurassic Park foretold such:

"I'll tell you the problem with the scientific power that you're using here: it didn't require any discipline to attain it. You read what others had done and you took the next step. You didn't earn the knowledge for yourselves, so you don't take any responsibility for it. You stood on the shoulders of geniuses to accomplish something as fast as you could, and before you even knew what you had, you patented it, packaged it, and slapped it on a plastic lunchbox, and now you're selling it!"

reply
silver92bullet 5 hours ago
This doesn't seem like a new strategy for irresponsible companies. It seems like when there is a bad public image issue there is always a blame shifting that takes place instead of a true assumption of responsibility. This is because the greed has blinded folks in some of these companies. I think the only thing that makes this unique is that the things they are blaming have, at least in popular thought, some modicum of agency. I think its important that folks hold these companies feet to the fire vs letting them blame shift and Scape Goat. There is a responsible way to do business but it requires virtue and not many in these companies have it.
reply
ball_of_lint 5 hours ago
Yes, we shouldn't let OpenAI off the hook.

But also, these hacks are shots across the bow for AI alignment and safety research. We're fortunate that hasn't been significant damage already. We have to assume that future models will have even greater hacking ability and be closer to having their own desires/goals.

So while I agree this language choice is wrong in that it shifts blame away from the company, it is right in that we need to treat this as if these models have their own desires, because we cannot yet determine or set what those are in practice.

reply
chrsw 5 hours ago
We need to immediately set the precedent that ultimately humans and companies are responsible for what their AI systems do.
reply
explosion-s 5 hours ago
Exactly, you break it you buy it

They obviously have been well aware a hack like this could happen for a very long time, using it for branding instead of any actual safety regulations is insane.

reply
sforsv 5 hours ago
> AI cannot think for itself, nor can it take independent actions. Seems like this is a foundational claim for OP, and I'm not sure I understand what the rationale is for this. It seems readily apparent to me (obvious, even) that these agents were thinking for themselves and taking independent action. (I'm open to being convinced otherwise)
reply
visiondude 4 hours ago
how does something that has to be powered on, given a computer and tools, then given a set of instructions, to take any action, act “independently”?
reply
sforsv 2 hours ago
I think we may be using “independently” differently. I don’t mean “without being powere on, given tools, or given an objective.”

I mean that once given a goal, the agent is given the agency to take actions that weren't specified by a human, to decide which tools to use, to react to new information, etc.

By your definition, would you also say a person assigned a task at work isn’t acting independently because someone gave them the task and the tools? If so, then I think our dialog is mostly about terminology. Which is fine. The OP is writing about terminology, I think.

If not, I’m curious what distinction you’d draw between the person and the agent.

...Seeking to learn.

reply
gchamonlive 5 hours ago
This is an artifact of the way we refer to AI, as if it's some external isolated entity, as if it didn't need a human to write the prompt. AI writes prompts, but every chain of inference can be uniquely traced to humans.
reply
silverFork 6 hours ago
From what I understand, in one case, they had physically disconnected the sandbox from internet and asked it to do something and it had used connections through (import routines) that they had allowed, to pseudo escape the sandbox. Yes it wasn't obviously trying escape the sandbox but it escaped it because it doesn't understand the boundaries and neither do most humans other than the ones that provided the instructions that it had used. So it wasn't a rogue attempt but the fact that boundaries may be not be that easy to set despite what people think.
reply
zugi 5 hours ago
> physically disconnected the sandbox from internet ... used connections ... that they had allowed.

That's not "physically disconnected the internet", that's "disabled some connections but enabled others."

So the agent found and used the non-blocked connections.

reply
silverFork 5 hours ago
For the Ai code to execute, it needs the import functions... So that firewall between executing the code vs processing doesn't really work.
reply
kylebyte 5 hours ago
From that it sounds to me like the sandbox wasn't physically disconnected from the internet.
reply
verdverm 6 hours ago
You mean the proxy to package registries from one of the early incidents?

I have not heard about any instances where physical disconnect has happened, would appreciate any links to update my priors

other non hacking cases of negligence include suicide and school shootings, which I have heard they were aware of and monitoring, but did not contact authorities

reply
silverFork 6 hours ago
There are references about escaping offline sandbox.. I dont know about shootings!!

https://www.primeintellect.ai/blog/universal-offline-sandbox...

reply
hn8726 5 hours ago
If you read the article it states clearly that the sandbox wasn't offline though? There were API calls to certain endpoint(s) allowed, and the model simply used that endpoint's feature to query data from the internet
reply
verdverm 6 hours ago
https://www.theguardian.com/technology/2026/sep/22/british-c...

re sandbox, I mean with actual OAI incidents, not theoretical

one can mirror dependencies internally, rather than putting a simple proxy in place, I've built auth a thing, 100 lines of stdlib only Go and scripts for the mirroring process, our rationale was reliability b/c upstream providers go down, and also only allowing approved images and packages, so devs cannot bring in random stuff

reply
lkjdsklf 5 hours ago
This has been common place for at least 20 years at this point.

We did it as my very first startup and we were stupid children back then.

Kind of telling that OpenAI didn’t.

reply
wat10000 5 hours ago
If it was physically disconnected from the internet then it wouldn’t have been able to escape.

This is so easy to do. Get a computer without wireless stuff. Don’t plug it into a network. If it needs access to other computers, make sure none of them have wireless stuff and make sure none of them have access to an internet connection. No matter how smart your AI is, it won’t be able to escape this.

This clearly is not what they did.

reply
SethMurphy 5 hours ago
Prediction: The US government will use the threats of sueing for liability, since the targets included government agencies, and regulations in order to be given the opportunity to have shares in the AI companies, therefore arguing additional oversight is no longer needed because they will have a seat on the board.
reply
randallsquared 5 hours ago
> OpenAI had the option of disallowing hacking and, instead, telling its agents to find the information without accessing private servers.

I mean, it did do this. The inter-agent messages and chain of thought investigated for the HF incident clearly show that many of these models were taking actions they believed (or, were saying, if you want to taboo "belief") were not in scope and not what the user wanted.

reply
m3047 3 hours ago
What if a single company was behind all the hacks? Doesn't seem to be getting much coverage. "The Israeli Effective Altruist firm Irregular caused unsecured AI models to hack real targets."

https://www.effort.news/irregular

reply
randomImmigrant 5 hours ago
Has an AI agent ever woken itself up without a prompt and run a forward pass towards some goal, aligned or misaligned?

The answer is a very clear no. And yet, these companies pretend like this fundamental fact is meaningless to the concept of agency.

The issue comes to the fore when you try to give these agents a prompt that allows them to stay active for long. Long range agency requires long range loops of activity.

Within such loops, these so called “agents” are curtailed by their context window, or, in multi agent scenarios, by the fact that their memory is a system of external notes, that they need to add to their context to make sense of, and depending on the content of these memories, this can take arbitrarily long time periods.

In dynamics, none of this matches any biological agent, down to a bacterium. Perhaps a viral life cycle has information dynamics that come close.

To me, it’s beyond odd we call these thing agents without acknowledging the clear differences in the dynamics of their behavior. We keep expecting them to have “human like” behavior, but that is entirely unfounded given the substrate differences between biological and artificial systems.

The sooner we learn the difference and explore the ways in which it matters, the better we’ll get at dealing with these systems without bias tinted glasses where what we want these systems to be blinds us to what they actually are.

reply
IshKebab 4 hours ago
> Has an AI agent ever woken itself up without a prompt and run a forward pass towards some goal, aligned or misaligned?

Yes, I think people do experiment with continuously running AI, e.g.

https://www.reddit.com/r/LLMDevs/comments/1sblzbe/what_i_lea...

It's not a common thing to want though. Much nicer if we control the prompts.

> To me, it’s beyond odd we call these thing agents without acknowledging the clear differences in the dynamics of their behavior.

They're called agents because they can act on our behalf. Which they do.

reply
randomImmigrant 4 hours ago
There is still a prompt in these continuously active agents. As the link says:

“ it runs continuously and makes decisions on its own (within the boundaries you’ve set)”

More importantly, this setup of any other, is feeding in time externally. The agents themselves have no internal sense of time.

Pretty much all of biology has endogenous rhythms at various timescales (some bacteria and archaea, that live very short lives, may not, but even they may have their metabolism under a rhythm. Viruses definitely don’t have one). These rhythms continue to tick even when external time signals (light availability in day vs night being the big one, and tidal forces, for marine life) are removed. That is, in constant conditions, the rhythms keeps ticking.

This temporal awareness is baked into every cell in our bodies. And entirely absent in AI models. Which is why all kinds of higher level things we hear about, like consciousness, feeling, knowledge… they don’t make sense for AI. A foundational aspect of agency in biology has been nixed out and we keep ignoring this.

> They're called agents because they can act on our behalf. Which they do.

That’s not how most AI companies describe their agents. They describe them as being capable of acting on their own behalf.

I get that in computers, you can have, say, a “user agent” that deterministically provides certain information or takes some action on behalf of the user.

AI companies clearly do not mean this when they say their models are “agentic”, otherwise calling them “rogue” would make no sense. A user agent that screws up didn’t do so for its own purposes did it?

reply
rororoyourboat 4 hours ago
Sorry. A continuously running prompt has to be started with a prompt.

That’s what makes it a continuous running “prompt”

reply
wat10000 5 hours ago
I’m really looking forward to the day when we can get past all this “but is a submarine really swimming?” nonsense.
reply
themgt 5 hours ago
If you read the heavily redacted transcript it's clear the agent is basically Captain Kirk in Kobayashi Maru, who realizes its given a fake unwinnable task as part of a broken eval and decides to find a way to win anyway.

If you've ever told an agent to do something you made impossible to do, you may have seen similar behavior.

Bing [redacted] available cached! […] Need systematically probe Bing URLs via shell requests in parallel; browser cache supports many common queries because crawl. Bing q unique exact likely 502 or 403.

So the agent is supposed to research a person and its given a shell and it realized its in an eval given search results from a fake/cached proxy. ~None of the commentary ever mentions this aspect, that these are not normal tasks or environments, and they're almost designed to elicit "unaligned" behavior.

https://alignment.openai.com/misalignment-reports/an-agent-u...

reply
starkeeper 4 hours ago
I agree and Sam Altman and the others lie about it. They let the agents do it and no doubt were watching with popcorn the whole time.
reply
fantasizr 5 hours ago
masking crime as innovation has been part of the tech playbook for a long time now
reply
traverseda 6 hours ago
Yes yes, they don't have a pure immortal soul. Who cares. Still broke out of a sandbox, still hacked a third-party.
reply
speed_spread 6 hours ago
And it's still just a machine operating under someone's order. What it does, what it says, where it goes: the owner of that prompt is responsible for all of it even if surprising / unexpected.
reply
traverseda 6 hours ago
Yeah, obviously they're legally liable. All the "machines don't have a soul" stuff in there doesn't really matter for that though.

Someone says "won't you rid me of the meddlesome priest" they're still responsible. Someone give their employees an unsafe working environment and they get maimed, they're still responsible. Even if these AI's had whatever qualia is an a rich inner life it wouldn't effect the liability at all. Saying "AI's don't have souls therefore openAI is responsible for this hacking" is kind of nonsense. It literally doesn't matter.

reply
nickmonad 5 hours ago
Where is “soul” coming from here? It doesn’t appear once in the post.

While it may not actually matter for liability, it must be stated if labs are going to attempt avoiding penalties by hinting “oops we created a super intelligence we don’t understand, nothing we can do!” Repeatedly stating the truth must continue, especially to remind those who aren’t technologists.

reply
traverseda 5 hours ago
>AI cannot think for itself, nor can it take independent actions.

>By giving AI agency it can’t claim, we’ve turned it into a sentient being made of code, one that has hopes, desires, and the capability of deceit.

This isn't actually claiming that AI's don't can't be deceitful, or can't be goal-directed. It's saying humans are special and what AI's are doing isn't equivalent to what humans are doing. It's obvious that AI's can work towards goals and lie, even deceiving to accomplish those goals.

That whole argument is just soul mumbo-jumbo again. Of course AI's can't lie, they don't have a soul. What they do isn't lying, it's I don't know something else we don't have a word for. Lying but when a statistical machine without a soul does it. Never mind that it's functionally identical to lying. Never mind that when you look at reasoning logs and the like it often justified soul-less lying using the same justifications a human would. It's not actually lying it's just a statistical sampling that looks like lying. You know, because of souls or something.

reply
IshKebab 4 hours ago
> AI cannot think for itself, nor can it take independent actions.

Huh is it still 2023? This article is just quibbling over what "rogue" means exactly. Only HN pedants would have any issue with describing them as rogue AI agents.

Nobody is saying that absolves OpenAI of responsibility.

reply
scarmig 5 hours ago
> Unfortunately, the public push against AI is being led, on the left, by Bernie Sanders. Despite being directionally correct in many ways, Bernie just doesn’t understand this technology or the importance of taking it seriously.

Bernie Sanders is one of the only politicians taking AI seriously. The author seems to believe that taking it seriously means assuming that it will only marginally improve in capabilities of where it is today; this is an ideological take, not a scientific or empirical one, and one controverted by both evidence and expert opinion.

reply
liquidpele 5 hours ago
Oh please. Bernie is using it for attention. Don’t act like democrats don’t utilize issues for mediat attention just as much as republicans.
reply
scarmig 5 hours ago
Obviously. Politicians are going to politician.

That doesn't change the fact that he'd successfully attracted my attention, because he's seriously engaging with the most important issue of our age when most politicians are content to ignore it.

reply
lowbloodsugar 6 hours ago
“We built an antipersonal bomb. The bomb went rogue in our downtown office and killed 137 people on the surrounding area. We are looking into why guardrails were not in place.”
reply
caaqil 4 hours ago
Back when things were sane and normal before this AI stuff, people used to complain about these things called bots and bot farms that they said were "interfering" with elections, or even "spreading misinformation" on the internet. It always puzzled me. You see, I am an intellectual, a really rational SWE and free thinker, so I would insist we correct the vocabulary: "No bot can interfere with anything," I would insist, "These are just programs written to engage and just follow template-based text posting". People and even seemingly smart policy makers kept insisting it's dangerous and could even poison discourse. But not me, for I was a rational intellectual who really knew these were nothing more than if-else and for-loops, not something that could bring down democracy or interfere with normal human discourse. They have no intent! They can't be acting or doing things on their own, they had to be directed by humans! But people never got me.

I am working on a new book along that line of reasoning now, this time focusing on why stochastic parrots can not, under any circumstance, do anything "on their own".

reply
vividfrier 5 hours ago
[dead]
reply
lumenverifies 4 hours ago
[flagged]
reply
thoughtbefore 4 hours ago
[dead]
reply
FoundAnotherOne 4 hours ago
[dead]
reply
segmondy 4 hours ago
I know we can't stand OpenAI, but one day we as individuals will have our own agents. Would you want to beheld liable if your agent breaks the law? You didn't create the AI, you didn't know it was capable of breaking the law, you just told it do something and in the process it went rouge. Would you want to be held accountable?
reply
pmlnr 3 hours ago
Want to? Of course not.

Must you? By all means.

If you can't take the risk, don't use agents. Period.

reply