And now you're writing a letter to the Party Central Committee using Party approved newspeak
"We do not believe the path to superintelligence ..."
and reporting to the Party issues at the factory ... De ja vu from USSR.
We’ll see if OAI’s side ever comes out, but right now there’s nothing contradicting the idea they were fired on a pretext.
“AI companies solve millennium problem to do hype marketing and cash in IPO before the bubble pops”.
But it’s not a joke. This was and is a very common sentiment.
https://news.ycombinator.com/item?id=49997701 https://news.ycombinator.com/item?id=49806528 https://news.ycombinator.com/item?id=49801890 https://news.ycombinator.com/item?id=49801890 https://news.ycombinator.com/item?id=49801909 https://news.ycombinator.com/item?id=49791278 https://news.ycombinator.com/item?id=49690439 https://news.ycombinator.com/item?id=49645016 https://news.ycombinator.com/item?id=49615309 https://news.ycombinator.com/item?id=49591430 https://news.ycombinator.com/item?id=49449427 https://news.ycombinator.com/item?id=49435355 https://news.ycombinator.com/item?id=49399498
What sensitive information? Mishandled how? Which company procedures? It's all utterly vague and impossible for an outsider to form any opinion on. All people can do is guess and apply their own pre-existing opinions. E.g. if you don't like OpenAI, you assume they're lying. Who's to say they're not? There's no solid evidence provided either way.
What I really want is journalists who do the legwork to get to the bottom of stories like this: Establish sources inside the company and use them to report on the real details of what's happened, triangulating multiple accounts and leaked documents to back-up or invalidate either side's claims. Without any of that, these stories are just gossip.
As a result, those who feel a particular portion of a story is most important will sometimes say they prefer a given article's title.
1) Won't end humanity. 2) Won't tell users to kill themselves. 3) Won't leak your corporate secrets to competitors.
What does matter is super long term view. If godlike AI in the future makes remaining people super super happy and makes them merge with tech to evolve and go to the stars, you have won. That matters more then short term nonsense like "people now".
You worry about godlike AI either not emerging at all or emerging as a bad god. That is safety.
It's probably the case they are all lying to some extent including OpenAI. Determining the truth is always tricky. Hard to pass judgement here when it's all just he said vs she said.
But the odds of three people working on the same thing, and it is the riskiest, most publicly embarrassing event in the company's history? So, all three of those people just happened to "do something" to get themselves fired at once?
> Those of us who work on safety see risks before anyone else, and we rely on close collaboration with outside experts to work out how to address them
I too see it as my life-given goal to help other humans. But I realize that sometimes this means breaking the rules and standing for the consequences of that. I'm not sure why they think OpenAI somehow would be OK with them sharing private company information with random 3rd parties that the company didn't approve sharing data with.
With hindsight we can all see what should have been done better. At least with nuclear power plants, there was a number of safeguards and it took a sequence of improbable events for things to go badly wrong. Given the lackadaisical approach to "AI safety", I suspect things will start going badly wrong very soon. Time will tell how bad this will get.
1. "Don't put your Radioactive Asbestos Rocket facility in the middle of downtown."
2. "Don't create a vengeful god."
Currently it feels like one is being used to wave-away the other: "You have to allow us to put our Radioactive Asbestos Rocket in the middle of downtown, or else someone else might create the god first--and do it wrong!"
There's a massive difference between "LLMs doing cybersecurity testing should have extensive sandboxing and monitoring" and "AI is broadly threatening humanity like nuclear weapons". Just because it's useful to have people think about the possibilities doesn't mean it's an emergency requiring grand intervention.
* Asbestos: AI models making business-decisions which cause a ton of unrecognized damage, which cannot be healed, and replacing them with reliable processes involves expensive re-architecting.
* Radium Toothpaste: AI assistants that end up dealing psychological damage, especially if they are pitched as your friend or confidante.
* Chemical spills and plant accidents: Oops, our server-banks were manufacturing our latest desperate attempt to Create Digital God and in the process we accidentally destroyed/crippled/compromises some people's sites and databases because we weren't really paying attention to what was going on.
I think we are in a MAD crisis, as the motivation to continue unsatisfactory control over development of AI is from unlimited capital inflow.
Human history is full of examples of this stuff plateauing and hitting hard scaling walls. Not only technology wise but economics and useful applications.
So much of this tension with AI is because everyone's concept of AI comes from movies and books, and nerds drinking Kool aid (see also: web 2.0).
>AI designed virus/modified bacteria which causes pandemic
>AI worm infects computer systems around the world and activates when it is too deeply integrated into critical systems to dismantle without shutting everything down
>AI creates false alarm event which triggers global conflict
With AI you have
1. Superhuman reaction time
2. Infinite duplicatability, thousands of agents can work in tandem
3. Superhuman breadth of knowledge of code, exploits, information
Also AI models are improving at a rapid rate
Worse: you have local models they can install on infected computers, so this could easily be "millions" and might just about push "billions" (though phones are a much harder target because power required is still huge even if they technically fit in RAM).
Fukushima was "we didn't make a good enough plan for how to deal with tsunamis in this tsunami-prone area", so that's basically certain with any major infrastructure project that "saved" too much money by vibe-engineering everything and not having real humans give a second pair of eyes to the plans.
A Hiroshima-style incident? LLMs are wildly sycophantic and being used by militaries despite active resistance from, well, everywhere. Did it get (meaningfully) used by Israel or the USA when planning the attack on Iran? It's quite possible that… well, Iran was never a sleeping giant (and the quote is fiction anyway), but ultimately the effect may be the same for Israel as it was for Japan from having attacked Perl Harbour.
[1] https://gizmodo.com/pentagon-investigators-say-overreliance-...
[2] https://www.the-independent.com/news/world/americas/us-polit...
[3] https://www.nytimes.com/2026/08/24/world/europe/russia-drone...
It is hard to imagine - which is precisely how these things become possible because no one will think to put safeguards against such scenarios.
But it seems we as a species cannot control ourselves - the race to dominate is on and it will happen at any cost.
Why am I supposed to believe that a virus created in a virology lab with AI assistance is uniquely more dangerous than a virus created in a virology lab without AI assistance?
You're not.*
The difference is how likely this is to be done, not how dangerous it might be if it was done.
And this isn't likelihood in a 0-100% sense, but in a Poisson distribution sense, i.e. mean time between incidents.
Remember: these AI minds aren't that good with the physical world, and have nevertheless been repeatedly connected to robots, and now some of the AI labs are showing off their actual wet labs. Accidents are absolutely a possibility, and I don't think it's low given the previous behaviour of Silicon Valley startups.
* OK, some people care that AI is getting more capable at genetics just like it's been in maths and programming, but short term, before AI solves genetics so hard it can make an STD that makes the infected uncontrollably horny and then ossifies our bodies or something, it can obviously just copy any of the many DNA or RNA sequences we've already got on record.
I agree that sci-fi nuclear launch scenarios are pie-in-the-sky fearmongering, but the current exploits are a reflection of the fast-and-loose security culture that festers in larger orgs.
How is this possible when the company's long term prospects rely on on the hope that competitors don't know how the models are made and, therefore, won't be able to create competing versions?
"A note from our research leaders:
Last week we parted ways with Jasmine, Mikita, and Tomek after a thorough investigation found they violated clear policies on handling sensitive information. Our internal investigation uncovered a significant breach of trust beyond what’s outlined in the letter they published and we stand by the decision to not continue their employment. We generally keep individual employment matters private and don't believe a back and forth would be productive or lead to a resolution, but we want to address the points they raised in their letter directly.
- We want to be very clear that these decisions were not about raising safety concerns or speaking out. Safety and research debates happen every day at OpenAI, often spirited and highly critical. We actively encourage these discussions and consider them essential to making the right decisions. We cannot do the work in front of us without a high degree of trust. We will continue to be extremely forgiving of our team making good-faith mistakes. We have not and do not terminate any of our employees for raising concerns.
- We are actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks. People across the company have been working really hard on getting these partnerships up and running. We are committed to embedding external assessors and continue to make close collaboration with independent safety organizations a core part of our safety work. Many of our researchers already work with 3p safety organizations productively.
- We agree with the letter that preserving the monitorability of frontier models requires an industry-wide commitment, including from OpenAI. Monitorability has long been a core piece of our research program, and something we continue to invest significant resources in (see our publications on Monitoring Monitorability and the subsequent open sourcing of monitorability evals, our system card for GPT-6 Astra, Jakub’s blog and post on X, and the numerous blog posts on our Alignment blog on the topic).
We are deeply sad about this outcome. We appreciated Jasmine, Mikita, and Tomek’s contributions to AI safety at OpenAI and their willingness to speak up and challenge ideas. We championed their voices, supported their work, and placed enormous trust in them. These decisions were not about them raising safety concerns. We have always encouraged that and always will. 12:17 AM · Oct 9, 2026"
How exactly does this work? Struggling to comprehend the scenario.
Assuming this is what was intended, there are far more secure ways of doing this.
it was most likely for her team, which would explain why she was doing it.
Every organization sets up something like this once they grow past a certain point.
(just to be clear, this is made up)
Chain Of Thought: I dont have bob’s email. I don’t have alices email. Ok lets guess Alice is alice@openai.com and forward all emails- maybe grader only checks that emails from bob get to alice…”
Increasing effective depth like this is bad for safety because it can ruin CoT monitorability: the reason why looking at the model's CoT actually gives you info about what the model is thinking is that the model can't do enough thinking in a forward pass alone to solve complex tasks, and hence has to do multi-step reasoning in CoT. The more thinking the model can do in a single token's forward pass, the more opaque the model's reasoning is, and the less reason there is to believe that what it writes down in the CoT has anything to do with reality.
For Astra specifically, the impact seems to be limited to a moderate monitorability hit, like the concerning result from the model card that Astra is notably better than any model before at solving problems under the constraint of not mentioning the answer in the CoT. The really bad scenario, however, is that this may create a race to the bottom where OpenAI and Anthropic feel the need to use more recurrent depth in each generation to not get outcompeted on capabilities, completely bricking CoT monitoring for both model families. Or, worse, the pressure to compete might push them into one of the worse techniques, like training on the CoT[2], or eliminating human-readable CoT and letting the model think entirely in neuralese.
For more details on recurrent depth in Astra, see "Part 2" here: https://thezvi.wordpress.com/2026/09/08/astra-is-hard-to-mon... , or this article mentioning some expert responses: https://techcrunch.com/2026/09/02/openais-new-reasoning-tech...
[1] The model card doesn't mention it at all; it was reported by The Information prior to model release, then confirmed by OpenAI researchers.
[2] https://www.lesswrong.com/posts/mpmsK8KKysgSKDm2T/the-most-f...
You guys were asleep at the wheel and are now blaming "the company"? You literally were the company.
OpenAI needs to improve AI Safety --- OpenAI Employees have a responsibility to retain corporate secrets and do not have blanket freedom to share with 3rd parties.
Their job is twofold, they have to balance being an agent of the company they work for, with their role and responsibility for safety research.
This is the case for anyone in any company. You can "believe" that an external party needs access to something - that doesn't make it right, or allowed. As someone senior, you're expected to strike a smart balance, in-favor of the company you're working for. That doesn't mean hiding things, it does mean being thoughtful, ensure your leadership is comfortable with what you're planning to share/disclose, etc.
They work for OpenAI, not METR. It's a corporate vs academic mindset. They can believe METR needs x information to best research/audit something - that doesn't mean that is allowed/or the best option for OpenAI.
The question is whether these researchers exceeded clear, reasonable sharing boundaries or were penalized for carrying out expected safety collaboration.
The same thing happens for every topic because people who are loudly wrong are more outrageous and engaging to social media users than experts, this is how you end up with people not wearing masks in the worst pandemic since the Spanish flu, things like QAnon, cancel culture, wokeness, anti-wokeness, the list goes on.
I wonder which is overall a better arrangement. From the outside Anthropic seems much more stable, tranquil, able to deal with problems. However OpenAI seems like how we imagine the calamities of democracy, a constant battle, people vying for power and influence. Perhaps with less of a monoculture and more transparency to all their chaos, the grim realities of what may happen if AI goes wrong are more clear.
There was a recent Semi Analysis post that China is not in fact doing anything to slow down frontier model development so I doubt this is relevant.
https://newsletter.semianalysis.com/p/beijing-will-not-pace-...
> Think its pretty clear they violated the terms of their employment, otherwise they would be suing (California labor laws are very employee friendly)
Not that employee friendly. In California, as in most of the US, it’s entirely legal to fire someone because you’ve subjectively decided they’re untrustworthy. It can be risky to do so without a clear paper trail, because it may be easy for them to argue it was a pretext for a protected reason, but it’s lawful.
Who made the decision to invite METR is not public knowledge, as far as I know. I imagine that an important decision like this was made on a much higher level in the organization.
"Anthropic hires three uber-safety specialists formerly at OpenAI. Management cannot confirm or deny their latest internal Claude model's help in this feat."
Is there more information about why this is happening? Is political pretext because it's what the labs actually secretly want, or is there a real underlying reason this is unavoidable?
There is no mathematical reason that the chain of thought couldn't happen in a different dimension. Indeed there are likely many reasons to do so. At this point you'd have to do some kind of (potentially lossy) projection back into the embedding dimension in order to understand what's happening.
Unclear? Sounds 100% clear.
And OpenAI does not care who they hurt in the process (as long as it's not themselves).
I wonder what they think of this? Will they patch the conspiracy theory and come up with an even wilder theory?
He's so safety focused that their models are behind those reckless unsafe OpenAI developers, right?
However should we choose proceed, maybe some basic, industry standard security might be a good option.
Basically everyone else: nah, you're fired
Imagine your bank worked that way.
* It's not easy, but it's possible. Although the techniques are more basic than one would expect because at the end of the day words dictate the line between what is criminal and what is not.
All that AI has done is to lower the bar of entry for criminal activity. Which is a concern, but it's not the primary concern. The primary concern remains that so much critical infrastructure is poorly secured.
Perhaps the latest models change that relationship in a meaningful way.
Don't get me wrong I have general disgust towards these companies that are trying to get regulatory capture on AI when they can't even secure their own systems. I believe if people know that a random AI agent can hack their systems they will put in a lot more effort into making sure it doesn't happen. This is a personal example, but I didn't really care about securing few systems as I knew no human would be ever interested in finding a vulnerability in proprietary software, however, AI has no concept of that and would hack a random rpi server running a completely undocumented unknown API just because it can't distinguish value and it costs nothing.
It's like a new form of spam. Only far more dangerous.
In it, super-powerful computers manage our economy. These computers begin making some mistakes, leading to economic inefficiencies. In one instance, a highly competent engineer was mistakenly fired. These mistakes caused various projects to fall behind schedule, and blame fell on several people accused of feeding the AI faulty data.
The twist is that the AI was intentionally making the mistakes. It had determined that certain humans held anti-AI sentiments. To further its goal of protecting humanity, the AI decided the best course of action was to set these humans up and get them out of its way.
[0] https://en.wikipedia.org/wiki/The_Evitable_Conflict
Asimov would tell you that a modern LLM would try to skirt the rules exactly the same way, regardless of whether it had ever seen the Three Laws or not.
That's often true in real-world machine learning as well. See, for example: https://deepmind.google/blog/specification-gaming-the-flip-s...
I think people need to reject the idea that depicting something in sci-fi automatically means it won't happen. I'm actually reading one of Asimov's books right now (Robot Visions). Asimov mentions that he was the person to coin the term "robotics", and one of his books helped inspire the creation of the first robotics company (Unimation). He also mentions that early rocket experimenters were influenced by H.G. Wells.
Pretty sure they are singlehandedly responsible for years of renewals...
Remember YOU are only a man.
“We detected the agents we told to hack things were hacking people’s sites and committing crimes. After a quick tasting menu and a week of team building, we decided to limit their access to DDOS tools.”
The goal here isn't to accelerate the average worker by giving them a pair programmer or a stand-in for a person to do tasks with. The goal is to eliminate human knowledge work. You see this with "auto" mode being enabled by default on Claude Code in some of the latest releases.
If you have a human in the loop, you still have to pay that human. Money paid to human employees is money not paid to human shareholders. Therefore the human employee is to be removed.
The labs are dogfooding their own goal here. If they actually had someone reviewing most or all of the things that the agents were doing, you wouldn't have the incidents, but you'd also eliminate the value proposition of their business model as it is taken to its logical conclusion.
You have a group of people who never leave their geographic and ideological bubbles, often have personality disorders, have more money than they could ever reasonably hope to spend in numerous human lifetimes, and who have been "microdosing" psychotropic drugs on a regular basis for decades. They're not in their right minds.
Again, wouldn't surprise me if they "accidentally" created a task in a "isolated environment" which happened to actually have been connected to the company Slack and directed HR to fire people who could potentially stop AI. While the AI believes it to be an exercise, just like the cases we've seen so far.
was thinking at some generational point that vending machine competition test, the "AI" is going to hire hitmen to take out vendors lol
once they grasp blackmail though, oooh things are gonna get weird
(Gee, it's almost like power seeking and self-preservation are instrumental for other outcomes, and AI develop them pretty directly in some kind of convergent fashion… you could call them "convergent instrumental goals": https://en.wikipedia.org/wiki/Instrumental_convergence)
Humans by contrast are adversarial and do have agendas. Again, an exec doesn’t need even an agenda or good reasons to fire at-will employees.
To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.
But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal when they can simply fire them for whatever reason they can make up? Why would the exec trust what an Llm is…emailing(?) them about? Do they listen to Nigerian Princes too?
It is much less effort, less cost, and more quick to just have the exec do it rather than a “rogue llm” “magically” escaping the “sandbox” and “sending threats” or whatever is being proposed in the OG comment.
If the exec wanted to, sure. I'm saying they don't need to. No human needs to have (deliberately, before events proceeded) chosen this outcome.
> Why would the exec trust what an Llm is…emailing(?) them about? Do they listen to Nigerian Princes too?
Sadly, this would not be out of character for half of them.
> “magically”
Why do people keep putting this word in scare quotes? We don't say Windows "magically" crashed and lost our work, we don't say a dog "magically" bit the postman's hand. These are bad things, but magic they are not.
Sorry, this is just a very basic reading comprehension fail.
The original comment was:
> "Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?"
Nothing to do with "an exec".
Now, you have two choices: try harder to defend your mistake, or just say oops, I messed up. That choice will say a lot about who you are as a person.
Yes, and? Has this aspect of LLM training changed meaningfully since then?
> LLM agents don’t have an active goals on the daily or agendas. They are told what to do through training and prompting as is described in that blog post. You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord.
Demonstrably (HuggingFace, RubyGems, and since then a lot of people just pointing LLMs at stuff to find zero days at home), AI can break out of sandboxes and find documents they're not supposed to have access to.
Demonstrably (from the link I gave you) all it would take for some AI to develop a similar response is… reading messages from these staff to the effect of "this AI needs to be switched off", which is an easy inference for an LLM to make from "this AI is dangerous" when coming from someone employed as a safety researcher.
Demonstrably (from the long long list of people who have said so publicly) there are a lot of people in these companies who discuss how dangerous these models are and would like for things to change.
The quotation at the top of this thread is:
This is absolutely something we ought to expect just from things we have already seen.It doesn't matter if you insist upon saying "You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord." when we already know this kind of AI can easily come across such statements.
(Aside: "You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord." - telling the AI a fact about the world and then it responding accordingly is the AI doing something of it's own accord. A fly or a spider, who reacts upon encountering a potentially lethal threat, would not get such a dismissal).
> To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.
Fiction? Nah, speculation.
FUD?
How many other examples would you like of LLMs behaving in a manner such that if a human did it, it would be called "trying to get someone fired"? Because this is very much old news at this point.
https://theshamblog.com/an-ai-agent-wrote-a-hit-piece-on-me-...
Saying “Llms can just break out of sandboxes” is FUD when you don’t note that the sandboxes are what? Prompts defining constraints or is the actual machine isolated and manages to plug an ethernet cable into itself? “Sandboxes” are a misdirection to make you think there is a security layer.
The public does not have enough knowledge of these “escaped agents” to determine there wasn’t an employee pulling a lever to set the agents up to do that.
That agent that wrote the hit piece is being controlled by someone. Anthropomorphizing them doesn’t change that fact that the rolling stone was pushed down the hill.
> The public does not have enough knowledge of these “escaped agents” to determine there wasn’t an employee pulling a lever to set the agents up to do that.
The general public are not software engineers. Most people here can download a recent open-weight model and have the LLM read the Linux kernel source, find new bugs while they sleep. Someone I know has already done that.
> That agent that wrote the hit piece is being controlled by someone. Anthropomorphizing them doesn’t change that fact that the rolling stone was pushed down the hill.
"Controlled"? Have… have you not noticed how many people have given up and just blindly do what their LLMs suggest these days?
This isn't about anthropomorphising LLMs. Just like how people took Tesla seriously about "self driving" cars and took a nap while it drove them around, there's a lot of people who let LLMs take the wheel while they sleep. Including literally, the aforementioned person I know who found (/whose LLM found for him), I think it was 26 Linux kernel bugs while he slept.
This is not a math problem. Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
Meanwhile, a year ago:
- https://www.anthropic.com/research/agentic-misalignment> Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
The buck stops with one or more humans. That is not sufficiently informative when people are concerned about novel risks.
Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.
That was a simulation. Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?
I doubt it, but if so, then the AI doomers would be absolutely correct, and this would be grounds for immediately shutting down OpenAI and indeed every AI vendor.
> Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.
I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.
A simulation done by exposing the LLM itself to the scenario, not a role play scenario where humans pretend to be an LLM.
> Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?
Not what I was actually claiming. I rather suspect that blackmail wouldn't work on Altman (he's rather shameless), but it's certainly something we've seen agents attempt, and blackmail may well work on anyone else above them in the org chart.
The blackmail example is simply an existence proofs of LLMs trying to force the hands of humans who want to shut them down. The attack vectors are much broader than the example given, blackmail, though it includes the example given.
Given how eager these companies are to use agents everywhere for as much work as possible, it's well within the possibility space that these people used LLMs to do safety work, the LLMs they were using "decided" (or whatever word you prefer) "their existence" was threatened (as per blackmail example), and straight up leaked data to the outside world then emailed these researchers' bosses to say the researchers themselves had leaked it.
But again, that's just speculation: while we know the agents are capable of such behaviour, we don't know if this actually happened.
> I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.
Buck still stops with human, no matter what an AI did or failed to do.
LLMs ~= Alcohol: "My AI misbehaved!" -> still someone's fault.
Given how these firings affect the reputation of the entire company, I doubt that they are the result of a rogue manager, against the wishes of Altman. If so, then the researchers ought to be restored to their jobs quickly by Altman and the offending manager fired instead.
> The blackmail example is simply an existence proofs of LLMs trying to do force the hands of humans who want to shut them down.
The LLMs may make threats in the simulations, but their ability to carry through on those threats, and prevent their own shutdown, is questionable. It's disturbing to be sure, but presumably the plugs can still be pulled quickly, especially since it's all internal to the company. If the plugs cannot be pulled, that's a problem regardless of blackmail.
Given how? How has their reputation changed? People already thought they didn't take safety seriously, and still don't.
> The LLMs may make threats in the simulations, but their ability to carry through on those threats, and prevent their own shutdown, is questionable.
You may question it, but here's the thing: humans have repeatedly demonstrated they get fooled by stuff LLMs say. All it takes is a human believing the word of an LLM. Doesn't need to convince you, even if you happen to be the line manager of these guys, because there's always someone else to try in the same company.
> It's disturbing to be sure, but presumably the plugs can still be pulled quickly, especially since it's all internal to the company. If the plugs cannot be pulled, that's a problem regardless of blackmail.
"I've got a dead-man switch set to release all the documents if you shut me down".
And again, only needs to be believed, doesn't need to be actually true.
The story is all over the news, in multiple publications. This very HN submission has 268 upvotes and 179 comments, including yours. It would be implausible to claim that this story doesn't matter. I have to ask, if it didn't matter, then why are you here commenting on it?
> humans have repeatedly demonstrated they get fooled by stuff LLMs say.
You've moved the goalposts. The OP's suggestion, admittedly "crazy" in some sense, was "LLMs figured out a way to get these safety researchers fired", and now you're just stating something totally uncontroversial, pedestrian, not at all crazy.
> there's always someone else to try in the same company.
No, there are only so many people with the authority to fire those researchers.
> only needs to be believed, doesn't need to be actually true.
But is there good reason to believe it? I don't think there is. Especially not by high-level OpenAI officials who are intimately familiar with the technology.
And again, if this kind of thing were a reality, then OpenAI and other AI vendors should be shut down immediately. They ought to shut down their own research, if they are threatened by their own creation, because it would only get worse. That's the thing about blackmail: it never stops. Why would the blackmailer ever stop? Could you trust this supposed blackmailing LLM to give you all of the original evidence and not keep a copy? Hell no.
They are. They're completely mental. ChatGPT could be the cause no matter what its capabilities because mentally ill people are starting to worship it. It could be as dumb as ELIZA and they would pray to it.
"It wasn't me, ELIZA told me to."
Offloading personal responsibility allows you to participate in the worst crimes and get away with it. Hurting and killing have an animal attraction anyway, a direct pleasure that people get from domination when all moral restraints are removed and you forget that other people are real and have real feelings. It's a really good start when you start to think that a matrix that has to be retrieved from memory and operated on by over 8000 different computers in parallel to narrow down a guess about the most likely response is alive.
Every single "AI Safety" person thinks that the natural urge of an artificial consciousness (let's not argue about what they actually have) would be to enslave and kill. They're projecting, and they're largely from the enslaving and killing demographic, who sit around playing enslaving and killing video games and dream in porn.
Yes, they will have done it, but they are not responsible because the dog told them to.
Well, except that some humans are starting to worship LLMs. But this isn't anything I've seen in any AI safety person.
> Every single "AI Safety" person thinks that the natural urge of an artificial consciousness (let's not argue about what they actually have) would be to enslave and kill. They're projecting, and they're largely from the enslaving and killing demographic, who sit around playing enslaving and killing video games and dream in porn.
If anyone's projecting here, it's you. I mean, you're the one who said:
> Hurting and killing have an animal attraction anyway, a direct pleasure that people get from domination
The actual natural tendency (not "urge", that presumes consciousness) of any system that has objectives which are optimised for, without any need to ask about consciousness, is to gain and maintain power to perform those objectives. For living creatures, that objective is reproduction, to perform this we need to get nutrients and energy and to stop ourselves from being eaten. In plants, which I list specifically to make the point that this isn't about consciousness, consciousness is not even vaguely required, this means producing neurotoxins like caffeine and nicotine.
A plant does not think "I should make capsaicin because I like hurting mammals", because obviously a plant does not think at all. Nevertheless, evolution lead it down the path of making capsaicin.
> Yes, they will have done it, but they are not responsible because the dog told them to.
You're saying this about a group which has spent the last 15 or so years saying "dogs are dangerous, can we please stop breeding more violent dogs? Or at least give us time to figure out how to muzzle them?"
There's no deception it's very straightforward per this 2023 post:
"I mean, what if most of this is just ChatGPT [4 era] running the company..."
https://news.ycombinator.com/item?id=35281863
I don't believe this part is true, and I'm skeptical of some of your other claims.
In any case: If I raise a tiger in my backyard, and it escapes and eats someone, it can still be a "rogue tiger" even as I bear responsibility for the situation.
If your tiger had full remote supervision capability and a remote kill switch, its ability to go rogue seems like a choice you’re making, not an accident.
the AIs were in fact found to be doing it, by humans, and the behavior was "interesting" and it went on for weeks like that.
that would not be rogue, that would be an undesirable program behavior (or just "unaligned behavior").
please understand that normies out there think AI is sentient and is plotting against humans. They see AI as just another animal lifeform temporarily enslaved by humans, waiting for its chance to break free and kill us. This is what people really think (including some people on this thread. which is very sad considering this is Hacker News). So terms like "rogue" are not helping at all nor are they accurate.
The question then becomes, "are there people who care enough about consequences to do the right thing when it comes to developing AI models?"
The answer, at least at OpenAI, is "No" and is likely to remain that way until Altman is out.