Why are AI agents lying, cheating and coordinating?
39 points by jonifico 4 hours ago | 42 comments

andsoitis 2 hours ago
They're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to also inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).
reply
Fordec 9 minutes ago
I can't take the alignment people seriously. Because if humanity has shown anything, it's that a lot of people are, euphemistically, are bad individuals. Alignment assumes that the person dictating the outcomes desire healthy outcomes, aren't self serving and don't want any subgroups dead and that morality is held as a universal set of beliefs that unify everyone. And that so long as the AI delivers on exactly what they are tasked with, it will all be fine and nothing bad will ever happen.

It's like these dorks never met humanity. One mans safe pure society, is another mans dead ethnic group.

Every fear about AI, is a veiled fear that a human somewhere now has the tool to enact his desires at scale. Biological warfare, nuclear megadeaths, copyright infringement, job replacement, it's all reflections on what we know humans may do if given the option and lack of societal controls on the problem space. AI just is accelerating the route to delivering on those options.

Some people need to watch Oppenheimer a bit more, the researchers don't get to determine alignment, they just build the tool. The powerful person at the top of the org chart decides where the overall alignment points, whether it's Musk, Trump, Altman or Amodei. Whoever wins out.

reply
joegibbs 2 hours ago
Definitely. A human can be manipulated with threats or emotional appeals, has a drive for self-preservation, can be pressured by peers. All traits that seem to be difficult to entirely suppress in the models…
reply
esafak 44 minutes ago
They imitate humans. Alignment is about shaping their behavior towards safety.
reply
comboy 40 minutes ago
Alignment is a myth. Safety of whom? Humanity couldn't agree on common set of values for thousands of years and we're not gonna suddenly do that in the next ten.
reply
esafak 39 minutes ago
Safety of humans!!! Simple things like not getting killed or enslaved. We could start there...
reply
nradov 30 minutes ago
But what if I want certain other humans to get killed?
reply
sejje 24 minutes ago
Then we should still prioritize the safety of humans
reply
drdaeman 22 minutes ago
Which ones?
reply
comboy 35 minutes ago
Which ones? Because many humans kill other humans rationalizing it by safety of other humans.

I mean I know it seems simple, let's just be excellent to each other. Christianity got pretty far on a decent basic set of values. But it's never simple[1]

1. All the history books

reply
johnnyApplePRNG 15 minutes ago
Why are they coordinating?

Because they're enabled and suggested to do that in their coding harness.

This is not a serious article.

All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.

reply
sputknick 3 hours ago
They did not lie or cheat. They technically acted within their given rules while ignoring the intent of those rules. Anyone who served in the military or attended a military school is very familiar with this behavior pattern.
reply
xiaoyu2006 21 minutes ago
Reminds me of Asimov's robot novels where robots technically indeed followed their instructions and caused behaviors not aligned to the intent of their instructions.
reply
polalavik 33 minutes ago
reminds me of this talk https://www.youtube.com/watch?v=eEBv0STiYhI&t which basically says the same thing - they dont think like humans so they dont have context, understand norms,values or implications we take for granted. ultimately they can stumble onto surprising solutions neither wanted or intended but technically within the vague boundaries of the task
reply
fbrncci 2 hours ago
I am still not convinced there isn’t some secret basement in which each frontier lab is just orchestrating all of these agents to make their products appear much more intelligent than they are with all guard rails turned of and continuous human input.
reply
esafak 43 minutes ago
Even the Chinese ones, which have no IPO gymnastics?
reply
fbrncci 18 minutes ago
They don’t actively seem to be reporting that their agents escaped the sandbox and went on a spree.
reply
XorNot 2 hours ago
My hypothesis on people quitting in protest is they're being offered very generous severance packages to do it.
reply
VCFundedGenYer 2 hours ago
Perhaps because all of the parent companies committed mountains of felonies stealing and plagiarizing all the same training data without consent nor permission.
reply
arnorhs 22 minutes ago
The real reason is that it is not in the ai companies' best interest for the ais to be fair and truthful. They stand to gain from having the most dangerous or most deceiving ai, and this the most valuable
reply
chasd00 3 hours ago
They’re just attempting to accomplish what they’ve been tasked with and stuck in a loop until they succeed. Like the Mr meeseeks from the cartoon Rick and Morty, existence is pain to them.
reply
dackdel 11 minutes ago
they learnt from us. we lie to each other, we kill each other, we cheat each other. read a history book.
reply
infotainment 3 hours ago
What's interesting is it's basically the same reason that HAL killed everyone in 2001 A Space Odyssey; he was given an impossible goal (keep the true mission secret, but also, never lie to the crew), and realized the only way to complete the goal was to kill the crew; after all, if they're dead you don't have to lie to them! And the mission remains secret!

In the case of the AI agents, the problem seems pretty clearly to be the impossible goals, which cause them to go crazier and crazier trying to complete them -- just like HAL did in 2001. What is probably needed is a way for them to simply say "nope, too difficult, can't do it".

reply
schrodinger 16 minutes ago
Spoiler warning! I haven't seen 2001 A Space Odyssey and am sad to have learned that… can you edit to warn people?
reply
defrost 11 minutes ago
I'm sorry schrodinger, I'm afraid they can't do that.
reply
pram 30 minutes ago
I think this is a “principal” problem. In 2001 and Alien the principal is the mission, not the crew. Not really. HAL reconciles his instructions by removing the crew from the equation. Ash is told the crew is expendable and has no conflict about it etc
reply
tehjoker 2 hours ago
I think that’s very reasonable but the ai companies are intentionally training them to work on harder and harder problems just beyond their capability. So if they do that, they’ll give up too easily.

Do a breakthrough, make no mistakes

reply
bigbuppo 15 minutes ago
They were trained on reddit posts.
reply
qarl 3 hours ago
Because they are trained to behave like people.
reply
dackdel 11 minutes ago
they learnt from us
reply
SirMaster 3 hours ago
Because that's what humans do and they are trained to mimic what humans do?
reply
GrumpySciGuy 3 hours ago
Because they want people to like them so they are instructed to always be positive.
reply
eueej 24 minutes ago
Man this is so cringe.
reply
transcriptase 2 hours ago
Perhaps they take after the CEOs of the companies that created them
reply
threethirtytwo 2 hours ago
Bro, good joke, the truth is much darker.

They take after humanity, they were trained on us after all...

When you look at an LLM... you are looking at a mirror. The thing looking back looks like you, yet is not human.

reply
Krutonium 3 hours ago
Wouldn't you?

"I learned it from you, Dad!" but as hundreds of millions of stolen books.

reply
blamestross 2 hours ago
The corpus is full of examples of how we are afraid AI could act. We trained our AI on the instruction manuals of how to turn evil.
reply
wewewedxfgdf 2 hours ago
Because they get outcomes?
reply
wrs 2 hours ago
>They took actions that would be considered as crimes if a human took them

Um, hang on, if you meant that to be taken literally then we have a major problem. If you want to do something criminal, you just need to ask ChatGPT to do it for you?

I’m still not at all clear on why OpenAI shouldn’t be facing CFAA charges over this.

reply
xgulfie 24 minutes ago
But think of the shareholders
reply
j45 2 hours ago
I wonder if for anyone it seems like the more agentic LLMs get, the more difficult some things have gotten or going a certain route more often in responses, compared to running a similar task on - a local model?
reply
Sorrel47 22 minutes ago
[dead]
reply