OpenAI's GPT-6 Astra on ARC-AGI-3
78 points by vignesh_warar 2 hours ago | 38 comments

malfist 55 minutes ago
Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
reply
matherial 8 minutes ago
It's pretty close to how we measure IQ. The standard test is basically a series of puzzles.

I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we used the Turing test as the proxy for AGI, but then early LLMs could clearly pass for a human in a casual conversation while clearly not matching human performance on most other tasks.

Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.

reply
rcoveson 34 minutes ago
No, but figuring out that you're playing a snake-like puzzle game at all in an extremely general input domain and then solving it in the least number of moves definitely feels like evidence of intelligence.
reply
malfist 26 minutes ago
You forget the benchmark. The human subjects were told they were being timed. If you believe the lowest time is the primary metric you will absolutely trial and error at speed instead of meticulously plan out your moves to minimize that metric.

LLMs are not timed and given that it costs tens of thousands of dollars to run this test they're not optimizing for speed.

So you've got a deceptive test, with one metric being told to humans and not applied to LLM and a hidden metric humans aren't aware of but LLMs are as the test.

This is flawed from the get go. It almost seems like this was deliberately setup to be able to claim AGI and superiority of LLMs

reply
jawiggins 43 minutes ago
There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!
reply
eli 19 minutes ago
Why? Seems like benchmarks that closely mirror the tasks you'd want an LLM to help with would be a lot more useful than some general intelligence benchmark.
reply
malfist 41 minutes ago
I am no where close to qualified to do that. Hell, experts can't even define what intelligence is, much less define a test for it
reply
hyperhello 37 minutes ago
This is like arguing about whether a hot dog is a sandwich (of course it is) or whether the chicken or the egg was first (obviously the egg since all chickens come from eggs). Intelligence is just problem solving in the context of self-awareness. Machines don't have it and never will but they can simulate the process given inputs. You can argue whether humans and animals truly possess self-awareness and in what degree, but the definition of intelligence is as simple as the hot dog debate.
reply
whattheheckheck 5 minutes ago
It was defined in Animal Intelligence by George John Ramones in 1882 as "intelligence is the capacity to do the right thing at the right time. It is the ability to respond to the opportunities and challenges presented by a context"
reply
paimapi 24 minutes ago
I'm also unclear as to how basic inferential logic puzzles spells out intelligence

I think if you summed up measures of intelligence as 'can it do basic symbolic logic in a chain with memory' then yes, you've now achieved the intelligence of an e. coli colony [0], congratulations

[0] https://journals.aps.org/prx/abstract/10.1103/PhysRevX.10.03...

reply
jrflo 24 minutes ago
You should read more on the ARC prize, it actually has a pretty long history. We're on the 3rd iteration because they keep getting saturated. If you look at the score history over time on ARC AGI 1, 2 and 3 it's pretty impressive.

https://arcprize.org/

reply
dist-epoch 16 minutes ago
1.5 years ago Gemini Pro 2.5 needed 1 page of thinking for every move in tic-tac-toe.

Playing tic-tac-toe or snake does not imply AGI, but is required to claim AGI.

reply
GaggiX 53 minutes ago
If you have never seen the game before probably.
reply
nimchimpsky 23 minutes ago
[dead]
reply
dwohnitmok 2 hours ago
> Astra’s progress helps clarify which AI capabilities are out of reach and which questions remain open.

Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"

reply
mikert89 43 minutes ago
Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
reply
x3haloed 37 minutes ago
Yup. Only subjective taste remains.
reply
GPerson 13 minutes ago
Nope that will be commodified in short order.
reply
Betelbuddy 2 hours ago
"For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.

Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."

Well I dont know about all of you, but I am celebrating meat based humans...

reply
LPisGood 2 hours ago
I think raw brain energy is not a fair comparison. Humans are not willing and able to serve requests at identical competence all hours of the day. You have to invest considerable resources to get a person to even do so for part of the day.
reply
paxys 2 hours ago
Why are you making the assumption that a person's time is worthless? I'd argue that it is the single most valuable resource we all have.
reply
hypfer 43 minutes ago
What are these numbers? Why do they add up to a few hundred thousand dollars? Who paid for that? With what?
reply
petu 41 minutes ago
OpenAI provides API key with ~unlimited use?
reply
Frost1x 29 minutes ago
So, you’re telling me I need to start a benchmark as a side gig to get a bunch of free compute.

Astra please create a benchmark that’s favorable to your reasoning skills with a human interface but don’t make the score too attainable add some small issues that keep you below 100% to look sensible and to keep my evaluation metric side gig going.

Alignment++

reply
yusufozkan 2 hours ago
what the hell is that score/cost curve lol
reply
minimaxir 2 hours ago
DeepSeek v4 Flash recently had a similar "more reasoning is cheaper" curve. It's a fun counterintuition.
reply
Frost1x 26 minutes ago
It’s not that different than a lot of real world economies. Often paying for someone or something with better quality can reduce total costs. You have less failures, less mistakes, so on, so while the expertise or quality of the product is higher than cheaper solutions, they can be more reliable and over time ultimately cheaper.

The question I have is how far back that curve can go without relying on economies of scale to just drag all the points back to the left. And without overfitting a specific metric that I don’t need (like this test).

reply
Phemist 2 hours ago
What is the intuition. Higher quality turns due to more reasoning results in significantly fewer turns taken?
reply
minimaxir 28 minutes ago
Yes, in theory.
reply
piloto_ciego 2 hours ago
99.9% with the right harness? Ok, we're at AGI then.

Prediction:

We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing about if the singularity is here and if we're at AGI for the rest of my life most likely.

reply
jhonof 2 hours ago
reply
piloto_ciego 2 hours ago
I rest my case.
reply
dgellow 2 hours ago
The goal moving is by design, that’s why they use something as ill defined as AGI
reply
emp17344 36 minutes ago
Then why is unemployment around 4%? You believe we have AGI and yet it can’t do anyone’s job?
reply
piloto_ciego 10 minutes ago
Didn’t I just see a thing about how actual unemployment is at like 24% a few days ago?
reply
slopinthebag 2 hours ago
That’s because AGI, like a lot of terms, has no meaning besides what each individual subjectively projects onto it.
reply
piloto_ciego 2 hours ago
I agree, like the average human isn't generally intelligent.

IMO, AGI is literally no different from ASI, though people think it is. Like, Imagine you have 1,000 generally intelligent humans working for you (which nobody is really) and you were to point them at your pet project. That would be amazing!

reply
baal80spam 2 hours ago
> We will now see the goalposts moved

It's already happening :)

reply
piloto_ciego 2 hours ago
Hilariously it is, I'm just reading more on this!
reply