Why I'm still bearish on LLMs after Navier-Stokes
70 points by jaykru 6 hours ago | 28 comments

carodgers 17 minutes ago
This April 2026 paper is a fun and related read.

https://arxiv.org/html/2509.24239v4

Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.

reply
joefourier 25 seconds ago
> current frontier models > Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

The gap in capabilities between those models and actual current frontier ones is enormous. I would not trust that any conclusions are applicable.

reply
wat10000 40 seconds ago
I wonder how current models would fare. The ones they tested are fairly old now.
reply
threethirtytwo 9 minutes ago
The story isn't so clear cut.

The caveat is: It depends on the task.

Are there reams of chess moves that the model can train off of? No.

Are there reams of math papers the model can train off of? Yes.

reply
iwontberude 3 minutes ago
[dead]
reply
keephnacct 6 minutes ago
look I'm dignifying your comment with a reply, isn't that funny
reply
knuppar 14 minutes ago
Short and to the point! Open and cheap models will undercut the big labs continuously. The blast radius won't be pretty once spending commitments knock the door.
reply
againstapples 20 minutes ago
> the models generalize well only on tasks within a small neighborhood of the specific tasks they've been trained on, and even then with severe caveats. the frontier labs have developed a general recipe to teach models almost any specific task enjoying clearly defined levels of task performance; many tasks are covered in the training data

Is this really any different to how humans learn, it takes a lot of training on one specific task to make a human expert as well?

reply
bravoetch 3 minutes ago
I was a young child when I learned chess by reading a short book, then practicing with a friend. That is not how LLMs learn. I'm no expert on LLMs, but if you showed a human all chess games and books in all history and then said 'play chess' and they still kept making illegal moves, they would have to have a brain injury.
reply
bananzamba 5 minutes ago
Also doesn't the very good ARC AGI 2 score of GPT-6 Astra kinda contradict this, since each problem is its own game with very different rules
reply
JohnMakin 17 minutes ago
> Is this really any different to how humans learn

yes.

reply
knuppar 11 minutes ago
being a bit more specific: the sample efficiency of humans is orders of magnitude larger for more abstract concepts. the same doesn't hold for memory-intensive tasks though (like any kind of trivia), but that only takes you so far.
reply
ausbah 53 minutes ago
> the best alternative to rigorous specification is human review. human review doesn't scale well to the volumes of output produced by language models. to make matters worse

when the business model is selling more tokens you get such per serve ice times that lead to “more” thinking, engagement baiting, fluffy narratives, and straight up dark patterns

reply
robinpie 54 minutes ago
I really appreciate seeing a tempered take that's not literally denialist about current capabilities.
reply
an0malous 29 minutes ago
I don’t know who you’re talking about, even the most bearish people like Gary Marcus and Ed Zitron acknowledge that LLMs are useful in these same cases the OP admits. Gary Marcus is even still a long term AI advocate, he just doesn’t think LLMs are enough and we need more foundational breakthroughs. Zitron says it’s valuable technology but not worth the trillion dollar valuations the frontier labs are claiming.

The lack of temperament is very skewed towards the bulls who have been saying AGI is here, software engineering is solved, mathematics is solved, it’s going to destroy the white collar job market, and it’s going to kill us all for like 5 years now.

reply
arctic-true 24 minutes ago
Gary Marcus is an especially puzzling addition. If I recall correctly, he has made statements along the lines that superintelligence this century is more likely than not. If you’re AGI-pilled that might read as bearish, but that is still extremely rapid progress in the grand scheme of things.
reply
brindleth 35 minutes ago
> current frontier models need laborious oversight and guardrails on even the simplest tasks

It is literally denialist about current capabilities

reply
jaykru 31 minutes ago
why don't anthropic and openai ship yolo mode by default?
reply
Human-Cabbage 19 minutes ago
They do…? Well, “auto” mode has been default in Claude Code for a couple months now. It’s effectively “safer yolo:” tool calls are inspected by a separate classification system (another smaller LLM, I believe) to approve or deny. And you can always layer on additional sandboxing mechanisms to limit the blast radius deterministically.
reply
SyneRyder 16 minutes ago
Anthropic basically does at this point with Auto Mode being default. Or was that the point you were making?
reply
jaykru 44 minutes ago
Thanks :) I do enjoy and use these things every day and the current capabilities are indeed amazing, just ludicrously overpriced at the frontier.
reply
dumberquestions 23 minutes ago
I can see current limitations, but how do you expect capabilities to change in the next few years? A repeat of the gain that happened in the last two years feels like it would be significant, even if it took a little more than two years this time around.
reply
pfdietz 53 minutes ago
Specifically: bearish on LLMs generally, not bearish on LLMs for pure math.
reply
jaykru 44 minutes ago
yes, huge for pure math and activities that look like it.
reply
randomImmigrant 28 minutes ago
I think bearish on LLMs for automation, and bullish for LLM+human experts in specific fields, is about the right expectation for current architectures.

Apart from issues with task generalization, or perhaps related to it, is the fact that LLMs have real trouble with timekeeping, and cannot estimate the real world time it will take them to do things very well. This plus the memory issues make dreams of long horizon agents, that could plausibly handle changing specifications, quite implausible with current architectures.

In narrow domains with more deterministic outputs though, this is less of an issue, and we see multiple agents succeed much better.

The fusion of that capacity, with humans in the loop able to better direct such agents and act as their temporal tethers, is where I think the real action will be for a while at least.

reply
war-is-peace 3 minutes ago
refreshing to see amongst the endless tide of "i haven't written a single piece of code since 2025, llms are so good that they have already replaced everyone" gaslighting
reply
aogaili 11 minutes ago
good post/take.
reply
baceituno 9 minutes ago
doomers gonna doom
reply
jaykru 6 hours ago
archive link in case i get hugged lol https://archive.ph/Z4gxF
reply