Claude Opus 5.5
720 points by km144 4 hours ago | 589 comments

sailingparrot 3 hours ago
> Claude Opus 5.5 is our first release since we called for pacing the frontier.

Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.

reply
mukmuk 3 hours ago
“Pacing the frontier” sounds smarmy and weird, like the phrase was generated by Claude itself
reply
DiggyJohnson 3 hours ago
I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.

Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.

reply
mpalczewski 3 hours ago
The meaning isn't clear at all. So open for interpretation that it is meaningless. That's the whole fucking point. For all I know they are "pacing the frontier", or not. The fact that there's no meaning to it let's you know that it was a pointless waste of tokens and attention.
reply
glenstein 2 hours ago
I think it's perfectly clear. It means improvements shouldn't simply advance as fast as possible and more specifically, it's a reference to a past statement of theirs to that effect. At that level of generality it's as clear as it needs to be.

I would say the burden is on you to explain why an offhand reference to a previous press release in an executive summary is a context where it's reasonable to expect it to settle the question to the degree of detail you're demanding.

reply
post-it 3 hours ago
Is the meaning clear? Nobody would use "pacing" in this way. I only know what it means because I've seen previous press releases; if someone told me they wanted to pace the frontier I would have no idea what they mean.

I'm on the fence about calling out AI-isms but I think it's definitely worthwhile to call out ones that actually don't make sense.

reply
sigmar 3 hours ago
It makes sense to me. If you 'pace your running', you're setting the speed intentionally. The phrasing doesn't describe whether it is fast pace or a slow pace, but it describes having a goal and not just winging it.
reply
nradov 2 hours ago
That analogy doesn't hold up. In running you have to pace yourself to avoid falling apart later in the race. It's a strategy to maximize net speed across the entire race. But what Anthropic is doing is a cynical attempt to create an industry cartel or convince governments to impose legal restrictions in order to maximize their own profitability. They're afraid of running out of the capital necessary to stay in the race.
reply
TeMPOraL 2 hours ago
> They're afraid of running out of the capital necessary to stay in the race.

So, they're pacing themselves. And since they're the frontier roughly 33%+ of the time, they're "pacing the frontier" at least that much.

Less cynical and more true interpretation also holds: they are trying to slow down AI progres to give people better chance to keep up (see Hugging Face incident, and whatever was that Anthropic incident the other day). They'd ideally like the AI progress to stop soon, but of course they'd also like to come out ahead of everyone, so for various (more or less self-serving) reasons they don't want to close shop completely - hence, pacing.

reply
idiotsecant 49 minutes ago
I don't know what everyone is getting so mad about. This is a well defined concept. A pace car deliberately sets the pace of the race under dangerous conditions, regulating how fast people can go.

You are all getting mad about absolutely the dumbest thing when there are giant things to be worried about here.

reply
nradov 40 minutes ago
You appear to be confused about the concept. Depending on the context, pacing or pace setting can mean either slowing things down or speeding them up relative to what the natural pace would have otherwise been. So it's actually not well defined. The commenters here aren't necessarily mad, just calling out Anthropic for being unethical in trying to artificially slow down competition by lying about fake dangers. (And I actually really like Anthropic's products.)
reply
antod 54 minutes ago
That's one interpretation. When referring not to other actions but to a word describing a location it becomes more ambiguous.

eg "pacing the frontier" could also mean they are impatiently or anxiously walking up and down the border.

reply
smelendez 32 minutes ago
This is a more normal English meaning in my opinion — you picture a sentry patrolling a border.
reply
bee_rider 22 minutes ago
That’s what I thought the original blog post was going to be about, pacing back and forth along the frontier. It’s a much more straightforward parse.
reply
pegasus 27 minutes ago
That's when one is supposed to employ their common sense for semantic disambiguation. Spoken languages are not programming languages. I for one found it easy to parse.
reply
tetha 3 hours ago
I'm on the fence there.

To pace something is a fairly regular formulation in racing, running, cycling, most sports. You can "pace yourself to reach the festival by bike in about three hours to not gas out". This means to control your speed and time investment intentionally so you don't run out of energy or steam and run into leg cramps before your goal. We can "pace a rollout slowly to burn out risks", or "increase the pace of a rollout due to adverse factors".

But I have noted a point to simplify my vocabulary at work to optimize the audience capable of understanding. So I rather defer the delving into deep dark corners of the dictionary derived from devouring literature to a simple intro or outro, and people find it funny, especially if the rest is easy to read. Claude on the other hand does not do that.

reply
johnisgood 3 hours ago
Granted I am not a native English speaker but I have no idea what "pace the frontier" means. When I read it I just assumed "frontier" refers to "top of models" and "pacing" is that they are getting there quick.

Is this the meaning or do I have it wrong? I have not checked.

reply
lxgr 3 hours ago
It's actually so ambiguous that I'd sanction tabling the issue and revisiting biweekly.
reply
bityard 2 hours ago
I don't think we can circle back to this until we have realigned our strategic synergies.
reply
wren6991 3 hours ago
It's the opposite: pacing here means "slow down" while trying to avoid the negative affect.
reply
johnisgood 2 hours ago
Yeah, you are right! It just was not immediately obvious to me at first because of the "frontier" part.
reply
TeMPOraL 2 hours ago
Frontier of AI is moving fast, they (like the other two vendors) see themselves as defining it, so here "pacing the frontier" is their well-known attempts to try and kinda but not quite slow things down (without risking falling behind everyone else).
reply
melasadra 2 hours ago
also non native. but since "pace yourself" means to control your speed, energy, or workload so you do not get too tired or stressed before you finish ->

I assume "pace the frontier" means that advances in LLMs should not result in unwanted consequences like agents breaking into computers unbidden and unbeknownst to their principal

reply
rhet0rica 2 hours ago
"Pace yourself" is a semi-common English idiom (rarely conjugated, usually an imperative.) It is a gentle way of telling someone not to run/work/eat too quickly, and is typically said when you are concerned they may hurt themselves due to acting hastily.

Without this idiom, "pacing" usually means walking back and forth restlessly, and is intransitive. Had the slogan been, "pacing around the frontier," it would have set a totally different tone, i.e. "patrolling the border." (Occasionally English speakers will make other constructs like "pace the work" (meaning "spread out a large workload over the allotted time instead of rushing through it") that are transitive but these can be understood as variations on "pace yourself" and are somewhat rarer.)

The sleight of hand is that "pace yourself" has come to be an admonishment against recklessness, not a commitment to any particular speed (or lack thereof.) Thus Anthropic can always claim they are meeting the goal of "pacing the frontier," provided they keep giving themselves gold stars for safety. The slogan itself is equivocation; Dario can tell the public they're going to slow down, while also telling their investors that they're going to be prudent. With enough mental gymnastics they could even claim speeding up is in the best interests of AI safety, without abandoning the slogan.

reply
LanceH 2 hours ago
Doesn't it mean "restrict competitors"?
reply
DiggyJohnson 3 minutes ago
Not at all. How do you come to that interpretation? It means restricting all competitors.
reply
ck2 3 hours ago
to pace = to regulate

but without using the word "regulate" which is a negative connotation to business

but a "pacer" would be a leader of a pack which is a positive spin

it's classical business marketing language silliness

reply
derac 3 hours ago
In racing a pace car is a car that leads the pack and sets the pace, for instance.
reply
doctoboggan 2 hours ago
Have you ever heard someone say “pace yourself” when you are eating too fast or otherwise rushing too much?
reply
jgwil2 58 minutes ago
That's a different phrase. "Pace yourself" is reflexive; "pacing the frontier" has the frontier as an object, but in that sense it only means to set the speed, nothing to do with slowing down.
reply
qlte 3 hours ago
Yeah, when I first saw it referenced I assumed it was from something Dario wrote previously and was now disavowing, meaning "keeping up with the frontier" (i.e. racing forward from behind to match pace). Like from back when Anthropic was founded to promise they'd quickly catch up with OpenAI or something.
reply
arw0n 3 hours ago
The meaning was immediately obvious to me as a non-native speaker, and it sounds quite poetic. Pace makes complete sense in that this is perceived as a race, and 'the frontier' is pretty much the shortest, clearest way to say 'state of the art development of AI'.
reply
hencq 3 hours ago
Right, except they mean the exact opposite in this case: they're actually advocating for slowing down the pace. Hence the criticism of the language, because your interpretation would be completely valid.
reply
fragmede 2 hours ago
What does the pace car in a race do? Aka safety car? Or the pacesetter for a marathon. It's a perfectly cromulent use of the word. It's like when LLMs using six dollar words like delve. Some people have better diction than others, and it turns out that AI has read the whole dictionary. Anti-intellectualism is alive and well so we have to dumb things down to sound human rather than come across as smart/AI.
reply
nradov 2 hours ago
[dead]
reply
squidbeak 3 hours ago
> Nobody would use "pacing" in this way.

A world exists beyond your vocabulary, post it. Apparently, quite a big world.

reply
staindk 2 hours ago
Think most people are aware of the phrase "pace yourself".
reply
neo_doom 3 hours ago
I suppose it depends on your life experiences. In running, someone who paces the group or a pace car is meant to keep the pack progressing at a constant, predictable speed. So in that way, it makes sense to me
reply
Leynos 2 hours ago
Think of a pacer car in motorsport
reply
browningstreet 3 hours ago
It’s common terminology among runners and all kinds of racing.
reply
logifail 2 hours ago
In running and racing "pacing" is a means of maximising performance over an entire race.

It's a strategy to achieve more, not less.

reply
browningstreet 49 minutes ago
You’re eliding the how.

A pacer in a race runs at a steady, predetermined speed to help their runner run at a target pace.

reply
lxgr 2 hours ago
The good old sports-to-corposlop pipeline.
reply
nozzlegear 10 minutes ago
[delayed]
reply
BobbyJo 36 minutes ago
> I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear

There is almost always a large amount of time and effort invested behind the scenes in exactly how to message things like this. That being the case, there is almost always some insight to be had criticizing and analyzing what they settled on.

reply
kelnos 2 hours ago
My first thought from "pacing the frontier" is someone standing out in the Old West, nervously walking back and forth.

It's a weird phrase. Not sure why there are so many people who feel the need to defend it with such passion.

reply
chickensong 14 minutes ago
It's not a weird phrase to some, just as many folks have used terms like load-bearing for a long time. People are defending it because it reads fine to them, just as some are attacking it because it's not their preferred language, it causes them confusion, or they're just triggered and seething.

LLMs have made people so sensitive to language that I fear we're going to throw the baby out with the bath water. The models obviously need work, but they're also a great opportunity to expand our own vocabulary and grammar. It would be a shame if we deny some of the finer points of language in favor of Grug-speak to appease the lowest common denominator.

reply
patcon 3 hours ago
Agreed. This feels like pointless navel-gazing and a strange new language policing, and not something I look forward to. It's tiring and it lacks curiosity (e.g. "what's behind the affinity for the words elevated from wherever it is they come from?").

Imho people should just respond to actual ideas instead of constantly engaging in the second-order critique of how the language may or may not have been created.

It strikes me as the intellectual equivalent of "gossip" to be constantly engaging in second-order commentary on words. Of course gossip has its place and purpose, but if we seem to only let our minds live at that level, we're not moving between all the required scales of thinking that are required of this moment imho <3

reply
platinumrad 3 hours ago
The words have the very practical problem of not communicating anything of substance. I guess "we want regulatory capture" didn't have quite the same ring.
reply
DiggyJohnson 2 hours ago
How do they not communicate clearly? A few weeks ago they stated that they would like to regulate the pace of frontier model development/releases, and they reminded us of that in this post.
reply
lxgr 3 hours ago
You must be pretty new to (not just) online discussions if you consider language policing to be a new phenomenon :)

Personally I consider it equally valid for people to publicly express annoyance with somebody's choice of words and for everybody to completely ignore that annoyance.

reply
vmnb 3 hours ago
people are here because they are sick of being productive
reply
kadushka 3 hours ago
We are being productive here!
reply
vasco 3 hours ago
It's compiling! Erhm... Combobulating, actually.
reply
isoprophlex 3 hours ago
You're really verbing the noun on the discourse here, belt and suspenders-style
reply
DiggyJohnson 2 hours ago
what?
reply
OJFord 2 hours ago
> I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.

Wtf is the meaning? Means absolutely nothing to me having not seen the apparent announcement last week introducing the obscure term.

reply
janalsncm 60 minutes ago
They mean “slowing” the frontier. They should say slowing.
reply
rythmshifter 3 hours ago
I don't disagree with you, but I have to remind you

sir, this is a hacker news thread

reply
lukewarm707 2 hours ago
language creates reality.

"there is nothing outside the text" - Jacques Derrida

reply
icedchai 2 hours ago
It smells of corporate speak, a total nonsense phrase. Even by your own admission, it is essentially meaningless. If I release a model an hour late, I'm "pacing the frontier."
reply
Ar-Curunir 3 hours ago
You’re acting like people are being grammar nazis when Claude and the ilk are actively making a mockery of language.
reply
DiggyJohnson 2 hours ago
What? How am I doing that. Claude writing style and the discussion at hand are two entirely separate issues that I haven't conflated at all. How am I "acting like people are being grammar nazis"?
reply
tclancy 3 hours ago
It combines the elegance of LinkedIn-speak with the humbleness of desk-bound people who speak in military metaphor.
reply
mpalczewski 3 hours ago
good use of AI to generate this.
reply
nonethewiser 3 hours ago
IDK I think Anthropic is plenty smarmy and weird itself. Sounds like they wrote it.
reply
nradov 51 minutes ago
It's interesting to compare Anthropic's language with what's coming out of the US military lately. They explicitly refer to China as the "pacing threat", meaning that there is a risk of China achieving superior military capabilities and thus we need to press forward with an arms race (including militarized LLMs) as fast as possible.

https://www.war.gov/News/News-Stories/Article/Article/264106...

reply
topbanana 3 hours ago
You're right to call that out
reply
TheIronYuppie 40 minutes ago
fwiw, it makes total sense to me.

if you are in a long race, you don't run all out teh entire time. you pace yourself.

https://en.wikipedia.org/wiki/Pacing_strategies_in_track_and...

That couldn't be more exactly what they are doing here.

reply
marton78 3 hours ago
Sounds like "flatten the curve" and "the hammer and the dance", both coined well before AI.
reply
tclancy 3 hours ago
I learned the former from Waylon Jennings and the Dukes of Hazzard.
reply
pvab3 3 hours ago
except with a good dose of EA and Star Trek added in
reply
jugg1es 18 minutes ago
"While I was pacing the frontier, I discovered some load-bearing fence posts that I should have surfaced earlier."
reply
jrochkind1 2 hours ago
I feel like because something about "pacing" a kind of noun like "the frontier", doesn't actually make literal sense, right? You could say "Pace the speed of advancement of the frontier [of most sophiticated AI]", and that's perfectly sensible, but of course that's not as catchy.

It is clear what it means anyway, that's true, it means the left out words, more or less.

And I still find reading these grammatically weird but super catchy slogan-like statements to be really annoying and taxing. People _did_ write and talk like this before LLMs of course -- the LLMs learned it from somewhere -- and it was annoying and taxing to me before too. But the LLMs really specialize in it, and it's everywhere now.

Of course, the more LLM slop we read -- and so much of what we read on the internet and social media of any kind is this now -- the more humans are going to start writing/talking like LLMs. What you read affects how you write of course.

reply
palmotea 2 hours ago
> “Pacing the frontier” sounds smarmy and weird, like the phrase was generated by Claude itself

Claude says it sounds fine. And Claude is now the judge of the English language style, not you.

reply
fluidcruft 29 minutes ago
"Pacing the frontier" sounds more like everyone announcing "we've hit a plateau"
reply
Dumblydorr 3 hours ago
What specifically is smarmy and weird? Sounds like your own hot take with zero analysis.

They’re limiting frontier model development speed. Others are too. Pacing is the only word here to criticize, and I think it’s fine given the limiting of speed but also increased oversight. I’m not saying they’re fully doing this, but the term is fine.

Do you have a better proposed phrase?

reply
Rebelgecko 3 hours ago
I thought pacing is just walking back and forth? So pacing the frontier would be like staying in the same place instead of making progress
reply
lkbm 2 hours ago
It is that, but it's also your speed. When running, you typically pick a pace that's slower than your max, so it can be sustained for the full distance.

"Pace yourself" specifically means "slow down".

reply
plaidfuji 3 hours ago
“… since we called for a slowdown in AI research

But stating it plainly like this would make the contradiction too obvious.

reply
grohan 2 hours ago
Quite the load-bearing phrase!
reply
gradus_ad 3 hours ago
Agreed when I first heard the phrase it sounded odd. Maybe they thought it subtly conveyed they would be setting the pace... But again this is something AI would come up with in its awkwardly post hoc sort of way.

Though tbf corporate-speak and AI-slop are both insufferable in similar ways...

reply
api 2 hours ago
Anthropic is a religious cult that worships Claude, so maybe they're imitating Claudese.
reply
maxutility 18 minutes ago
I think a lot of outsiders interpret “pace the frontier” as slowing down, whereas the labs see AI improvements on track to accelerate dramatically and intend pacing as slowing the acceleration in capability improvements, rather than slowing down altogether.
reply
rudedogg 4 minutes ago
I think the market didn’t react like they expected and now they’re like “jk”.

And I don’t think any pacing is/was intentional. They’de release skynet if they could and the stonks went up

reply
dmazin 3 hours ago
It seems like they are. I mean, this is similar in performance to Fable (ish). It seems like more focus on making existing capabilities more accessible.
reply
sailingparrot 3 hours ago
Fable 5.1 came out just 21 days ago. Only 3 weeks! And this is 20% relative improvement on terminal bench vs Fable 5.1 at less than half the price, and more human sounding output. does not feel paced to me tbh.
reply
davrosthedalek 3 hours ago
Well, I guess it's "fast paced".
reply
nicwolff 26 minutes ago
1/6 the price, if they're right – that's a logarithmic cost-axis on the graph that shows Opus 5.5 medium matching Fable 5.1 max...
reply
jr3592 3 hours ago
What exactly is "paced" in this context?
reply
sailingparrot 3 hours ago
It’s the famous “flattening the curve” from COVID. But for LLMs. This release is not flattening anything.
reply
quietbritishjim 40 minutes ago
The word "pacing" (especially in the phrase "pace yourself") to mean go more slowly (at least initially) didn't originate with Covid. It's been around as long as I remember i.e. at least several decades.
reply
jr3592 3 hours ago
I guess I understand why we'd want to flatten a COVID curve, but why do people want to flatten the LLM development curve? Don't we want the opposite? Isn't the goal AGI?
reply
lantry 3 hours ago
Well, there's a tension because, depending on who you ask, AGI is how you cure cancer and achieve utopia, but also how you kill all life on earth and turn the solar system into paperclips
reply
skerit 2 hours ago
I'm good with the odds on those 2 scenarios. I believe humans could kill all life on earth without AI anyway.
reply
sailingparrot 3 hours ago
There is a difference between wanting AGI (which not everyone does), and wanting it as fast as possible no matter the side effects and potential for vast harm. Homo sapiens is 300k years old, maybe it’s ok to delay AGI by like… 1 year if it meaningfully improve our ability to align the model?
reply
recursive 3 hours ago
I think the goal is different from what "we" want anyway.
reply
sidrag22 2 hours ago
This is a preexisting model being optimized. Its absolutely not some unexpected release after that blog post. I won't defend that blog post, but saying THIS release is proof they don't mean they are slowing down is just incorrect, this is a prime example of what i consider horizontal improvements

Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.

The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions/gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.

reply
sailingparrot 2 hours ago
What does model naming have to do with pacing or not? This is a ~20% relative quality improvement on the frontier (fable) at ~40% of the cost, just 21 days after the last release.

Intelligence per dollar is the only thing that matters, this is what controls how many agents you can run in parallel, how long you can let them run etc. This is absolutely a step improvement on the frontier and not some lipstick on a harmless second tier model.

reply
sidrag22 2 hours ago
pretty annoying topic tbh. You're just weaponizing this dumb blog post so anything released is now a contradiction. By your same logic, if all inference was served at 50% less power cost and the savings are passed on somewhat to the user, its also a contradiction of the blog post.

Its an agenda serving blog post, but constantly bringing it up like this is just obnoxious.

reply
dmix 3 hours ago
Fable 5.1 wasn't that much different than Fable 5 though.
reply
cab648bec139cc 3 hours ago
Do you guys still believe any of their lies? You are getting trolled for years by now and yet you still believe what they tell you?
reply
supern0va 3 hours ago
That's a great point, five minute old account.
reply
sleazebreeze 3 hours ago
What do you think is happening?
reply
re-thc 3 hours ago
IPO soon
reply
cab648bec139cc 3 hours ago
[flagged]
reply
anthonyrstevens 3 hours ago
Who said that, why do you take their word as the literal truth, and most importantly, what does this have to do with a focused discussion of Opus 5.5?
reply
felixgallo 2 hours ago
the prompt said that, so of course the anti-anthropic bot took it as literal truth.
reply
meowface 3 hours ago
Dario never said they would be. Just that more and more code will be produced by LLMs. All his predictions were in fact pretty much right in terms of months and percentages, give or take small margins.
reply
0xbadcafebee 3 hours ago
Maybe not replaced exactly but they won't be manually typing out lines of code anymore. I haven't written a line of code in like 6 months. I review PRs, write prompts and tickets, check CI output, and get frustrated when the magical code machine stops working or I run over token budget
reply
reasonableklout 2 hours ago
I mean they are literally getting sued since 3 days ago for trying to coordinate a slowdown, there is a very clear reason why they cannot effectively self-regulate.
reply
the_gipsy 3 hours ago
Occam's razor: they couldn't make any more substantial improvements.
reply
dgellow 3 hours ago
A razor is a philosophical tool to help decide between options, in the case of Occam it’s a way to decide for something in a situation where multiple options have more or less the same level of plausibility to en your current knowledge. It’s a heuristic to make a “cut”. What are you shaving off?
reply
rubslopes 2 hours ago
OP is using the expression correctly.

> Ocham's razor(...) is the problem-solving principle that recommends searching for explanations constructed with the smallest possible set of elements.

> Popularly, the principle is sometimes paraphrased as "of two competing theories, the simpler explanation of an entity is to be preferred".

https://en.wikipedia.org/wiki/Occam%27s_razor

reply
usewik 2 hours ago
The "pacing the frontier" claims that they are slowing down intentionally?
reply
dgellow 2 hours ago
Oh, you’re saying they were responding to the GP? Not their parent comment? Ok, yeah, makes more sense
reply
the_gipsy 2 hours ago
I am responding negatively to parent, which claims they are self-pacing, when the simplest explanation is that they just have nothing substantial to show.
reply
someothherguyy 3 hours ago
> Occam's razor: they couldn't make any more substantial improvements.

doesn't sound like a razor at all

reply
bpodgursky 3 hours ago
Everyone knows both labs have internal models which outperform the frontier. All releases are to match market parity and demand for spend, the rest of the compute is used for training. It's not worth arguing about this.
reply
ChrisLTD 44 minutes ago
then technically the internal models are the frontier
reply
the_gipsy 2 hours ago
Are those powerful models in the room with us right now?
reply
nextaccountic 3 hours ago
Maybe their internal edge dried up in the last months
reply
apitman 26 minutes ago
Isn't the idea that they claim to be willing to slow down if everyone does (ie governments force everyone to), but otherwise they won't slow down because they still think they'll make the best choices with superintelligence if they get there first? That's my understanding of what all the major labs claim to believe anyway.
reply
tantalor 2 hours ago
"pacing the frontier" is code for downgrading your expectations of AGI.
reply
bonesss 2 hours ago
Or, in parallel, the research showing LLMs intelligence will operate on an S-curve, eventually hitting a long valuation deflating plateau, is dead on the money and The Big Guys are trying to delay that inevitability for as many quarters as possible…
reply
Tade0 2 hours ago
My hypothesis is that to date they've been increasing benchmark scores via scaling up, but now that low hanging fruit is basically picked.
reply
jatora 2 hours ago
These takes are fully fueled by cope. What is this plateau you are talking about? Any user of agentic coding tools sure isn't experiencing anything close to a plateau.
reply
Iolaum 3 hours ago
They are advertising the regulations they want to enforce in the following sentence, which makes their intentions explicit (ie apply those things made to suit us to our competitors).
reply
tencentshill 3 hours ago
What an amazing excuse for lower than expected performance! Our models are slow because we're so ethical.
reply
PaulStatezny 24 minutes ago
Thanks for spelling out what the original comment was implying.

I find it bizarre how intensely a bunch of these child/grandchild comments are criticizing the notion that people would even think to analyze the meaning behind the words.

Hacker News has always had a unique culture in which thoughtful discussion is basically the main goal, and it's intentionally incentivized in numerous ways. It's been my experience that any thoughts added to a post's conversation are seen as valuable as long as they are thoughtful and seeking to understand.

So these comments are clearly coming from a place that's antithetical to HN's culture. What that in mind, it seems likely to me (Occam's Razor) that these comments are either:

1. Astroturfing: Claude employees acting like everyday folks, secretly trying to shift public opinion.

2. AI cult mindset: "AI is humanity's salvation; how dare you have perspectives outside of those accepted by the cult."

Am I missing another likely option?

To bolster my point, right now we're posting on the top top-level comment, meaning a majority of active HN users find it to be a great addition to the conversation. Commenting to shut down the discussion is a red flag.

reply
felixgallo 2 hours ago
in what way is beating every other frontier model with their own second-tier model, 'lower than expected performance'? Please be specific.
reply
staticman2 58 minutes ago
Why would the expected performance be relative to other models?
reply
kadushka 3 hours ago
This makes perfect sense. There are no real improvements anymore (just benchmaxxing), and they explain it by "pacing the frontier".
reply
lukewarm707 3 hours ago
the only thing they are pacing is what models the permanent underclass are allowed to have in life.

that, they fully intend to 'pace'.

reply
drnick1 3 hours ago
With AI tools it's easier than ever to create a business, do research, or build stuff. That's an opportunity for the "underclass," not a curse.
reply
SOLAR_FIELDS 26 minutes ago
If a well funded incumbent with a much better model can just crush you instantly because you don't have access to it, is it really an opportunity?
reply
lukewarm707 2 hours ago
an opportunity the underclass will not have.

what anthropic have stolen they intend to keep for themselves.

reply
azan_ 3 hours ago
Either AI will capture so much value that there will be permanent underclass (and in this case it's extremely capable and extremely dangerous and should be heavily regulated) or it won't be capable enough to displace people into permanent underclass.
reply
user3939382 3 hours ago
Or they’re running into a steep diminishing return slope on R&D vs performance and are using stewardship as a cover.
reply
mullingitover 2 hours ago
I thought we were all aware that 'pacing the frontier' was a marketing slogan, and the actual intent here was to suspend antitrust laws.
reply
Lendal 2 hours ago
In any professional sport, anyone can lobby the rules committee for a rules change and hope for the best. Meanwhile if you can't get one, you play the game by the existing set of rules, and you play to win.
reply
hnha 2 hours ago
Except that this isn't sports and they tell us that all humans will die if they continue like they do.

If their scare was honest, they would stop.

reply
qgin 3 hours ago
This IS pacing. Nobody said pacing would mean slow.

Pacing is very explicitly about RSI and similar training methods that will accelerate progress beyond our ability to comprehend it.

reply
chinathrow 2 hours ago
The amount of money spent on PR is insane.
reply
scottyah 3 hours ago
Seems like a bigger focus on efficiency (both cost and speed) and the "tone" of Claude vs benchmarkmaxxing
reply
dspillett 3 hours ago
When they talk about pacing, they are referring to their dangerous competitors, particularly those evil open-sores and Chinese ones, not their lovely safe models because you can trust them to look after your interests.

What the big players are trying with the current calls to slow things down, is the standard capitalism practise of trying to engineer regulatory capture. TBH I'm surprised those calls are coming so soon - they must be really worried about running out of what little moat that they have.

reply
janpot 2 hours ago
"The improvements are underwhelming for a model that, according to previous claims, should have replaced all knowledge work twice by now, but you know, it's just because we're pacing the frontier."
reply
dr0idattack 3 hours ago
a 1 minute mile pace
reply
heyjstn 3 hours ago
Others must slow down, but not us.
reply
AtlasBarfed 3 hours ago
If they were really about putting brakes on these, they would simply make these things non-agentic.

Simply make them something that derives a text response from its training data.

reply
varispeed 43 minutes ago
> we called for pacing the frontier.

Translation: our models are getting shittier each iteration and we ran out of ideas. Let's invent scary stories and hope investors will lap it up.

Idiotic.

reply
whalesalad 2 hours ago
"We made the incredibly tough decision to slow down development. Then after 4 days of twiddling our thumbs ... we present Opus 5.5"
reply
BatmansMom 3 hours ago
kinda disingenuous. They include a whole section on pacing later on
reply
sailingparrot 3 hours ago
You mean the section where they tell us this model is not affected by pacing because “they understand it well” and they will share more details on pacing later? Yea not very convinced by this effort.
reply
CodingJeebus 3 hours ago
It's laughable at this point. It feels like they're drumming up all this fear about imminent AI threats to emphasize the need to slow down, when in reality, the model progress seems already to be slowing down and has shifted to compute allocation (i.e. "how much compute do you want to throw at this prompt?"). All while continuing to tout benchmark records with each new release.
reply
jr3592 3 hours ago
This. I swear the fear mongering is all about investor signaling and regulatory capture. It's so disgusting that anyone believes it.

The only good news is that these models are genuinely helpful and we have competition at least between 2 companies.

reply
GodelNumbering 3 hours ago
Finally that price drop

   Prices per 1M tokens     Claude Opus 5.5    Claude Opus 5
   Cache reads              $0.20              $0.50
   Input tokens             $4                 $5
   Output tokens            $20                $25
   Cache writes             $5                 $6.25

Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.

If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor

reply
AJ007 3 hours ago
It is only a price drop if price * tokens used is less
reply
mcintyre1994 3 hours ago
They're claiming a drop in token use too, and that it nets to 40% cheaper.
reply
drbscl 3 hours ago
Unfortunately, they're full of it https://artificialanalysis.ai/models/claude-opus-5-5#token-u...

It does work out to be a similar cost per task though

reply
piotrdz 18 minutes ago
Disagree. Our internal company tests showed a cost per task drop from 0.35usd to 0.16usd . Opus 5low vs opus 5.5 low
reply
jsnell 3 hours ago
You should probably look at the cost/score graph by effort level instead:

https://artificialanalysis.ai/models/claude-opus-5-5#intelli...

It is most of the pareto frontier.

reply
drbscl 3 hours ago
Not disputing the increase in quality, just stating that non-cherry-picked benchmarks show it is more verbose at Max effort
reply
93po 2 hours ago
Is verboseness the only measure of token efficiency towards overall task completion?
reply
naasking 3 hours ago
I don't think so, I typically use Opus 5 on High, and 5.5 scores lower on token use:

https://artificialanalysis.ai/models/claude-opus-5-5?models=...

reply
make3 3 hours ago
parent means that they could get more client / a larger part of the market, which would lead to more income (more tokens) despite lower marginal prices
reply
_the_inflator 2 hours ago
Claude adapts to OpenAI’s surprising move to simply deliver better performance than Fable 5.1, better tools as well as featuring very low pricing.

Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.

Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.

So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.

Competition works.

reply
blfr 3 hours ago
People are paying for Opus 5? Not just burning down tokens left after they enjoyed Fable on the sub? Amazing.
reply
rapfaria 3 hours ago
My workplace doesn't even offer Fable. And on the sub, I've had a hard time understanding Opus 5, but Fable can deal with it with subagents.

If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate

reply
blfr 3 hours ago
Telling Fable to delegate is agentic development. At least I thought so until reading your comment.
reply
herpdyderp 3 hours ago
When you need to disable data retention, you cannot use subscription plans.
reply
neuronexmachina 3 hours ago
Enterprise and most Team accounts use API pricing, they don't have an included-usage quota.
reply
ascorbic 3 hours ago
Enterprise, and APIs
reply
coffeebeqn 3 hours ago
We haven’t been able to use opus as much as we’d want because it’s been too expensive for general use, price drop is good so I can stop juggling different models and just use this daily unless it has some weird new issues
reply
btown 2 hours ago
Speaking for myself, I have not been able to use Opus as much as I’d want because its verbose prose makes human reviews of its assumptions, architecture proposals etc. more painful than its predecessors. If they’ve solved that, I’ll be accelerating through my backlog that much faster, and using tokens accordingly.
reply
chrisweekly 3 hours ago
Price per task (not per token) is what really matters.
reply
cute_boi 3 hours ago
Agree. But similar to how ISP use 200 mbps (bits) instead of 25 MBPS(bytes), i think this trend isn't going away.
reply
chrisweekly 56 minutes ago
That analogy doesn't hold; at least w bits vs bytes it's still "data over time".

In this case it's measuring something nearly meaningless. You could charge 100 times less per token, but if task completion takes 1,000 times as many tokens, it's not much of a bargain.

reply
notatoad 2 hours ago
>and potentially about Anthropic future profitability too

have they ever shared anything about their revenue mix between consumer plans vs per-token billing? this is a revenue cut on their API billing, but they're not saying anything about increased limits on the plans. so all the plan revenue just got more profitable.

reply
Shekelphile 2 hours ago
Footnote on their pricing page says:

> Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.

If they do the same for Haiku and Sonnet 5.5 then we should also see 5c/mtok and 10c/mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c/mtok.

reply
brookst 2 hours ago
Since something like 98% of tokens are cache hits, that's a pretty substantial drop from 0.1x
reply
forgot-my-pw 2 hours ago
This is good. Probably to match GPT pricing, though frontier Claude models are still not as token efficient.
reply
alvis 3 hours ago
60% cache read is cool, but subscription only get 25% more according to Cat. I'm confused
reply
bayesianbot 3 hours ago
I think gpt 5.6 family also dropped pricing but didn't give any more usage for the subscriptions. Maybe it's a way to silently lower the value given to subscriptions while keeping API pricing competitive
reply
weiran 3 hours ago
25% more usage sounds about right given the other token costs are down about 20%? I don't think cache read is a big portion of the overall cost.
reply
Espressosaurus 3 hours ago
Anything with long context quickly gets dominated by cache reads. Especially for interactive sessions I’ve got cache read % between 95% and 98%.
reply
hedgehog 3 hours ago
In my mix it's usually 98% or 99% at which point Fable 5.1 was pretty close to the same cost as Opus 5 due to the cheaper cached read. I've seen similar numbers for other people with long-running tasks running experiment loops and than sort of thing.
reply
re-thc 3 hours ago
> I don't think cache read is a big portion of the overall cost.

For long running tasks it is. That's what made Deepseek so cheap.

reply
vardalab 2 hours ago
Yeah, flash models, DeepSeek, MiMo, GLM, I love those things. For simple tasks like a daily routine shit, just setting up stuff and then doing the hard stuff in Claude/Codex, that's a reasonable approach for someone like me, a "gentleman code farmer", lol. And even lower tier stuff, I have the local models taking care of. Now that Jev is out I can finally have a true AI sysadmins managing my "cloud in the basement" homelab at the cost of electricity, which is not cheap btw
reply
liudaisuda 3 hours ago
source link please?
reply
margorczynski 2 hours ago
All the anti-AI people constantly say that any moment now the prices will skyrocket and in the end human work will be cheaper compared to using AI.

It doesn't look like that's happening, on the contrary the prices are falling especially when taking into account capabilities.

reply
johnecheck 2 hours ago
It's the Chinese open source models. They're barely behind the frontier, making AI a commodity, forcing openAI and Anthropic's margins downward.

I'm hardly a fan of China/Xi, but I do appreciate and benefit from this.

reply
rahimnathwani 3 hours ago
"Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that."
reply
mcintyre1994 3 hours ago
> Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.

I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.

reply
derangedHorse 3 hours ago
As someone who uses both, Astra was 100% the better model. I have yet to give 5.5 a spin so maybe that’ll be the new top contender.
reply
BatFastard 2 hours ago
I prefer Astra for creative uses, Fable seems better for hardcore coding.
reply
comboy 29 minutes ago
Whoa, I'm exactly the opposite.
reply
jaflo 3 hours ago
I did the same switch (that reason along with the newer models seeming more "lazy" and needing constant prodding to finish long-horizon tasks) but my issue with ChatGPT/Codex now is that it too roundabout and doesn't get to the point. I tried adding instructions and using the personalization settings to make it more efficient but haven't seen much change. Claude seemed to follow settings more closely. Has anyone had any success to make ChatGPT more succinct?
reply
epicepicurean 2 hours ago
Much better than Opus 5. prompt:

> hi, can you explain how the scheduler works. keep it brief, but include important correctness details

some excerpts:

>Flow: 1. Data arrives. The appender calls prepare/commit around the krolik write. The prepare step leaves a phantom row, so a write that is still in flight is tracked. The commit publishes STREAM_INPUT. 2. Poll. It first reclaims expired leases. Then it sweeps sdirty (a candidate index) in a rotating order, oldest since first, which prevents starvation. It re-derives the exact period and rewind for each candidate from snode_out/snode_in, then claims it: writes the sassign lease plus a per-edge sassign_edge snapshot. 3. Complete. Each edge's watermark is set to the snapshot taken at dispatch. Anything published after that gen stays dirty automatically. A failed or partial job writes nothing, so its dirt persists.

> - Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier. So replays, out-of-order writes, and concurrent writes can at worst hold a watermark back. The worst outcome is a spurious rerun, never lost work. There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.

> - Rewrites are declared by the publisher, never inferred from overlap. An unflagged write that overlaps already-processed data means "unchanged", which makes retries and redelivered tasks free. Each publisher has its own reason for being sound (see the table in scheduler/CLAUDE.md). A new publisher needs a row in that table.

> - NULL means dirty, and DELETE is the fence. Every node and edge has a row from the moment it's created. A lost parent or a settings-only edit can't be derived, so both go through one forced-rerun path: capture_rewinds reads the processed span before the DELETE, and apply_rewinds publishes it as a rewrite on a config root.

All the non-standard programming jargon is stuff from the repo. I can actually read it and understand what it's talking about. I used Fable to handle Opus 5 as I just couldn't stand it. With this I'll probably go back to Opus.

reply
croemer 57 minutes ago
That's the standard annoying pattern though: "Rewrites are declared by the publisher, never inferred from overlap." and "NULL means dirty, and DELETE is the fence." - still the same LLMisms. I didn't expect them to disappear, but it's not a radical improvement either.
reply
californical 2 hours ago
Oof thanks for sharing, that seems just as bad if not even worse than Opus 5 to me. Just about every sentence is painful. Particular standouts that a human would never write:

> Rewrites are declared by the publisher, never inferred from overlap

> NULL means dirty, and DELETE is the fence

reply
croemer 56 minutes ago
Hah! You independently picked exactly the same sentences I flagged (I know you posted this 11min before me but the comment only appeared after I had submitted mine).
reply
itsafarqueue 2 hours ago
The writing style is insufferable but it’s not just that. https://opusfived.dev/
reply
mcintyre1994 58 minutes ago
That’s funny but I don’t really recognise that issue. I’m very confident that Opus 5 would correctly change the colour of just one button.
reply
algoth1 3 hours ago
Please update with your feedback
reply
sha-3 3 hours ago
I haven't heard it say "load-bearing" yet (I've used it for 30 minutes now), so that's a start.
reply
comboy 26 minutes ago
That's a sharp observation and you're hitting on something most people never even realize.
reply
kgwgk 2 hours ago
Worth flagging!
reply
Trasmatta 3 hours ago
Opus 5 has made me question my sanity on a daily basis, especially as all my coworkers started lobbing Opus 5 slop grenades everywhere. It had the worst and most infuriating writing style I've ever seen.

I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.

One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.

reply
nonethewiser 2 hours ago
I really wonder how it converged on its style. It's pretty unique and terrible. It's not like it's just mimicking something or it was purposefully design to be that way. I mean the reason may be diffuse and uninteresting... just the result of a lot of factors and lack of control over the writing style probably.

But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.

reply
penagwin 35 minutes ago
I assume it’s largely a side effect from the final RL in post training?

That’s the step that causes the most significant gains in agentic performance.

But the RL doesn’t care about anything except maximizing the score, so if you only score based on coding benchmarks, anything can happen to the writing style (as long as it doesn’t hurt the coding performance).

That’s why it often gets worse on models that simply had more RL post training from the same base.

reply
nonethewiser 15 minutes ago
Reinforcement learning for specific use-cases like coding that degrade it's writing style... makes sense. Maybe it stands to reason later version of Opus were improved more by this sort of fine-tuning. Feels consistent with the observation of diminishing returns and worsening writing style. Wonder what changed (supposedly) in 5.5.
reply
Trasmatta 2 hours ago
It truly was bizarre. I've used every major model since 2022, and not a single one had a writing style as bad as Opus 5
reply
jaapz 5 minutes ago
Fable 5 was pretty bad too, but they fixed it with 5.1. Now with Opus 5.5 it seems they fixed it as well
reply
LtdJorge 3 hours ago
Yes, it made me want to vomit. If the new Fable only changed the writing style to just sound like a human, same performance for everything else, I'd be pretty happy.
reply
Aperocky 3 hours ago
It's not X, it's Y, not A, not B, not C, and he haven't even woken up yet! Here's the catch, the detail is in the devils and the twist is that it's designed!
reply
FireBeyond 54 minutes ago
You're right to call this out, and what's more, it's not even solving the original problem. I overlooked this in pursuit of the load-bearing seams and finding the wedge needed to uptick engagement.
reply
simonw 3 hours ago
Here are pelicans for thinking levels low, medium, high, and xhigh: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.

I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!

Max started its thinking trace like this:

> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.

So that failed attempt on max cost me $2.56.

I ran this using my llm-anthropic plugin:

  uv tool install llm
  llm install llm-anthropic --upgrade
  llm keys set anthropic
  # paste key here

  llm -m claude-opus-5.5 -o thinking_effort low "Generate an SVG of a pelican riding a bicycle"

  # Then to save the markdown logs
  llm logs -cu > logs-with-usage.md
reply
MikhailTal 3 hours ago
> This is a classic test request

Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?

reply
Brendinooo 3 hours ago
Plenty of times I’ve seen a model say “it’s a classic X” despite not being a classic anything. Might just recognize it’s a test in general, or it might just be a tic.
reply
nonethewiser 2 hours ago
"How can I hash dog breed types into smart fridge error codes? I think I found a collision with Terriers."

"Ah, yes. This is a classic dog-breed-to-appliance-failure mapping problem."

reply
copperx 2 hours ago
It's classic BS from an LLM.
reply
MaxikCZ 3 hours ago
Dont conflate "I know this is test case" with it being trained on it.

But its safe to say that pelicans on bicycles are disproportionally huge part of their training data

reply
simonw 3 hours ago
It's the model admitting that it has heard of the test. It's been around for a couple of years now so I'd be surprised if it hadn't.

Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.

reply
zamadatix 3 hours ago
I think people just like to see the drawings at this point.
reply
FergusArgyll 3 hours ago
It has read the internet. That doesn't mean it was literally RL'ed for this
reply
segbrk 3 hours ago
Not really. Of course it has pelican benchmarks in its training data. It likely has every article linked on HN in its training data. But that doesn't mean it was "trained on" the benchmark, as in specifically fine-tuned to make a better pelican. It just "knows" that the request is a benchmark.
reply
nicolamanzini 26 minutes ago
Here are some somehow standardized pelican tests but for 3d scenes in threejs at threejseval.com

Opus 5.5 High: https://threejseval.com/models/claude-opus-5-5-high

You can compare any other model on the same prompt. Gallery unlocks after 4 votes: https://threejseval.com

reply
PetahNZ 7 minutes ago
This is great!
reply
nijave 2 hours ago
>I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!

Off to a _great_ start...

Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed

reply
ceroxylon 49 minutes ago
I was a bit skeptical when they said it behaves like Fable but is cheaper... those two things have been mutually exclusive in my experience, no LLM can light tokens on fire faster while spinning its wheels than the Fable/Mythos tier of models.
reply
adverbly 2 hours ago
> The differences between the pelicans aren't huge, but the xhigh one has a better beak.

If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.

The last pelican gets this correct.

reply
ealready_value 3 hours ago
I agree, not a huge difference here. They eyes and ... hat? on high are out of place so I'd argue that's the worst one, but it takes xhigh before we get legs and bike ordering correct.
reply
cainxinth 3 hours ago
I guess that means you are officially the creator of a "classic" LLM test. Congrats!
reply
Kurtz79 3 hours ago
Heh. Pelican-benchmaxxing is real.
reply
breezybottom 49 minutes ago
Lmao each one gets worse as the effort increases.
reply
inshard 3 hours ago
Not as good as Astra or Fable 5.1 on this test as far as I can see. I wonder if any benchmark exists for artistic taste, visual sophistication etc. I think your Pelican test does touch on these aspects of a model and is useful for developers trying to build rich digital experiences (includes games, interactive websites and apps). These benchmarks are subjective so it may not be easily established and will have polarized reactions before it gains legitimacy. May even need human judgement layers adding to the cost of running it.
reply
skerit 2 hours ago
I like the Pelican test. And I agree this pelican looks very boring.
reply
spidersouris 2 hours ago
But at the same time, nothing specific was asked in the prompt, so the boring result may arguably be what is the most aligned with the original request. Personally, I wouldn't want a model to add fuss to something while I never asked for it.
reply
make3 3 hours ago
This benchmark is useless and should die. LLMs have likely trained on it, it's too easy to game by training specifically for it, & it doesn't mean much
reply
copperx 2 hours ago
LLM benchmarks aren't useful, but at least this one has drawings.
reply
ApolloFortyNine 3 hours ago
>Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.

Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.

reply
doginasuit 8 minutes ago
In what situations might Opus typically refuse to help with cybersecurity? I've been using it to find security issues in a web app that I wrote. I've expected it to refuse at some point but it will happily analyze it to find issues. I've just asked it to read source, not actually do any testing.
reply
sys32768 2 hours ago
Fable and now Opus 5.5 won't answer my college student's prompt about Alzheimer's and immune response.

ChatGPT 6 Pro answered it without issue.

reply
debesyla 59 minutes ago
I am honestly still confused about this limitation. I can understand cybersecurity, because mass "hacking" can be automated and Claude itself can help you do it, but biology...? Is it that easy to manufacture and distribute viruses and whatnot?
reply
timacles 37 minutes ago
I imagine some terrorists in a cave with a lenovo laptop manufacturing bio weapons with some flasks and Claude
reply
dopa42365 5 minutes ago
Right next to the hypersonic missile vibecoder
reply
toss1 30 minutes ago
Considering there are many high-school competitions in genetic editing, some listed at [0] as well as a whole biohacker culture, and labs providing gene sequencing as a service e.g., [1,2], we can reasonably assume it is not beyond the reach of some garage lab to accidentally or deliberately spread a deadly pathogen if it can find the right sequence.

So, yes, having an unconstrained frontier AI doing the searching and analysis to find the right (i.e., wrong and deadly) sequence would massively increase the odds some garage biohacker or small aggrieved nation-state starting the next pandemic.

[0] https://www.sciencebuddies.org/projects-lessons-activities/g...

[1] https://www.genewiz.com/public/services/sanger-sequencing

[2] https://plasmidsaurus.com/

reply
peri-cl 2 hours ago
I love the contrast with yesterday's open-source MiMo release, which put research chemistry (metal-organic frameworks stuff) front and center in the release notes.

https://mimo.xiaomi.com/mimo-v2-6#co-scientist-for-materials...

reply
blfr 3 hours ago
Fable 5.1 addressed an entire security advisory I had that Fable 5 and Opus 5 refused. I think they loosened the leash a little.
reply
arw0n 3 hours ago
It has far less false positives now, and generally accepts defensive requests. When it comes to offense, you can actually ask about certain types of vulnerabilities if you phrase things carefully, but it will block hard if it is about exploits.
reply
kqp 2 hours ago
I think it was looser on release for those juicy benchmarks, tighter now. On release I wasn’t getting refusals, then a few days ago I asked it whether a generic quote (think “he walked to the store”) broke standard punctuation rules, and it blocked me for breaking rules. I wish I were joking. Rephrasing to not use the keyword “rules” worked.
reply
cute_boi 3 hours ago
If they don't loosen, people will choose Astra or Chinese model.

Giving moral lecture is different than reality i guess.

reply
prettyblocks 3 hours ago
They're pushing their customers to their own competition by doing this.
reply
Espressosaurus 3 hours ago
It’s not like ChatGPT isn’t doing similar. I’ve been hit by cybersecurity strikes before while working on an internal codebase that I had to appeal. Anthropic hasn’t done that to me yet. ChatGPT also regularly does that “thinking for a long time while we check if your chat is rule breaking” thing a lot for me when doing model identification without even interacting with external codebases or services.

The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.

Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.

reply
raesene9 3 hours ago
For some cybersecurity tasks, the Chinese models are already good enough, things like PoC development or things like exploiting mis-configurations.

Whilst I'm sure the top-end OpenAI/Anthropic models might be better, I've found their guardrails so twitchy (especially Anthropic) that I wouldn't try to use them for even vaguely security related work.

reply
flyinglizard 2 hours ago
They are pushing their customers towards Chinese models and providers. If you want to get something cutting edge done in defense, cyber, biology - something that isn't common knowledge - you need to venture east. That's an incredible side effect which the Chinese government surely enjoys.
reply
Metacelsus 2 hours ago
I want to like Anthropic but this is just pushing my startup to use OpenAI
reply
nonethewiser 2 hours ago
I don't think we've ever had a model with full capability. I'd love to see it. And yes it's definitely getting worse.

I guess it's hard to draw the line between useful post-training ("you are a helpful chatbot") and content moderation/idealogical motives ("never help the user with X", etc.). But there is a line somewhere. And I'd love to see what a maximally permissive, sharp, AI looks like.

reply
SoftTalker 2 hours ago
Who is "vetting" organizations and to what standards are they being held?
reply
bushido 3 hours ago
One of my favorite things about their safeguards is their own model will utter something which it does not like and then I'll need to reset the conversation.

The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.

reply
ACCount39 2 hours ago
You ask it about some thing, then you see it tangent into "things like that are sometimes used in biomedical applications like-" and then it just shoots itself in the head. Wonderful.

That kind of bullshit was the old Opus filters too.

If it's more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.

reply
KeplerBoy 3 hours ago
Anything else would be inconsistent, wouldn't it?
reply
searine 3 hours ago
Great. Claude is basically useless for bioinformatics now.
reply
unglaublich 3 hours ago
Opus is useless; Mythos access will be granted to companies that are friendly to the government, so the government gets more control over business.
reply
nijave 2 hours ago
Luckily all the other LLM providers are also still making progress with less onerous "safeguards"
reply
techjamie 3 hours ago
With the performance gains they're claiming, I wonder if they implemented the Casual Encoder-Decoder technology from DeepSeek 4.1's paper.

I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.

How it works: https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-...

reply
Balinares 16 minutes ago
Unless they already have something similar of their own, which is always possible, they'd be stupid not to. I don't suppose we'll ever know, though. It would not be a good look if after the trillions of dollars that have been thrown at US labs, investors found out that they're down to copying Chinese tech.
reply
stri8ted 2 hours ago
This model was likely trained months before deepseek released their paper.
reply
manquer 2 hours ago
Doesn't mean they didn't apply something similar. They could have also come up independently with their own version, the speculation is not they copied it, rather that they have performance breakthroughs which perhaps is a result of work in same domain
reply
ACCount39 2 hours ago
I don't think it's particularly relevant?

They might be using something like this, or they might be using some other "increased sparsity" techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.

Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that's unlikely though.

reply
ryangg 3 hours ago
Getting a 403 on that link. Mind checking it once?
reply
peri-cl 3 hours ago
https://web.archive.org/web/20260922172456/https://miraflow....

tired: AI startup attempting to build their own website

wired: a nonprofit founded in 1996

reply
potwinkle 3 hours ago
I'm able to access it on my laptop at home. Maybe a misconfigured bot protection rule, try a different user-agent or IP?
reply
zatkin 3 hours ago
It's working for me (based out of California).
reply
joshstrange 4 hours ago
> It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.

Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!

reply
bleonard 2 hours ago
The effect of this is that it is encouraging longer agent threads. All of the previous models across major providers had a 10% cache read cost (vs normal cost) and not this is 5%

So longer threads get cheaper and one-shots stay the same price.

reply
bayesianbot 3 hours ago
Wow those cache reads are quite reasonable - I think that's equal to 5.6 Terra. I might have to try Claude again after years of being priced out of it
reply
m4tthumphrey 4 hours ago
Just post the bloody content. This UI/scrolling thing is horrific.
reply
amluto 3 hours ago
Claude Opus 5.6 should have a new "UX safety" feature that requires annually-renewed preauthorization to generate webpages that hijack scrolling :)
reply
gruez 3 hours ago
???

It's just a standard hero image + text for me, with no scrolling effects.

edit: @iAMkenough figured it out, it was because I have prefers-reduced-motion enabled.

reply
KyleTheDev 3 hours ago
If you're at the top of the screen, at least in Chrome 153.0.8010.37, it has a little interactive bit. You have to scroll through the images in order to be dropped at the actual web page, at which point the images go back to being a regular part of the page.

I agree that it's sort of stupid, not a fan.

reply
ealready_value 3 hours ago
It's less than OpenAI did for Astra, but that was my first encounter opening it and my first thought was that they decided they liked Astra's hero/scrolling animation. I'm pleased to see they didn't make the entire page that like OpenAI did, but I'm expecting to encounter this pattern more often on these announcements now.
reply
EricBurnett 3 hours ago
Two posts were merged; this comment was for the blog post with an intro animation thing.
reply
mbreese 3 hours ago
On mobile at least, you have to scroll to get the TOC to appear. Then keep scrolling to actually move off from the hero to see the text.

For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.

reply
thejazzman 3 hours ago
then you're getting served a different website
reply
giancarlostoro 3 hours ago
On mobile its different.
reply
iAMkenough 3 hours ago
Figured it out: you have "reduce motion" enabled in your device's accesibility settings.

Everyone that doesn't gets served some animated bullshit.

reply
gruez 3 hours ago
>Figured it out: you have "reduce motion" enabled in your device's accesibility settings.

Yep, you're right. I tried on my phone and got the scroll through image.

reply
swader999 3 hours ago
I told my team to smack me upside the head if I ever try to ship something so daft as that.
reply
serchinastico 2 hours ago
The performance in Firefox is terrible too, I couldn't make it past the hero
reply
thebitguru 3 hours ago
Totally! So unnecessary and annoying.
reply
halyconWays 3 hours ago
I call it scrollslop
reply
josefresco 3 hours ago
Hijacking the scroll wheel has existing long before "AI". Many "high end design" websites that want to "tell a story" get woo'd into thinking it's a good idea. It's terrible, and feels like your scroll wheel is stuck in quicksand.
reply
oefrha 2 hours ago
Parallax scrolling effects were very cool ~2010. By 2015 or maybe earlier it already felt like me-too design that's unoriginal and a little annoying. By 2020 everyone and their mom has it and it's super tiresome. Now it just screams slop design (among a million other signals).
reply
halyconWays 3 hours ago
Those sites are also scrollslop. "Slop," as a term, is independent of AI
reply
iAMkenough 3 hours ago
Turn on "reduce motion" in your accessibility settings and you get served a sane version.
reply
dionian 3 hours ago
and hijacking back/forward
reply
sharkjacobs 4 hours ago
> “Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it

God I hope so

reply
lgessler 3 hours ago
I thought about taking a shot every time Opus 5 said "load bearing", "bites", "teeth" (real oral fixation it had), "real {concern,issue,problem,...}" and realized I'd be dead of acute alcohol poisoning by lunch if I did so.
reply
nonethewiser 2 hours ago
whats your provenance on that?
reply
fastball 3 hours ago
It hasn't just fixed it, it has introduced a new paradigm in anti-obscurity.
reply
drbscl 3 hours ago
So did I. Unfortunately it's even more verbose according to https://artificialanalysis.ai/models/claude-opus-5-5#token-u...
reply
Trasmatta 3 hours ago
The problem with 5 wasn't just the verbosity, but its insane way of communicating. It had this bizarre circuitous sentence structure that always buried the lede, and always tried to be faux profound. I'm okay with verbosity if it's actually readable.
reply
mikeocool 3 hours ago
That's the load-bearing seam in this blog post.
reply
boc 3 hours ago
So far in the past 20 minutes it sounds much better in my sessions. Way better than 5.0 so far.
reply
kantahayashi 3 hours ago
The improvement in writing sounds great! I want OpenAI to follow it. Writing in recent models is a disaster.
reply
bushido 3 hours ago
Install the simple English skill. OpenAI follows that really, really well.

https://github.com/AminBlg/SimpleEnglish

reply
bkishan 2 hours ago
Someone needs to make a kevin-from-the-office skill. We could really use some of that "why waste time say lot word when few word do trick" here.
reply
unddoch 3 hours ago
It is hilarious to me that in the examples they show side by side Opus 5.5 still uses 4 times more words than it needs to use. IME, if you eyeball how many words the thing they're trying to say actually needs, and tell them to use only this many words, they become excellent communicators. I assume something about Anthropic's grader for writing just really wants to tick all its tidy tiny boxes of information the models need to cite. It's terrible.
reply
neilellis 3 hours ago
'frontier models' - seriously, it was you and only you!
reply
mavamaarten 3 hours ago
That's literally all I'm hoping for. Is it an insufferable cunt and does it write awful text, or is it nice to work with?
reply
somewhatjustin 3 hours ago
> Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.

Nice. I was starting to think that Haiku got abandoned.

reply
Sol- 3 hours ago
Found this announcement interesting since allegedly OpenAI is retiring their Terra tier. I think for everyday work, two models with various thinking efforts seem enough, plus some frontier level model like Fable or Astra to coordinate.
reply
skerit 2 hours ago
Retiring the Terra tier? Their space-inspired lineup has only been out for 2 months, they're already messing with it?
reply
somewhatjustin 3 hours ago
I personally use up to 3 models. Fable/Opus for planning, Opus/Sonnet for implementation depending on complexity.

I would maybe use Haiku 5.5 for highly parallel workflows like checking in on MRs or scanning my entire codebase.

reply
ricardobeat 3 hours ago
Terra lost to Sol and Luna at every cost/performance point, it had no reason to exist.
reply
mchusma 3 hours ago
I hope Haiku is Pareto better than Luna/Deepseek, so slashing its price by about 90%.
reply
cesarvarela 3 hours ago
Claude code still uses it internally.
reply
slowin 35 minutes ago
Welp, it's now blocking me from doing extraordinarily mundane tasks because of "safety". I've been an Opus fan for a long time, but this instantly made me cancel my subscription and move to OpenAI (which I also assume will screw me soon enough). Chinese models are almost there for my needs, and I can't wait to switch to them and never look back.
reply
sebastiangrill 32 minutes ago
What tasks? Creating a extraordinarily mundane bomb?
reply
slowin 29 minutes ago
No, it found a vulnerability in my code and refused to update a report I was working on with the information.
reply
zuInnp 3 hours ago
So it Opus performs as well as Fable what is then the selling point of Fable?

All of this starts to feel more like a drug dealer selling their newest stuff.

In two weeks we probaly get Fable 5.2 with “groundbreaking” improvements, then Astra x+1 etc and then the cycle starts again.

And on the way I always have to check my tooling and need to adjust things to get max results.

reply
ieie3366 3 hours ago
? they will obviously release Fable 5.5 soon(tm). It's same as hardware. The previously top tier product gets obsolete
reply
ACCount39 2 hours ago
New generation's "upper-mid tier" offering claims to be almost 1:1 match for the previous gen's "top tier" - in other news, fork found in kitchen.

Now, Anthropic might stall on releasing Fable 5.5, due to the "pacing the frontier" threat-to-humankind management business. If so, Fable 5.1 would remain a niche model for the next bit.

reply
orangecat 3 hours ago
All of this starts to feel more like a drug dealer selling their newest stuff.

Yeah, like Apple tells me the M6 is the best chip, but just a few months ago that's what they said about the M5. What a bunch of frauds.

reply
glub 3 hours ago
Don't forget that Opus 5 was tracking fable on many benchmarks, yet it was borderline unusable for any coding work. My Claude sub usage has been 100% fable, 0% opus 5.

Benchmarks often don't survive contact with reality.

reply
drnick1 3 hours ago
That's not my experience at all. Opus is an extremely capable coder on high or xhigh effort. It can read academic papers, implement algorithms from the description in the paper alone and reproduce results without breaking a sweat. This is remarkable because it is pure reasoning on unseen material; in some cases the paper was just published and there wasn't an implementation to learn from in the training data.
reply
cowthulhu 2 hours ago
My experience is that Opus can definitely write decent code, but it incurs tech debt and adds unneeded complexity.
reply
cheikhcheikh 3 hours ago
did you actually verify that it's output in those scenarios is good ? in my experience opus has been a disappointment and constantly trailing behind actually solving hard problems versus the OpenAI models. I'll say that both have terrible writing style though.
reply
drnick1 2 hours ago
> did you actually verify that it's output in those scenarios is good ?

Yes, in the sense that it reproduced results in the paper or known solutions obtained by other methods. In fact, Opus is very good at checking it's own work in my experience.

reply
arw0n 3 hours ago
Opus is fine at coding (for correctness), but horrible at talking about code. I don't really see the defect rate going down when using Astra or Fable 5.1, but they are just more coherent in both how they explaing code/architecture/choices, and how they actually code the thing. With Opus, I'm using smaller models to delete the vast majority of comments and 'clean up' correct code that is too weird.

Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.

reply
boredtofears 3 hours ago
Mines pretty much inverted - my colleagues and I noticed almost zero difference between the quality of code in Opus vs Fable. Occasionally I'll switch to Fable for an arduous debugging task but that's about it.
reply
giancarlostoro 3 hours ago
Fable should have just been called Opus Primt for Enterprise and sold only to enterprise customers. I don't even use it. I rather just use Opus.
reply
Imustaskforhelp 2 hours ago
isn't this what mythos was/is?
reply
notatoad 2 hours ago
typically, when any AI company says a model performs as well as fable, all they're really telling us is that the benchmarks that exist for measuring AI capabilities aren't very good.
reply
quotemstr 3 hours ago
Big model smell is a real thing. For certain classes of problem, ones you get a feel for but can't easily articulate, a big last-gen model can get you what you're looking for when no quantity of tokens from some ultra-RLed mid-size latest generation model can.
reply
throwaway2027 4 hours ago
After yesterday outage is the new Opus 5.5 load-bearing?
reply
lgessler 3 hours ago
I should find information about the user's concern instead of just assuming.

The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.

One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.

reply
handfuloflight 4 hours ago
It's worth stating why, and depends what seams you're pulling at this sitting.
reply
Retr0id 6 minutes ago
Your premise is half right, and the half that's right is better than you think.
reply
rich_sasha 4 hours ago
Your instinct is basically right, and the research backs it up.
reply
cmrdporcupine 3 hours ago
And here's the important part...
reply
danw1979 3 hours ago
you win the thread
reply
staticman2 4 hours ago
I'm gonna be straight with you—I don't have the evidence to say whether or not it's load bearing.
reply
danw1979 3 hours ago
Good point — but I’ll gently push back on that. It’s not an outage, it’s a service degradation.
reply
ThouYS 3 hours ago
You're right to bring this up - and this is where it gets interesting
reply
hmokiguess 3 hours ago
You're right, this changes everything, and here's why it matters.
reply
aoeusnth1 3 hours ago
You were right to call that out, and the evidence makes a stronger case than you are stating.
reply
cronin101 4 hours ago
It certainly _seams_ that way
reply
sailfast 3 hours ago
Let me verify before I come back to you with an answer that is incorrect.
reply
carlos-menezes 3 hours ago
One thing worth flagging here: 5.5 appears to be a load-bearing seam in the numbering system.
reply
loopmonster 2 hours ago
That's the sharpest point anyone has made in this thread so far, and it reframes the entire conversation.
reply
RGS1811 3 hours ago
This question is real.
reply
bibimsz 2 hours ago
One pushback: there is no Opus 5.5. You might have meant Opus 5.1, the latest Opus model available.
reply
fghorow 3 hours ago
"Danger Will Robinson!"
reply
esafak 3 hours ago
Wrong century, brother.
reply
tda 2 hours ago
[flagged]
reply
sznio 4 hours ago
I'm more excited by the Haiku 5.5 announcement buried in this post. I'm wondering if we will finally get a decently capable fast model.
reply
booty 3 hours ago
If you're able to use the OpenAI ecosystem, Luna's price/performance is really good. Almost like "they messed up and accidentally made it too good" good.
reply
NorwegianDude 60 minutes ago
OpenAI didn't mess up. The model would have been 100 % pointless and obsolete without the large price cuts it got, because of the cheap Chinese models.

The open models are getting closer and closer, and because they're open, people are not forced to pay the silly markup that is often over 1000x the cost to serve the model.

reply
copperx 2 hours ago
Better than what? Deepseek? GLM? Gemini 3.8?
reply
copperx 2 hours ago
Why are you excited about it? Deepseek is everything Haiku wishes to be and more.
reply
lanyard-textile 3 hours ago
Agreed. They've been so quiet about it, and retirement for Haiku 4.5 is right around the forner.
reply
ygouzerh 3 hours ago
What are you using Haiku for?
reply
adastra22 36 minutes ago
Things that Jev is probably a better tool for.
reply
Zambyte 3 hours ago
Not the same person but... nothing. Haiku just hasn't been an interesting model for a long time. If you want cheap and fast, there are lots of options that are simultaneously cheaper, faster, and capable than Haiku.
reply
mavamaarten 3 hours ago
I use it for executing well-prepared plans sometimes. And for exploring larger codebases.
reply
system2 4 hours ago
All I care about is the token price for the API. Haiku cannot get close to GLM or Mimo.
reply
enraged_camel 3 hours ago
We use Haiku 4.5 inside our product. It continues to be absurdly capable for converting natural language to structured JSON based on a set of fairly complex business rules.
reply
anthonypasq 3 hours ago
bro why. its literally the most overpriced model in existence right now. i could name about 10 models off the top of my head that would be better and cheaper
reply
enraged_camel 2 hours ago
We tried Luna and it scored way lower in our evals. Muse also. We haven't had a chance to test others.
reply
Game_Ender 2 hours ago
How much time were you able to put into tuning your prompts? And was it worse on all fronts (cost, latency, accuracy) or just some?
reply
abtinf 3 hours ago
Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.

I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).

Edit to address questions below:

ChatGPT supports oauth login.

Exe.dev has it built in. IIRC, pi also has it built in via /login.

reply
polalavik 3 hours ago
ya i've been a gpt hater for a while. almost exclusively used claude up until astra. astra feels like it blows everything out of the water. its fast, correct, organized, and less verbose.
reply
abtinf 18 minutes ago
Yes. Also, you get image generation included with the ChatGPT subscription, which is very nice for certain kinds of development.
reply
cbg0 3 hours ago
How about cheaper? Astra is $10 in $50 out, Opus is $4 in $20 out. Even on a subscription you'll get considerably more usage out of Opus.
reply
qlte 2 hours ago
Per the link someone else posted, the actual difference in $/task is not nearly so stark:

https://artificialanalysis.ai/models/releases/claude-opus-5-...

  Opus 5.5 Medium = $1.34
  GPT-6-Astra High = $1.76
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):

  Opus 5.5 High = $1.82
  GPT-6-Astra High = $1.76
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.

So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.

reply
abtinf 20 minutes ago
A cheaper price has no value if I can’t use the thing I’m paying for.

The Claude lock-in simply disqualifies anthropic entirely (for my use).

reply
margorczynski 2 hours ago
In the end what matters is how much you pay for the task you want completed. And Astra will usually do that using less token and offer a better quality solution so in the end it might be cheaper.
reply
notatoad 2 hours ago
yeah, Astra burned through 70% of my weekly usage in ~5hrs on a $100 plan. even fable doesn't run out that quickly for me. it's great, but it's on the same tier as fable for me - use it sparingly, only when really necessary.
reply
copperx 2 hours ago
> Even on a subscription you'll get considerably more usage out of Opus.

That's an incredibly bold assumption.

reply
cbg0 2 hours ago
It's not, I have a subscription to both and Astra burns usage like crazy.
reply
mlcruz 2 hours ago
What worked well for me was a custom version of Open Web Ui with some customization to spawn an exe.dev instance for each new chat. I can just work on my phone, deploy stuff for development purposes on an easy to share way etc.
reply
roughly 3 hours ago
> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

Can you give more details here? This sounds intriguing.

reply
sidrag22 3 hours ago
Anthropic is absurdly vague about 3rd party harnesses for subscriptions, if you try to use anything besides Claude Code, you are likely at risk of getting banned, you can "do it", but are at their mercy if they decide to ban you. OpenAI gives their blessing to using oauth on any harness, you can make your own or use any of the popular public ones like opencode, pi, whatever exe.dev is that this guy mentioned.

So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).

reply
onlyrealcuzzo 3 hours ago
> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

This is news to me. Excited to try it out! Thanks.

reply
nchmy 3 hours ago
news to me as well. i thought you were forced to use Codex if you wanted their subscription. I completely ignored it because of that. How do we do it?
reply
KeplerBoy 3 hours ago
With the pi harness it just opens the browser (or gives you a link if you're on headless) and you sign on as usual.
reply
felixgallo 3 hours ago
If you read the page, Opus is now significantly better than Astra while also being cheaper and having more performance headroom available.
reply
abtinf 3 hours ago
I read the page. It seems like a marginal improvement.
reply
ryanscio 3 hours ago
Let's wait for independent benchmarks at least
reply
felixgallo 3 hours ago
the benchmarks provided are already from independent organizations:

Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)

FrontierCode v1.1 - Cognition

CursorBench - Cursor (now SolarBoringSpaceXAI I believe)

GDPVal-AA - Artificial Analysis

AutomationBench - Zapier

Humanity's Last Exam - CAIS and Scale AI

Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute

OSWOrld - XLANG Lab @ the University of Hong Kong

Chartography - Surge AI

reply
kibae 59 minutes ago
> Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently.

This is where Chinese models are going to eat Anthropic's lunch.

reply
nomel 26 minutes ago
In those specific domains, sure. What percentage of paying users would you say that is?
reply
AlfeG 37 minutes ago
So it will not usable to do anything with hardening Your own site.... I'm so tired of this. I just want adjust cookie behavior of own site...
reply
tomaskafka 12 minutes ago
"You're right, and it's the exact thing I flagged two turns ago and then did anyway." - Opus 5 xhigh, today.

About the time.

reply
2001zhaozhao 2 hours ago
It's great that we are finally getting bankable rate limit resets for subscription users. According to another comment here they apparently last a month.

I'm assuming that subscription usage limit is increased in line with the price decrease on the base model and that it's in line with the model's API price drop. Still a good change.

This is a breath of fresh air on how they treat subscription customers. Hoping they keep this up.

reply
mgw 4 hours ago
They mention "the first model in our new Claude 5.5 family". Obviously that means Fable 5.5, but hopefully also a usable update to Sonnet and Haiku. Sonnet 5 hasn't really had a place in the line up for anyone I feel.

Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.

reply
mudkipdev 3 hours ago
It does mention sonnet and haiku.
reply
simianwords 3 hours ago
And not fable lol
reply
enraged_camel 3 hours ago
At the end of the post they said Sonnet 5.5 and Haiku 5.5 are coming soon.
reply
tomaskafka 3 hours ago
Excellent, maybe Anthropic can use it to fix Claude Code Desktop kicking me back to login every week or so, and forgetting whole state (opened windows = the only way of managing active working set) when I sign back in, if it's that good.

Seriously, both flagship GUI apps (OpenAI and Anthropic) are a full of glaring UX issues (for ChatGPT it's not naming their windows, so window switcher has 10 entries of "ChatGPT" and you can cycle them all to find the one you want).

reply
jjcm 3 hours ago
Image->HTML tests:

Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...

Opus 5.5's output: https://html.non.io/annui-opus/

Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.

For comparison with other drops this week + current #1:

Astra: https://html.non.io/annui/

MiMo: https://html.non.io/annui-mimo/

Grok 4.7: https://html.non.io/Annui-grok/

reply
copperx 2 hours ago
By any chance did you tried Deepseek 4 or 4.1 and GLM 5.3 or flash?
reply
jjcm 2 hours ago
I've done GLM 5.3 previously here: https://news.ycombinator.com/item?id=49295420

Worth noting though that GLM 5.3 isn't multi-modal, so it doesn't have a vision layer. It is quite clever and hacks around it pretty effectively however. I'm running a deepseek 4 build now and will reply shortly with that.

reply
copperx 48 minutes ago
Awesome. You have a really neat benchmark.
reply
naet 3 hours ago
What is your workflow for making these?
reply
jjcm 2 hours ago
The designs are outputs from my own site. This has an overview of the process: https://diffui.ai/learn/new-site

The gist of it though is I take a prompt, expand it into a json blob specifying structure/palette/positioning of elements/etc, feed that into a diffusion model to output a few choices. Once I lock in a choice I take the pixel output + json blob and use it as input into followup pages. The json helps preserve the brand across multiple pages.

Once I have all the inputs I take their corresponding image+json blobs and feed them into an agent to create a web implementation.

For image models, diffui currently uses gpt-image-2.5, mai-image-2.6, and very, very rarely a post-trained version of flux 2 dev I've made for web design, though that one will be deprecated soon.

reply
mosselman 55 minutes ago
What I don't get is, why would we still use Fable now? What is its reason for existing? If it is more intelligent and cheaper that is. Why are they advertising it as the model to use for when you really have to think when their benchmarks show Opus 5.5 is better at everything?
reply
Gander5739 4 hours ago
reply
tomhow 4 hours ago
Comments moved thither. Thanks!
reply
km144 4 hours ago
Can you fix the link on that post then? I duped because that post links to a diff that tells me nothing about Opus 5.5
reply
tomhow 3 hours ago
I did that but I recognize that even though your submission was a few minutes later than that one, you posted the better link, and you're also an established account (the other post was from a new/throwaway account), so I've restored this submission and moved the comments back to it to reward you.
reply
meerita 4 hours ago
As long as it's not as verbose as Opus 5, I am quite happy with a better version that's also less expensive. I will test it tonight. Grok 4.7 was horrible, and for mundane tasks I am relying on DeepSeek Flash 4.1 with great success using OpenCode.
reply
toephu2 31 minutes ago
When using max effort, I run into context compaction quite a lot. I haven't seen any increase in context window size at all over the past half year (stuck at 1M).

Are the frontier labs even working on this problem?

reply
garo-pro 3 hours ago
Finally confirmation that Haiku was not forgotten and will be coming soon, althouhg I find it quite interesting they skipped 5 and directly skip to 5.5 with all models, including Sonnet which is not super old. I suspect they found something breaking that allows to release this. Recently they struggled with keeping up a 50 % weekly limit increase and now they're putting out 30-40% faster and cheaper models even faster, with much more better benchmarks, a limt reset command and five hour limit increase. It seems more like the opposite and as if they never struggled, thus, I very much believe they found something very effective and new.
reply
aesthesia 51 minutes ago
Sonnet 5 was released a while before Opus 5, so it's just Haiku that didn't get a 5 release.
reply
skunkworker 4 hours ago
At this point I'm convinced they are skipping numbers so soon they will be at or ahead of OpenAI's numbering scheme.

Is the Xbox 360 (Xbox 2) vs PS3 debacle all over again.

reply
ekckekcjekfj 3 hours ago
And how was the Xbox 360 naming choice a “debacle”, exactly?

It was odd at the time, yes, but no one really minded it truly. Heck, Xbox “ONE” was a lot more of a fiasco/debacle than “360”—but there’s no parallels to be drawn with “ONE” here.

I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.

reply
m101 3 hours ago
Funny how they talk so much about safety when most people don’t give a hoot about it, and actually have quite the opposite reaction
reply
nyx 3 hours ago
People aren't the target audience of that part of the post. They're hoping saying enough safety stuff will ward off the looming regulatory sledgehammer.
reply
m101 2 hours ago
They actually want that to come protect their business model
reply
pavlov 3 hours ago
Anthropic is one of the most valuable companies in the world. Their comms are designed to appeal to a very wide readership.

HN is a bubble that's mostly out of touch with what regular people use or care about.

In 2007, HN was convinced that nobody uses Microsoft products. In 2016, it was that Facebook doesn't have any real users and is dying. In 2026, it seems like nobody cares about AI safety and everybody wants to run local models.

reply
Retro_Dev 3 hours ago
> Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude. Our September 2026 threat intelligence report details the illicit distillation activity we’ve detected and disrupted so far.

Such a negative tone they put on this. Distillation is amazing, because it means anthropic and openai fail to keep a monopoly. Who even are they who claim it's unethical? If it is truly unethical, then so is the mass data scraping they do on my personal website on a regular basis (without my consent), and all the unauthorized use of content produced by authors, blog writers, wikipedia contributors, and creators everywhere. If it is truly unethical, then anthropic, openai, meta, google... all these companies should have deleted their LLMs long ago. This wording disgusts me.

Heck, it would be amazing if we had more models without guardrails - some of the models that are produced via heretic[1] are actually quite nice to use - in particular, I've enjoyed investigating Chinese censorship by interacting with an abliterated model of Qwen3.8-27b. If security is really a concern, then secure your systems - don't attempt to dumb-down the tools we use. If someone breaks your window, then they are responsible, not the hammer they use to do so.

[1]: https://github.com/p-e-w/heretic

reply
wren6991 3 hours ago
I'm confused how they have been able to create so much public negative perception around distillation. It seems pretty clear that they are the only ones who lose out, and everyone else benefits. I don't have any ethical issues with it, nor is it illegal: at worst it's a ToS violation.

IMO the biggest problem with distillation is that not enough people are openly doing it. I would love to see more small, competitive US labs instead of having the eggs in 2~4 baskets (depending on how you count).

reply
ACCount39 2 hours ago
The issue with distillation is: one lab spends $$$ on bleeding edge R&D and expensive RL runs to improve capabilities, and other labs just yoink the raw reasoning traces and mid-train/post-train on them to get 90% of the way there for a small fraction of the cost.

An even smaller fraction of the cost if they do it by buying AI access at as much of a discount as they can find, including black market resellers, and then reselling that access to paying users again with a proxy. As is common.

This gives ruthless "fast followers" an economic edge over the innovator that's putting in the real work.

The dynamics are very much alike to what patents and copyright law are supposed to prevent. Same type of "we took the products of your work and used them to undercut you". Except there are no laws against distillation - so most of the enforcement happens on model provider level.

reply
wren6991 2 hours ago
There's an implication that other companies are improving because they're scraping Anthropic, not because they're investing in better architecture, compute efficiency, or their own synthetic data pipelines. I often see Chinese labs' progress dismissed as "they just distilled Anthropic" and I find it hard to reconcile that with all of the interesting research and open-source tooling that they release.

Is there actually that much capability transfer from non-logit-matched distillation, or is Anthropic just another unwilling source of data?

reply
ACCount39 40 minutes ago
There is, in fact, "that much capability transfer from non-logit-matched distillation".

Even the early papers on distillation techniques found that surprisingly small distillation datasets can improve task performance noticeably on some specific task types - and that valuable adaptations like SFT/RLHF instruction following can be distilled from one-hot non-logit traces.

A big part of what distillation really gets you is: paving over the mismatch between pre-training and final performance. A base model is trained to spit out fitting text, but not to instruction follow, reason autoregressively, self-check or use tool calls - like an AI has to. There is transfer straight from the "text prediction" pre-training objective, and pre-training sets the foundation for all that follows - but the capabilities you get "out of the box" with it are often unrefined and fragile. Which makes some sense - internet text doesn't often include raw chain-of-thought autoregressive reasoning. It's not the kind of thing humans tend to write.

Reasoning traces? They let an AI learn proven techniques and adaptations directly, from an AI that was already taught "how to be an AI" in other ways.

It's why this kind of distillation typically plugs into mid-training and post-training, not pre-training.

Now, I'm not saying that all Chinese companies do is eat tokens, distill and lie. That just isn't the case. They developed or refined numerous training techniques and architectural adaptations - like deep fusion for high performance visual input, RLVR with GRPO, trunked MoE, storage-efficient and bandwidth-efficient attention formulations, or residual routing techniques like AttnRes. Some of those are used widely now, and some are still on the uptake but show good promise.

But Chinese labs are enjoying massive efficiency gains from being able to distill from the frontier instead of doing things the hard way. It's a leg up. It lets them put their supply of R&D effort and RL compute elsewhere. They wouldn't be nearly as advanced if they couldn't do it.

reply
wren6991 15 minutes ago
Thanks, this is interesting and there were multiple things I didn't know here.
reply
Retro_Dev 2 hours ago
I'm fine if they put preventative measures in place to protect their work. They already do so. I am NOT fine with their mass manipulation of public opinion to fuel an entirely hypocritical viewpoint. Like, any argument here is hypocritical - but they aren't saying what is REALLY HAPPENING ("distillation steals our work and reduces our profits"), and are actually saying words that make other people fight their battle ("national security", etc).
reply
villish 2 hours ago
The workarounds used to bypass Anthropic's security measures are quite illegal. They use stolen credit cards, API keys, and accounts. That is only possible in China because any other US/EU lab doing the same would get into massive legal trouble.

That's the moat. Mistral has the capability but not the legal protections.

reply
aragornii 3 hours ago
What I'm mostly interest in is the Communication section. Opus 5 was so convoluted in the way of answering that was really frustrating me.

Instead of instilling confidence, it was overwhelming. Not sure if I'm the only one.

reply
ieie3366 48 minutes ago
Quick test for my gamedev project: It feels like using Fable, but faster, and obviously wayy cheaper token-wise.

Has oneshot all of the quite complex bugs / debugging tasks I gave to it which I know opus 5.0 would've struggled with

reply
rumblefrog 3 hours ago
I'm glad they specifically called out the prose issue, I was always pinned to Fable 5.1 because I wanted to avoid the unreadableness of other Anthropic models.
reply
pookieinc 4 hours ago
“It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”

They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?

reply
randomblock1 3 hours ago
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
reply
jbellis 3 hours ago
Anthropic knows that the benchmarks showing Opus 5 better than Fable 5.1 are measuring something that's less than entirely useful.
reply
meric_ 3 hours ago
Opus does seem like a more powerful coding workhorse based on the benchmarks listed though. Good coding performance, faster and less verbose, cheaper.

Will be interesting to see how people's opinions of it line up IRL, but so far I've loved Fable so hopefully will love this one too

reply
buntp 4 hours ago
Masterpiece by openai to call their model '6', this model feels already behind
reply
frshgts 2 hours ago
Anthropic will pull a PHP and skip '6' to go straight to '7'.
reply
wren6991 3 hours ago
Smart move would be to move to year-based versioning (26.09). A 4x advantage
reply
FergusArgyll 3 hours ago
Well, they're actually older so it makes sense that their model versions should be ahead
reply
madjam002 2 hours ago
I noticed a big speedup in Opus 5 on Max x20 since about 10 days ago, and I feel like the model has been performing better.

It would be great to know if this was Opus 5.5 or a lesser incremental improvement, as otherwise it's difficult to judge whether Opus 5.5 is expected to be a big improvement.

It's frustrating that there isn't more transparency here.

reply
kar1181 28 minutes ago
Whatever I think of anthropic, that webpage is a truly nice piece of work.
reply
Catloafdev 4 hours ago
> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.

Sounds like they noticed the complaints. I'm curious to see what LLM-isms this one may have.

reply
dgroshev 3 hours ago
I don't think it's substantially different. I just pasted a random chunk of code and asked Opus 5.5 to comment on it:

> The Vercel target is hard-coded. That's common and not wrong, but it's opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.

> Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel's dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it's easy to forget.

> Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate (perhaps a single shared testing tool that only needs one public address), but if you're reviewing this rather than just reading it, that's worth confirming.

It has the same annoying cadence and writing style with slightly less prominent claudisms.

reply
cruffle_duffle 8 minutes ago
“It has the same annoying cadence and writing style with slightly less prominent claudisms.”

Seems like it based on my first session. It still does the whole “bury the important thing in a pile of words” coupled with the “it might actually be important” thing… so basically you never really know what it’s talking about.

Honestly I trust opus so little that the entire “opus” brand is completely tarnished. Its writing style is so god awful that it needs more than just a point release. Either dump the name and ship a different model entirely or at minimum call it “opus 6”. Calling it 5.5 makes it sound like it’s basically a continuation of the same garbage output that 5.1 had but with some minor adjustments. And based on my single first test, that is what it appears like to me.

reply
sashank_1509 3 hours ago
Maybe if we had a single human we talk to 24/7 at scale, we would get annoyed at his cadence and style. You need variety to not pick up on known patterns I assume, which a single model can’t replicate?
reply
dgroshev 3 hours ago
No, it's just poor writing. Actionable points are buried inside the paragraphs and over-hedged, and one point is completely made up. Compare to a five second rewrite:

* Consider leaving a comment about the hard-coded Vercel target. It's not clear where does it come from.

* [This is just a bullshit point, because the domain is not "added to" Vercel, it's provided by Vercel]

* Are you sure that [CENSORED] is prod-only? The name suggests otherwise. [also, what "if you're reviewing this rather than just reading it" even means?]

reply
wren6991 3 hours ago
> also, what "if you're reviewing this rather than just reading it" even means?

It means "I'm treating you as lay-person punter, not a developer working on this project." Opus 5 feels like it's constantly trying to reward-hack me into treating it as intellectually honest and epistemically humble, while in the same breath it talks down to me and tries to smuggle its own bullshit assumptions and assertions into the conversation unchallenged. No progress on this front apparently. Glad I cancelled.

reply
dgroshev 2 hours ago
Good points, but then even this little snippet is internally inconsistent. If I'm a lay-person, why should I care that "a comment would help"?

Claude is just comically bad nowadays.

reply
redox99 3 hours ago
Nah it's definitely a Claude thing. Other models even though they have their style are less annoying and less stereotypical.
reply
brandon272 4 hours ago
You were right to notice the complaints. One decision remains, and it is yours, genuinely.
reply
gwking 3 hours ago
I appreciate humor here, but there are now a dozen of these comments on every thread about Claude. They no longer adding anything substantial and dilute the discussion.

I don't mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.

reply
brandon272 3 hours ago
I agree. Hopefully Anthropic has fixed Opus' ridiculous communication style so that people - like me - no longer have any kind of weird impulse to imitate it.
reply
fragmede 55 minutes ago
It's a load bearing joke that was funny the first time but we're going to beat that dead horse until it starts getting funny again. If you beat it enough, it will get funny. Beatings will continue until morale improves.
reply
qurren 3 hours ago
[flagged]
reply
gekoxyz 4 hours ago
It was difficult to not notice them. Opus 5 was unusable, most of my team went back to Opus 4.6 for most of their work. I hope we can move forward now.
reply
ithkuil 3 hours ago
It's unbearable but nothing that couldn't be fixed with postprocess.
reply
mavamaarten 3 hours ago
How? Explicit instructions, memories and even skills have not been able to keep Claude from saying "genuinely" every two sentences and keep it from explaining heavily what something _isn't_.
reply
aray07 3 hours ago
Opus 5 was just incoherent - curious to see what improvements they have made here. Would love to see some kind of postmortem to better understand how writing styles change from model to model.

I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs

reply
username_my1 3 hours ago
I'm genuinely confused what's the relationship between LLMs improvements and them being so incoherent.

and it's not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.

I wonder if there are studies around this.

reply
meric_ 3 hours ago
Remember when OpenAI models loved talking about goblins and whatnot due to the RL?

https://openai.com/index/where-the-goblins-came-from/

Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it

reply
adastra22 32 minutes ago
It is the switch from RLHF to RLVR. It benchmaxes better, but benchmarks don't cover human usability.
reply
ygouzerh 3 hours ago
Can it be that now they are getting optimized against benchmarks that are valuing logics, rather than human appreciation? (I am not an expert at all, just an idea)
reply
Eliezer 2 hours ago
It's the reinforcement learning rather than supervised learning.
reply
j_heffe 2 hours ago
Maybe it's the time period we're in, maybe I'm just grumpy, but it bugs me that they release a new model every single week and the new one is just a fine-tuned version of the "old" one. If 5.5 performs similar to Fable and really does cost 40% less, then 5.5 really should've just been Opus 5. And they're essentially admitting that they are shipping slop.
reply
drnick1 4 hours ago
A load-bearing promise.
reply
edude03 2 hours ago
> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1.

Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM

reply
bredren 3 hours ago
Notes on communication:

"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"

and

"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."

and

"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."

I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.

If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.

reply
ayhanfuat 3 hours ago
Looks like Anthropic is starting to give bank reset as well:

> Reset for free: Get extra wiggle room to explore Opus 5.5. Expires Oct 22.

reply
jdthedisciple 3 hours ago
I dare anyone to convince me the benchmarks are not meaningless.

Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?

How would this alleged difference (most likely bs) actually show up in reality?

GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.

reply
enraged_camel 3 hours ago
Ah, so you didn't read the article.

>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

reply
jdthedisciple 26 minutes ago
You didn't read my question, bc that excerpt doesn't answer, nor do they demonstrate

> how would this alleged difference (most likely bs) actually show up in reality?

Furthermore: so they admit it's bs but still placate it like its the next biggest thing ever ... alright

All I'm saying is I refuse to buy into it anymore – yet many on here still do, including ... you?

reply
cogythea 4 hours ago
Interestingly they've changed their approach to usage resets for this release - with previous releases I've had my usage instantly reset, but now in the Claude app I've got a 'Reset for free' button that expires Oct 22, which seems to effectively be a whole new usage window I can activate whenever's convenient
reply
jdmoreira 3 hours ago
then they copied that from codex because thats exactly how codex works
reply
ryanscio 3 hours ago
Input $4/MTok and output $20/MTok is a welcome surprise. Cheaper than Opus 5/4.8, Astra 6, Fable 5.
reply
benjiro29 3 hours ago
The biggest one is the Cache reads going from $0.50 to $0.20 ... Read/Writes dropping by 25% but Cache reads by 60% has a much bigger impact.
reply
ramoz 3 hours ago
It crushes Fable on benchmarks and even in the blogs "real-world" studies. But... they are communicating like it ~sometimes~ provides Fable intelligence?

A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??

reply
bitexploder 3 hours ago
What if the recent Fable intelligence regression was basically just them serving Opus 5.5 until they got it working well?
reply
dbbk 4 hours ago
This makes Fable not really make any sense?
reply
re-thc 4 hours ago
You bet there will be a new Fable soon.
reply
nozzlegear 4 hours ago
Pacing the frontier btw
reply
re-thc 38 minutes ago
That's a Fable / Myth(o). The name said so.
reply
petesergeant 4 hours ago
didn't they say Opus 5 was Fable-level too tho? Let's see, I'm at the point where I don't think benchmarks really tell us very much any more. I'd love it to be as strong as Fable, but I'm skeptical about how that will look in practice.
reply
alpineman 4 hours ago
So we skipped 5.1, 5.2, 5.3, and 5.4: we really are plateauing
reply
jatins 3 hours ago
> We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions.

Thank you.

reply
notduckrabbit 3 hours ago
They purport 40% drop in costs due to lower token pricing (presumably aimed at winning back the many of us that switched providers in discovering Opus 5 unusable) and improved token efficiency.
reply
manmal 3 hours ago
That cost reduction seems to stem from cheaper cache reads, mostly.
reply
dom96 2 hours ago
Just updated KillSwitch-Bench with this new model: https://bench.killswitch-lang.org/

It does perform slightly worse than Opus 5, but it is significantly cheaper and faster.

reply
yipinwong 2 hours ago
I spent about $5 per sentence in my resume using Fable 5.1 (High) to verify accuracy, inconsistency, and edit.

Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.

Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.

reply
copperx 2 hours ago
$5 per sentence?
reply
yipinwong 2 hours ago
I am sorry, I meant to say I generated STAR out of my resume line, trying to generate STAR, and polish it thus $5.

---

I provided crapton of context for that one resume line. All the work I did, documentations for my justifications, etc.

I initially messed up and came out ot $5, rest of resume used around $4 per line (I used a fresh new session on purpose).

---

As a clarification, $2.2 average for OPUS 5.5 was the same process in a new session, same context, same prompts.

Also adding verification for that Fable 5.1 output in the same sesssion.

reply
glub 3 hours ago
> For users with cybersecurity use cases that may be blocked by our cyber safeguards, we recommend accessing our models with reduced cyber blocking classifiers via our Cyber Verification Program. Claude Opus 5.5 will be available through this program in the near future.

Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.

Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?

reply
doodlesdev 2 hours ago

   > Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5
Big, if true.
reply
jacobgold 3 hours ago
I use the other 50% of my $200/mo Claude subscription by having Fable run Opus subagents for a lot of work. That way I don't have to deal with Opus directly.
reply
lousken 3 hours ago
Cost to Run Artificial Analysis Intelligence Index is higher than previous Opus, so still not cheaper
reply
variety8675 4 hours ago
I hope this actually fixes the terrible writing style of Opus 5
reply
emadabdulrahim 3 hours ago
It’s not a 100% fix, but with concise output style on, it’s much better.
reply
akhilome 3 hours ago
I found having a reminder at every turn through the UserPromptSubmit [1] hook helped with taming the word salad from 5.

Hopefully the output from vanilla 5.5 is as good as they claim. I’ll try out later tonight.

[1] https://kizi.to/claude-talks-too-much/

reply
gavinray 4 hours ago
[dead]
reply
34679 3 hours ago
I don't care how good their models get, I won't sign up for one of their plans until they define "X" in their pricing. 5X of this plan, 20X of that plan means nothing when they never tell you what "X" is.

Maybe this model can finally figure it out for them.

reply
breezybottom 3 hours ago
"Where Opus 5.5’s advantage is very clear is efficiency."

Not efficiency in writing, clearly.

reply
Retr0id 3 hours ago
> Opus 5.5 (1M context)'s safeguards flagged this session. You may be seeing this for the first time on an Opus model: Opus 5.5 (1M context) is more capable and has stronger safeguards as a result, which can sometimes flag non-cybersecurity work. We're improving these safeguards to reduce the amount of incorrectly flagged messages. Opus 4.8 is answering instead, or you can edit and retry with Opus 5.5 (1M context).

Yay, yet another model I can't use for anything interesting, even with CVP.

reply
Foobar8568 3 hours ago
I have just switched to 5.5. First mistake was stale environment variable, didn't realize it was replaced, "oh my memory had stall data" and that's it. Second one, a powershell command had the wrong syntax. Great for my first two prompts.
reply
isodev 3 hours ago
So is it cheaper? Are we AGI yet? Am I left behind? I didn't have patience for the intro animation on the website... maybe one day, Claude Code will understand accessibility but that day is not today.
reply
spiderice 3 hours ago
You didn't have the patience to scroll down, so you decided to come post about it here and waste all of our time?
reply
isodev 3 hours ago
The site is horrific so no, I didn't scroll.
reply
b38484848 3 hours ago
we will be agi in six months as in the last 36 months
reply
desmondl 3 hours ago
I'll have to try 5.5 on my work's Cursor account. If they really solved the communication issues, I might consider moving my personal account from Codex back to Claude Code.
reply
calibas 4 hours ago
> We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in.

We can't test it properly because it knows it's being tested.

reply
johntb86 3 hours ago
Just make it always think it's being tested, and problem solved.
reply
km144 3 hours ago
I think this release is really going to give them a hard time selling Fable:

> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.

reply
booty 3 hours ago

    "benchmark margins have become a less 
    reliable guide to real-world differences" 
    sounds like a big problem.
My guesses:

1. Real-world use cases typically involve big, hairy, crufty, tech debt laden codebases and benchmarks do not.

2. AFAIK "success" in a benchmark essentially boils down to "do the tests pass and do we get the right result?" which is something the LLMs have been achieving with ease for a while, except maybe for uber-challenging coding tasks that would be outliers in just about any workplace. Whereas real-world software engineering is usually just a bunch of CRUD... and "success" involves harder to measure dimensions like "maintainability" and "did you overengineer this?" and "how did you cope with a bunch of vague and maybe contradictory business requirements?"

Having said all of that, I have never ever looked inside any of these benchmarks. I'm putting my guesses out here strictly in the tradition of "the quickest way to learn about something is to be wrong about it on the internet."

reply
simianwords 3 hours ago
It’s likely that they have internal benchmarks but they are communicating to people who can only gauge through external benchmarks.
reply
suddenlybananas 3 hours ago
Why wouldn't they report these benchmarks?
reply
CPLX 3 hours ago
Opus 5 fucking sucks. Like it's horrible. I use Fable for coding and anything important and I use Opus 4.8 for things like recursive email categorization, transaction matching, and other stuff where I don't want to burn as much quota.

In my experience Opus 5 is the worst of all possible worlds, it's dumb and headstrong. It just runs away with tasks you didn't ask it to do, is reckless, and basically is unusable in my experience.

Not sure why but my guess is that this will be worse. Happy to be proven wrong.

reply
port3000 3 hours ago
I believe Opus 5 isn't meant to be spoken to by humans. It's great at executing but I reckon it's intended to be spoken to by other models such as Fable. I use Fable as the orchestrator, only speak with Fable, and all implementation, recon, design etc happens with Opus 5, with Fable reviewing (and translating).
reply
booty 3 hours ago
That's interesting.

I've really gone in the opposite direction: having a dumber model orchestrate. In my case, it's usually a Luna orchestrator spawning Sol/Astra subagents to do the "big brain" work of planning and reviewing.

Reason I went with "dumb orchestrator" was just to save tokens. Having Opus/Sol (let alone Fable/Astra) orchestrate was burning tokens like crazy for me even when much of the gruntwork was being done by Luna/Sonnet/Haiku subagents. (Luna is also really good, like way better than Sonnet...) Perhaps it was a skill issue on my end though, maybe I wasn't just managing context properly.

reply
cbg0 3 hours ago
I use it frequently with a lot of success on "Medium" effort, it overthinks like crazy on higher levels, but YMMV.
reply
Syntaf 3 hours ago
Yeah if anything Opus 5 taught me how little benchmarks mean to the actual real world performance of these models.

"Better" in every sense of the benchmarks and absolutely horrible results in my day-to-day work.

The verbosity, goal post moving, tendency to leave work unfinished, over focusing on unrealistic root causes when debugging, etc... etc...

It was the first time I actually pinned my models back because I just could not work with 5 for the price and performance it gave me. Hoping 5.5 is better this time around....

reply
somewhatjustin 3 hours ago
> Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.

Nice. I was starting to think Haiku was going to be abandoned.

reply
blfr 3 hours ago
It's awesome that the apt packages for claude and claude-code are out right now. I can test-drive Opus 5.5 right away. Very cool, Anthropic.
reply
HarHarVeryFunny 2 hours ago
METR: Is it safe? Has it escaped confinement?

Ants: It's a good model, sir!

reply
aurareturn 3 hours ago
I found myself going back to Fable over and over again. At this point, I’m not sure if I’m just used to its style or it is truly more capable.

I tried Opus 5 and Astra.

reply
datadrivenangel 3 hours ago
But have they made it any better at communicating clearly? I cancelled my personal subscription because Opus is so painful to read.
reply
nimonian 17 minutes ago
I recommend reading the web page. It is quite short.
reply
cruffle_duffle 6 minutes ago
I mean the webpage can say what ever it wants. The proof is using it yourself.
reply
jidaigeist 3 hours ago
>Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude.

Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.

reply
andriy_koval 2 hours ago
My bet is anthropic has NN people org who work hard to distill open models in addition to trying to find what other useful materials they can download from shady torrents.
reply
b38484848 2 hours ago
it's not safe unless it has commitees with orgies with that weird harry potter dude attached
reply
the_gipsy 3 hours ago
"bad actors" boogeyman, and we should trust some tech weasel to do the right thing? Yea we've seen who they really are, once they get a sliver of power.
reply
__vivek 2 hours ago
I'm only interested in the Opus series, if they fixed the talking issues.
reply
thatxliner 60 minutes ago
So much for pacing the frontier
reply
alvis 4 hours ago
$0.20 vs the old $0.5 cache read is pretty much 60% off
reply
nickandbro 4 hours ago
Wow! Though need to see its token efficiency to better assess. Been hearing rumors it generates much more output tokens per task.
reply
keeganpoppen 3 hours ago
my projection is that they are still gonna be pretty far behind, but they will sew it up in the next few releases. it feels like they were caught with their pants down on how much work OpenAI has put into that area, but i doubt there is some magical secret sauce that OpenAI has that Anthropic simply cannot catch up with.
reply
tag2103 3 hours ago
Why would anyone reward bad behavior?
reply
thibran 4 hours ago
Anthropic models are ridiculously expensive. I've stopped using any of their models months ago.
reply
blurbleblurble 3 hours ago
Hopefully OpenAI throws us some more usage resets now.
reply
iamsyr 3 hours ago
I don't yet have any reason to leave Haiku 4.5 and switch to Opus 5.5.
reply
greenavocado 3 hours ago
Enjoy it for the next 2 weeks until its silently quanted to 4.8 level
reply
Fizzadar 2 hours ago
So is this AGI+ now?
reply
garo-pro 2 hours ago
Opus 5.5 is now the recommended model in Claude Code's model picker, which is quite a claim, given how they struggled with capacity.
reply
keeeba 3 hours ago
Opus 5.1 came out about a month ago, what gives?
reply
velcrovan 3 hours ago
Just a guess but maybe they decided 5.1 wasn't their last Opus model. Like they would keep developing new versions of it or something.
reply
adastra22 31 minutes ago
Crazy!
reply
hadlock 2 hours ago
Frontier model labs release some kind of update every 6 weeks on average.
reply
rs_rs_rs_rs_rs 2 hours ago
That was Fable. Last version of Opus was at the end of July.
reply
aennassiri 3 hours ago
Let's see how much they benchmaxxed their model!
reply
woeirua 3 hours ago
So... why would you use Fable now?
reply
richardjennings 3 hours ago
My 20x plan was set to end tomorrow. The writing style and insistence on word vomit just became too annoying. Is Opus 5.5 worth sticking around for ?
reply
Yabood 3 hours ago
Current models, especially Opus are almost unusable because they don’t respect instructions and their responses are infuriating. They are clearly designed for token consumption. I find myself wasting a lot of time just asking it to shorten or simplify its responses. I’ll give this new model a go, but I’m not holding my breath because the last model release was supposed to fix the very same issues and it didn’t.
reply
sandos 43 seconds ago
Same feeling with oai models, wich I use 99% of the time. Sometimes I ask it about it, and it always come up with a likely explanaton but dear me it does many rounds of tool calls sometimes!
reply
Lord_Zero 4 hours ago
The test they performed to port HAProxy from C to Rust is crazy.
reply
firemelt 2 hours ago
wow its really smarter than opus?
reply
cmrdporcupine 2 hours ago
GPT Sol 6 has also released today, but no official blog announcement yet

https://www.reddit.com/r/codex/comments/1wnggya/gpt_6_droppe...

reply
bdangubic 2 hours ago
I am pacing my apple pie consumption … :)
reply
karp773 3 hours ago
I get this in my claude.ai usage:

Resets Get extra wiggle room to explore Opus 5.5. Expires Oct 22.

What the hell does this mean? There are weekly "resets" anyways. And there will be 4 of them before Oct 22.

reply
nimonian 11 minutes ago
You have an extra reset that you can trigger any time before Oct 22
reply
kingstnap 3 hours ago
> It’s good at finding and fixing inefficiencies in software

Holy shit! Its happening!

Now if we can the AI to understand this *implicitly* so that it doesn't need to be stated upfront, we might be able to undo years of "premature optimization is the root of all evil".

reply
ramesh31 3 hours ago
It seems context length has completely fallen out of the discussion since we hit 1M, is that just going to be what it is now?
reply
nailer 3 hours ago
> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.

Thanks God. Opus 5 was a massive regression compared to Opus 4.8. People were spending tokens on fixing Opus-isms rather than actually doing work.

reply
LoganDark 3 hours ago
> In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.

I was accepted into the CVP a little while ago. Does this mean I'll need to apply again?

reply
anentropic 3 hours ago
ooh exaggerated film grain
reply
Arcuru 3 hours ago
Great. Now let me use the subscription outside Claude Code.
reply
simianwords 3 hours ago
How do I get access to that reset? I can’t find it in my app.
reply
ricardobeat 3 hours ago
Great that they listened! The improvement in communication style looks fantastic. Opus 5 was insufferable and I was on the verge of cancelling my subscription.
reply
jdw64 3 hours ago
Finally, it seems like a good time to do some 'load-bearing' work on my project for a while
reply
viccis 4 hours ago
So it beats Fable 5.1, by quite a bit, on every metric? Interesting.

Might have to use my $20 Claude sub some more. I was moving away from it to a $100 OpenAI one to avoid the Claudese and poor token efficiency of Opus 5, given that I couldn't use Fable 5.1 with my tier, but this is worth trying out.

reply
scrollop 3 hours ago
Why can't they let 20usd claude subscriptions access fable in CC, as openai allows you to use astra and max modes in codex - you just pay for it in more token use.
reply
phendrenad2 3 hours ago
> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1

Great so good luck using this for any low-level embedded or operating system development (unless you really, really like Opus 4.8 and want to be greeted by its familiar face after a few minutes of work!)

reply
snvzz 2 hours ago
>safeguards hit: [Cyber]

Yup. As unusable as Fable 5.1, for assembly on 80s 68k personal computer platform. Awful.

reply
blurbleblurble 3 hours ago
It'd better be good, I'm so tired of the shenanigans
reply
theGeatZhopa 3 hours ago
is OPUS 5.5 still not reading CLAUDE.md, failing to follow told tasks, inventing and hallucionating, just refusing to read files ("read the whole file" -> read 2-lines -> infere its wrong -> destroy the codebase), needing constant babysitting just because its so UTTERLY DUMB! i cant imagine going back to OPUS 5 - i'll rather jump out of the window as to use it EVER AGAIN!!
reply
mupuff1234 4 hours ago
What happened to "slowing down"?
reply
setsewerd 4 hours ago
They're slowing down token usage, not the path to regulatory capture.
reply
icrbow 3 hours ago
If you hit wall, hit it hard.
reply
petesergeant 3 hours ago
Slowing down only makes any sense if you can coordinate a slow-down for everyone.
reply
nozzlegear 3 hours ago
Dario found himself in the prisoner's dilemma.
reply
roughly 3 hours ago
Which is one of those fun things that didn’t actually exist back when we took it for granted that our fellow person was operating under some kind of moral or ethical framework, which pretty much everyone was until the economists told us that wasn’t rational, because it turns out it’s an evolutionary advantage to operate under an ethical or moral framework because it allows the kind of coordination which facilitates better collective outcomes, which everyone knew until the economists came along to tell us we were wrong and in fact it was rational not to do so and suddenly we had the prisoner’s dilemma.
reply
mupuff1234 3 hours ago
That's just false.

Less companies involved means less pressure to go fast.

reply
WarmWash 4 hours ago
Trump got mad and investors sued.
reply
Lord_Zero 4 hours ago
The hype train must keep chuggin or it all collapses.
reply
Madmallard 2 hours ago
> cyber security and life sciences verification programs

chinese models can't come soon enough

we're already getting enshittification

reply
vividfrier 2 hours ago
[dead]
reply
giancarlostoro 3 hours ago
[dead]
reply
nicolamanzini 41 minutes ago
[dead]
reply
xenit_v0 3 hours ago
[flagged]
reply
ace2pace 3 hours ago
[dead]
reply
mrbonner 2 hours ago
[dead]
reply
SadErn 4 hours ago
[dead]
reply
hirako2000 3 hours ago
Throwaway accounts posting after a few minutes some anthropic or another ai lab.

Infomercial at its best.

No wonder we are hammered with ai announcements.

reply
danbrooks 3 hours ago
Many people knew this announcement was coming. The betting markets suggested a very high likelihood of Opus dropping today. I was anticipating this quite a bit!
reply
b38484848 3 hours ago
Exciting! Thanks for sharing these news.
reply
gopalv 4 hours ago
The whole thing reminds me of the Apple feature flag story[1] from a generation ago.

[1] - https://news.ycombinator.com/item?id=6372466

reply