GPT-6 Astra on OpenRouter
100 points by Topfi 4 hours ago | 45 comments

simonw 2 hours ago
I posted this in the other Astra thread but it's just fallen off the homepage, so...

Pelicans from Astra, plus 5.6 Sol, Terra, Luna for comparison: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...

I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models.

Astra uses less tokens overall too, for better results.

Astra transcript here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

reply
vb-8448 18 minutes ago
I find curios that Astra's pelicans are basically the same (yellow sun top right corner, green bike, same bike shape, same legs style, very similar background) while in other there is more randomness.
reply
thimabi 12 minutes ago
This tracks with what OpenAI has been saying about Astra: that it tends to do things in a certain way and it’s up to you to prompt it to change its style.
reply
BrokenCogs 53 minutes ago
Interesting that all of the bikes are turquoise colored, except for the medium effort
reply
andai 25 minutes ago
Oh my goodness. I was not prepared for Luna on "none".

Reminded me of https://clocks.brianmoore.com/

reply
tyre 5 minutes ago
Haiku the GOAT
reply
Xunjin 59 minutes ago
Do you have other ideas of "combinations" for this kind of benchmark?

I'm wondering if this is being trained on by the models today.

reply
r_lee 18 minutes ago
I have a feeling they've been doing that for a while now, even if unintentional, as it's such a well known benchmark
reply
lspears 39 minutes ago
Luna's price seems off
reply
leoqa 2 hours ago
[flagged]
reply
satvikpendem 2 hours ago
It's simonw. It's interesting to see their pelican benchmark, another comment by a different author elsewhere here shows some very good SVG generation too.
reply
samuelknight 2 hours ago
How are we supposed to know if Astra is frontier without the pelican?
reply
ComplexSystems 2 hours ago
It's absolutely related.
reply
StopTheCringe 33 minutes ago
[flagged]
reply
XCSme 2 hours ago
That's some crazy SVG generation:

https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...

It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.

reply
satvikpendem 2 hours ago
That's quite shocking, at a sufficiently advanced level we can make all non-realistic graphics purely out of SVGs, as they'd have good scaling for things like logos and app icons. I know it was technically and theoretically possible before AI but most people weren't spending hours tweaking SVG HTML. I remember making an SVG dark mode toggle icon and it took days to get it right, I assume it's one shottable now.
reply
kulahan 34 minutes ago
That photo looks pretty hilariously stupid, so this appears to be more of a first toe dip rather than some indication we can one-shot a previously difficult process.
reply
appplication 18 minutes ago
Sure, it is a bit cartoonish but it’s relatively impressive. I do wonder what you would get if you asked for photorealism
reply
embedding-shape 2 hours ago
At the bottom it says "Score 98.58", what measure is used for this score? It's kind of horrible, the perspective is all off (legs of the table makes that very obvious), the mouse/hamster has two mouths, a stub for a right paw, looks like left hand holds a melon on a stick or something, and there are pluses in the background for some reason. Not sure it'd call it "close to perfect" which the score seems to want to indicate.
reply
XCSme 2 hours ago
Do you prefer the fable one?

It's more "correct" but looks a lot worse in my opinion:

https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...

reply
kulahan 33 minutes ago
No, both of them look awful. I genuinely am unsure if this is a meme you're making that's going over my head...?
reply
XCSme 2 hours ago
Yeah, that's confusing, the score is for the entire benchmark, not for SVG generation only.

Good point about the mouths, I just noticed, lol

Imo, it's still better than most models, I personally like the stylized perspective.

You can view here all generations for all models: https://aibenchy.com/showcase/

reply
XCSme 52 minutes ago
I've replaced "Score" there with model ranking, to reduce confusion, thanks for the feedback!
reply
jaesonaras 11 minutes ago
Anyone had success using Astra as a Foundry model via Github Copilot? The error I get is that tooling is not available if reasoning has a value.
reply
vb-8448 22 minutes ago
Played in codex app a couple of hours today: it feels much faster than SOL, even if the TPS is half of it.
reply
kingstnap 3 hours ago
Its also available finally to Pro users! Just took 24 hours.
reply
InsideOutSanta 3 hours ago
They gave out bankable resets for every day people on pro plans didn't get Astra. Given that, I wish they'd waited a few more days before activating it on my account :-D
reply
paxys 3 hours ago
That's a pretty genius internal incentive to move fast.
reply
wincy 3 hours ago
They haven’t activated Astra for me yet, I have two resets now. I’ve been using the opportunity to test out how good 5.6 Sol is at computer use asking it to generate stuff in Blender which has been… interesting

Edit: nevermind it JUST gave me a notification to use it!

reply
embedding-shape 2 hours ago
> They gave out bankable resets for every day people on pro plans didn't get Astra.

Yeah, when I saw that Tweet I knew the person was saying it because they knew it'll be available within 24h.

reply
algoth1 3 hours ago
Just got it. European plus user here. Only codex, no chatgpt
reply
gavinray 3 hours ago
I have GPT-6 access in Codex and OpenAI API now

I'm a Business plan user with Cyber verification enabled, FWIW.

reply
embedding-shape 2 hours ago
Same just got access literally this minute, Pro user here, no Cyber verification but have passed my ID over to them back in 2024 or something, maybe at the ChatGPT 3 API launch or something?

Has there been anything published about if Astra uses different amount of usage from your subscription plan compared to Sol? Don't recall coming across that in the press releases.

reply
r_lee 3 hours ago
is Azure for this actually ZDR?
reply
starik36 2 hours ago
What is the actual utility of using this model on Azure? It's twice as expensive, according to the link.

Do Azure offer something that simply hitting the OpenAI endpoint doesn't provide?

reply
hhh 9 minutes ago
It's the same price as regular processing. You get guarantees microsoft give you, which are ones OpenAI won't (or require dedicated spend,) and you can use azure identities for access.
reply
claiir 2 hours ago
They’re ZDR and the OAI ones aren’t
reply
paxys 3 minutes ago
OpenAI also has ZDR for enterprise accounts.
reply
itsjustkev 2 hours ago
Compared to OpenAI flex? I'm pretty sure that is their batch processing endpoint, which is naturally cheaper.
reply
cute_boi 3 hours ago
I got it, but sadly no resets...
reply
vatsachak 3 hours ago
Damn I am so hyped
reply
marsven_422 5 minutes ago
[dead]
reply
StopTheCringe 33 minutes ago
[flagged]
reply
noob0053 3 hours ago
[flagged]
reply
taywrobel 2 hours ago
You made an account just to post this?
reply
OpenGayEye_ 28 minutes ago
[flagged]
reply
6thbit 2 hours ago
Invoking simonw for pelicans pretty please.
reply
Osama0456 3 hours ago
is it finally available on openrouter and what about fable 5.1
reply
r_lee 3 hours ago
that's literally what the link is for...
reply