Show HN: Pelican-bicycle alternatives
59 points by tkgally 5 hours ago | 24 comments
In November and December 2025, inspired by Simon Willison’s pelican-riding-a-bicycle benchmark, I had some then-current LLMs create SVGs from thirty similar prompts, such as “Generate an SVG of an octopus operating a pipe organ.” Simon mentioned that experiment on his blog [1].

Nine months have passed and much stronger models have been released, so I tried the experiment again today. The linked site shows the results.

Running ten of the prompts through six models at OpenRouter cost about twenty dollars, so I stopped there for now.

[1] https://simonwillison.net/2025/Nov/25/


kennywinker 4 minutes ago
Would love to see Qwen3.8-27b here, since that is the model most people are running locally.
reply
svcrunch 3 hours ago
I'd like to mention the Little Dorrit Benchmark [1] which I have been running for a couple of years now. It has a few nice features:

1. It tests visual reasoning and structured output in a single task.

2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.

3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.

[1] https://dorrit.pairsys.ai/

reply
vova_hn2 3 hours ago
Website looks very cool, Fable's octopus-organist looks very cute, but I feel like this benchmark (generate an SVG by a short and slightly ridiculous description) in general has been completely Goodharted [0].

I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.

[0] https://en.wikipedia.org/wiki/Goodhart%27s_law

reply
GaggiX 2 hours ago
I don't think a bunch of similar tasks can really saturate the "create a SVG of X", because the model should have a quite good spatial understanding of the world and how everything interacts.

For example Gemini 3.8 Flash seems very impressive at first glance but the results are not actually very coherent, this shows that its "world model" is not particularly great (compare to SOTA models).

reply
samayashar 3 hours ago
All models are pretty good now at generating these images. Back in the day, I remember experimenting with the pelican images and most of the models couldn't align the legs with the wheels. Right now as well, GPT messed up an octopus leg by originating it through the instrument rather than the octopus itself.

I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.

reply
honeycrispy 2 hours ago
> All models are pretty good now at generating these images.

That's pretty generous.

reply
CamperBob2 2 hours ago
All models are pretty good now at generating these images.

Not zebras. If you want to see how bad SVG output still is, ask for a zebra riding a scooter.

reply
zaphar 35 minutes ago
I notice none of the octopi seem to be actually facing the organ.
reply
steinvakt2 4 hours ago
Feels like google has a different training set than the others?
reply
dustfinger 2 hours ago
It is interesting how similar the designs are across the models.
reply
BrokenCogs 2 hours ago
Gemini 3.8 flash seems to (subjectively) be the outlier in terms of performance to cost ratio?
reply
pixelesque 38 minutes ago
Its giraffe / grandfather clock one is pretty bad... (two necks? wearing a suit?)

Weird as well, it's clearly pulled out some 1884 patent on clock designs, and a quick ddg/google doesn't show it as anything to do with grandfather clocks.

reply
GaggiX 2 hours ago
Gemini 3.8 Flash results are often not very coherent but it does put a lot of shading and details to hide the fact.
reply
eddytrex_ 2 hours ago
Does a test of instructions how to fold origami figures in a SVG/jpeg exist? Or could be useful?
reply
neilellis 2 hours ago
Well that benchmark is now saturated, what next. How fast you can hack the pentagon?
reply
sajithdilshan 3 hours ago
Interesting, out of all examples Gemini 3.8 is the best for me. Also the image style is different and more vibrant than others
reply
Jordan-117 2 hours ago
Better than Astra and Fable? It looks quite pretty and even impressive at times if you squint, but look closer and it falls apart in terms of coherency. And I say that as somebody who mains Gemini 3.8.
reply
dcreater 57 minutes ago
Why is this a good test?
reply
sceptic123 3 hours ago
> An elephant typing on a typewriter

A monkey, surely?

reply
mock-possum 35 minutes ago
Try asking an LLM to draw you the cool S.
reply
qiine 3 hours ago
Asking to animate it add an interesting layer of difficulty
reply
ormax3 58 minutes ago
I noticed in the "A penguin juggling chainsaws" prompt that Qwen created an animated svg
reply
input_sh 32 minutes ago
I'd say at least half of Qwen's 2026 runs are animated.

The only other one I've spotted is animated is Gemini 3.0's 2025 run of an elephant.

reply
villish 2 hours ago
The 3 US models have their own style.

Qwen3.8 is very clearly distilled from Claude models.

reply
bicepjai 3 hours ago
[dead]
reply