Gpt is really good in vision stuff, or at least their MoE seems to be really cohesive. From my experience Claude models can be really good at language but the moment they need to look at a picture and decide why the design is not good what parts need improvement it degrades a lot. My easiest benchmark is giving them a screenshot of a feature in my app and tell it "identify non-normative UI blocks and improve readability and consistency". Sol does a great job at re-structuring the page into composable units that build upon each other and the general looks and feels of the app. Claude tends to over-focus one one part while completely forgetting about the rest or the cohesion as a whole.
Sounds like the kind of UI I like. (Take me back to Windows XP...)
FWIW, the summary-description[1] of "frontend-design"[2] gives me a few things to pick at:
> create polished code
Methinks only if you're using it with a very popular framework like React. What happens if you ask Claude to make the UI in WinForms or MFC?
> high-impact animations
That's bad UX 101 right there: animations in a UI exist as an affordance to the user, and never for its own sake (e.g. macOS's "genie" animation when you minimize a window to the dock exists so the user knows where they can restore the window from). The only people who actually want "high impact animations" in software are salespeople who want something for demo purposes.
> generic system fonts, predictable purple gradients, and cookie-cutter components.
This screams wanting to be different for the sake of standing-out, not because it results in a better software product; users benefit when their software fits-in with platform conventions: if you refuse to use a stock checkbox <input> or <select> drop-down and instead use your own entirely custom component solely for aesthetic reasons then you are producing worse software. There's nothing wrong with system-fonts, but your site will look ugly after your third-party font-host CDN shuts-down and turns into a walking CSRF factory.
> thoughtful typography with unexpected font pairings
The above fragment set my alarm-bells off. Yikes.
> scroll-triggered interactions
Not every web-page should be an Apple.com product brochure page. This is also a fantastic way to make your webpage horribly inaccessible.
------
The SKILL.md itself[3] grinds my gears too:
> Approach this as the design lead at a small studio known for giving every client a visual identity that could not be mistaken for anyone else's.
Claude has no way of knowing what designs are actually unique or not...
> For web designs, the hero is a thesis. Open with the most characteristic thing in the subject's world, in whatever form makes sense for it: a headline, an image, an animation, a live demo, an interactive moment
...this is exactly what everyone else's web-pages look like!
> For calibration: AI-generated design right now clusters around three looks: (1) a warm cream background (near #F4F1EA) with a high-contrast serif display and a terracotta accent; (2) a near-black background with a single bright acid-green or vermilion accent; (3) a broadsheet-style layout with hairline rules, zero border-radius, and dense newspaper-like columns
...I called this out weeks ago[4], lol.
and I could go on. This is all quite painful to read.
------
[1] https://claude.com/plugins/frontend-design
[2] https://github.com/anthropics/claude-plugins-official/tree/m...
[3] https://github.com/anthropics/claude-plugins-official/blob/2...
We have vision models for our pharmacy and I could never imagine taking the latency hit to use a Sol in our robotics, it would be likely 25-50x slower.
I’m evaluating these VLMs to figure out which ones are good enough to auto-annotate my data, so I can fine-tune my detector.
I wrote a bit more about this here: https://x.com/skalskip92/status/2080334344061694429?s=20
LLM needs to setup an image classifier to use as a tool call.
I remember around 2013-15 people were scoffing at uses of deep learning CNNs for various things, because why don't you just use an SVM on HOG features? Or face detection is solved, just use Viola-Jones.
What if you give the benefit of doubt and assume the author knows about alternatives and uses VLMs for their strengths? They use it to auto-annotate training data for regular deep learning models.
https://ieeexplore.ieee.org/document/1674847 G. J. Vanderbrug and A. Rosenfeld, “Two-Stage Template Matching,” IEEE Transactions on Computers, Vol. C-26, No. 4, pp. 384–393, April 1977. DOI: 10.1109/TC.1977.1674847
Another application where template matching would work brilliantly? Car counting in parking lots using satellite imagery.
Source: I did this [1] using OpenCV and template matching. Outperformed "Cars Overhead with Context" models.
https://abcnews.com/International/satellite-data-suggests-co...
For 3, pricing for image tokens was the same as text tokens. Since they don't indicate a difference on 3.7, I would assume the same holds. And as far as I know the number of image tokens is the same for both (depending on the detail level you pick, but it's generally around 1k per image).
So they're about the same, 3.7 is slightly more expensive. At least until the end of the year (when they raise 3.7's pricing).
Anyway, my point was that 3.5 tended to have worse performance and significantly higher costs. 3 and 3.7 are both better and cheaper than 3.5.
In the next bench, Sol looks like it’s correct again but the bboxes are rotated 90 degrees for some reason.
ChatGPT Pro with GPT 5.6-sol: https://chatgpt.com/share/6a834217-ca8c-83e8-a8e8-45d5b8797b...
The puzzle: https://activityvillage-files.s3.eu-west-2.amazonaws.com/s3f...
OP:
> do you think you did a good job there
ChatGPT:
> I spent 15 minutes, emitted several fake-sounding “tracing the puzzle” progress updates, and then gave a confident permutation without showing that I had actually followed the lines correctly. It reads much more like I guessed than solved it. The only part I did well was obeying the “no Python or tools” instruction.
My observations:
1) Sarcastic tone suggests pre-prompting, or frequent (and therefore stored in memories) denigration of the model in past conversations. I'm leaning the former - it sounds like it was instructed to read admission of defeat.
2) The part about "no Python or tools" is setting the model up for failure.
I mean, this task is, for a human, basically a game of "simulate a line following robot in your head". Pretty sure a VLM could solve that if it was allowed to do the same thing. Off the top of my head, an algorithm like:
1. Identify start and end points
2. Foreach start point, follow next pixel minimizing angle, until endpoint is reached.
3. Report answer
It's literally what every human facing this task does.
EDIT:
My attempt - same image, prompt altered to allow for code (but still no search/external checks), solved in 1/5th of the time, correctly, and (going by thinking trace summaries that I don't think show up in shared chats), basically the same way I'd approach it, by tracing the lines, coloring them as it goes.
https://chatgpt.com/share/6a834f76-8240-83ed-acff-0c67af399d...
INB4: I know this is now not a pure vision check, but it really doesn't make much sense to diss models for failing to solve tasks explicitly designed to teach humans to externalize computation that's hard to do in their heads (i.e. kids, crayons, coloring paths).
Still, if such things are becoming a benchmark for tool-less evaluation, it's only a matter of time until the models learn - much like humans learn in school - to follow algorithms mentally, essentially emulating an ad-hoc computer in their head.
With Python, it was able to successfully solve it in 9 minutes: https://chatgpt.com/s/t_6a8350ecddfc81919328caf68de74861
The real pain point is that at work, I use Codex and I'm currently working on a project that involves debugging some polyline topology, very similar to the path following puzzle. The vision is completely useless here.
Your VLM idea sounds good. Theoretically, the inverse problem (generating an SVG of a pelican riding a bike) can also be solved with a VLM that plans out how to draw it, not unlike a human planning out a path for their hand to follow.
When reasoning got introduced a year ago to GPT 5, on average the model performed worse than GPT4-o for short video clip captioning (Ie hallucinating actions that didn’t happen). The old GPT 5 was extremely finicky in terms of fps sample rate.
The other SOTA LLMs (like Gemini Pro) have clearly been optimized for long video understanding, since they can’t see almost anything sub-second (even if you up the frame sampling rate).
Sol is the first model we’ve seen to accurately caption complex sub-second movements (eg woman suddenly turns heard head to right). It’s robust to different fps sample rates so I can only guess that they trained on videos sampled at different fps.
Anecdotal but I've seen it use python to crop, zoom, and "enhance" (fiddle with sharpness and brightness) images to read sections of handwritten census data from the 1800s. Feels like that there might just be a mismatch of capabilities when it comes to straight outputting coordinates but I bet the model is better at actually finding the answer given any tools available. Which I get is a bit of an apples and oranges situation.
I've also tried to use it to identify an old pair of glasses and it didn't stand a chance, so I do think it's not quite there yet when it comes to some vision tasks.
Overall I've been hooked on using agents from different companies for what they are best at (Thanks to Theo). Fable is expensive, but unmatched for planning and top level organization of other agents. Sol is fast, will persistantly go after goals (sometimes to its detriment), and does well with computer use.
Also it seems to be more capable, need to test more, but I think it's at least getting on par and it's fully open-source and open-weights.
Here's some benchmarks:
https://benchlm.ai/compare/gpt-5-6-sol-vs-qwen3-8-max
https://qwen.ai/blog?id=qwen3.8#full-benchmark-table (incredible UI/UX demos)
https://venturebeat.com/technology/qwen3-8-max-arrives-with-...
EDIT: Am I early to the discussion, or is none else using Qwen3.8-max?
In my personal mini benchmark minicpm-v-4.6 scores amazingly well. Its a 0.8B model which runs fine on many consumer hardware.
To use an LLM, you just prompt it with an image + text saying "count the pills in this image".
To use OpenCV, ... you just prompt an LLM with an image + text saying "count the pills in this image, using OpenCV instead of eyeballing it".
(I like to throw in "produce intermediary artifacts so I can see the process" for more difficult tasks; this helps the model avoiding making hallucination-prone leaps and gives more opportunities to self-correct. At a cost of extra time and tokens, of course.)
Using OpenCV without an LLM? Nah, not touching that, I don't have free weekends to waste anymore.
It's not an LLM, it's a custom thing we built. Here's a comprehensive list of support for various notation glyphs: https://www.soundslice.com/help/en/creating/pdf-import/294/s...
For example, a UI / UX professional being asked to appraise a website screenshot may determine that the image in question has "desirable" traits which are inherently not deterministically measurable. Such as, if the interface elements have strong information hierarchy, or if they are deemed to be "fashionable" with current UI trends.
...but that's an example of a UX/usability matter that can be assessed objectively and non-subjectively.
Is 16 px or 14 px a better font-size value for a subheading, in a hypothetical layout? Immediately that kind of decision, where both options are objectively good for 12 px paragraph text, suddenly becomes an issue of taste that cannot be evaluated crudely by an algorithm.
But, if you have a proper well documented design system and you tell the LLM to use the DS and to avoid styling hacks they can generally do it. Even the dumber ones than Sol 5.6.
Of course only if the design is achievable in the design system.
The performance as a general model is indeed really impressive and i think they might actually win compared to fine tuned models.
Their feedback loop of training on user data is incredibly strong. I've learned that lots of accuracy results depends on threshold configs, which llms should be able to dynamically set.
Or the future will develop in llms using fine-tuned models as tools? Inference cost and speed does still seem to be below user expectations.
But for being able to one shot with this accuracy... IMPRESSIVE
I recently used it at grocery stores in a foreign country. Photographed the whole aisle and told it to find Y (detergent, softener, glue, sour cream, whatever), at the same time recommend the best Y for whatever reason. Worked marvelously, including the cases where the object wasn't present and it told me there was nothing useful.
I asked then, can you crop the exact image of how does the item look like and where is it in the aisle - did that perfectly as well.
I will add that all frontier models were fine with such tasks from the early 2024's.
I usually go to https://arena.ai/leaderboard/vision/pareto for a nice overview of current models.
I hope that whatever was lost at GDM in the last few months, didn't include their extra focus on vision capabilities.
I keep waiting for these AI companies to assemble the parts into a great autonomous driving module.
The only quip is the default UI isn't very good. When changing that reaches the top of my priority list, I'll switch it since they don't force you into a walled garden. Plan is to run it through frigate into HomeAssistant and use a UI from them. I've never used frigate before though so it'll be a learning process if plug and play solutions aren't already available
I asked it to recognize and draw the very faint reflection of what I was wearing, visible in only a tiny black part of a very brightly lit poster behind glass.
In addition, the poster itself also happened to contain similar clothing.
You can see the reference images and its output in my writeup here: https://medium.com/@rviragh/gpt-5-6-sol-very-good-image-reco...
While a human can focus on the reflection easily, this is an enormous challenge for a vision model. It's very impressive.
After I increased the game's resolution, I asked it to increase the image's size while keeping the same scale and existing content, and gosh, it constantly keeps getting something wrong no matter what I tell it, even on Sol Max with the $100 Pro subscription.
An organically-grown meat-based pixel-artist could have recreated the image and more within 2-3 days, in exchange for food and shelter.
This 100%
The previews it generated were amazing but wouldn't really be possible as a grid-based tilemap, with lots of clusters and overlaps of elements of varying sizes.
So I just decided to use the preview as a static scrolling background, but it's been a pain to get it to add more content around the edges that still tiles with the existing image at the same scale.
Unlike when Apple says "it's the best iphone we've ever made", LLMs are more or less interchangeable. So "OpenAI's best model" means nothing if "Anthropic wipes the floor with them" or "[open weights model] is 10x cheaper for 1% less quality".
As a reader, it feels like these titles are click bait.
I mean I should hope so, as it is also the latest one
If that is the full quality image given to the model, I think it's not surprising that the model confused with 03/2022.
It seems pretty counter intuitive that we can't do vision significantly better with specialized techniques.
GPT 5.6 Sol was outperformed on all benchmarks by Gemini 3.5 Flash, apart from a single exception (OCR) where Fable was the winner.
Gemini 3.5 Flash not only outperformed GPT 5.6 Sol, but did so at 1/3 of the cost.
Here’s a comparison of the best low-cost models I put together last week. What’s crazy is that Gemini 3.7 Flash is now 50% off on OpenRouter, and this chart doesn’t even account for that discount. https://x.com/skalskip92/status/2088032652301304121?s=20
I run complicated, messy PDFs through these models. 2.5 Pro required a lot of kludgy hacks to get it to fully "see," but from 3.1 pro on I've removed many of them and haven't spotted problems.
3.7 Flash scores better than 3.1 pro on most benchmarks, leading me to believe that even if your OCR requires reasoning to interpret text or data, 3.7 Flash is probably going to be better.
It's not a technical problem, it's a commercial one. If Google can't ship a model to replace the one they deprecated, that tells you everything you need to know about choosing a Gemini model for whatever you're trying to do.
That link shows 3.1 pro listed as deprecated with no replacement model.
these models aren’t successors and barely have a common ancestor, they are independently baked in the training oven and assigned a semantic version randomly by someone trying to show initiative but not trying to do on the toes of the last guy who got promoted first
So 3 pro is outdated and will likely never exit preview
The “flash” and “lite” models are the real “pro” in colloquial ideas of fleshed out and capability, at this point.
they’re better, faster and cheaper, larger context windows keeping up with the industry and more
3.7 Flash is better at coding, sure, but AI is not just for coding.
3.7 Flash is better at coding, sure, but AI is not just for coding.
[0] https://playground.roboflow.com/evals
[1] https://playground.roboflow.com/evals/object-detection
Some other Chinese models are also fast and cheap, but a harder sell in a U.S. production environment.
Important to remember that json schema instructions take precedence over the normal prompt, so move as much into property descriptions as possible.
3.0 flash (not lite) handled it like a champ though, fwiw.
Over the last two weeks, Qwen released two new models. Qwen3.8-Max is totally insane, but it’s only available through the Alibaba Cloud API. I wrote a similar blog covering Qwen3.8-Max: [https://blog.roboflow.com/qwen3-8-max/](https://blog.roboflow.com/qwen3-8-max/)
If you’re looking for something you can run locally, Qwen3.8-27B might be a great option. On Friday, I did a quick comparison between Qwen3.8-Max and Qwen3.8-27B: [https://x.com/skalskip92/status/2088411215441621469?s=20](https://x.com/skalskip92/status/2088411215441621469?s=20)
For example, 3.7 Flash is #1 on MMLU Pro and AA’s agentic spreadsheets/docs benchmark, etc. Yes, beating Fable.
Agentic coding is only one dimension.
My worry is that this is a zero-sum game and when Gemini catches up on coding, it'll regress to the mean in other areas.