Nothing special about this model for overly-detailed work like mine.
It's been a while since I last tested (and discontinued my subscription), but the "pro" models from OpenAI dominate. Not surprising, given the price difference, but it would be nice if an OCR-specific model could perform better. It's worth mentioning that even the highest-end models do a pretty poor job with intricate text like mine.
Mistral's one advantage is that Anthropic now flags OCR, because they don't allow anything that could be considered "reproduction", even of work for which you own the copyright. So my new workflow is Mistral OCR for the actual OCR, followed by a proofreading pass by Claude (which is allowed). Claude is obviously more expensive, but it caught entirely hallucinated sentences created by Mistral OCR 4.0, so I was glad for the backup check.
That said, models have sometimes surprising weaknesses and a model could be terrible overall but magically work for one type of document.
I'd love to see some advancements in traditional OCR based on ideas and concepts we've learned from newer "OCR-like" models since traditional OCR is drastically cheaper.
I haven't been impressed with any of Mistral's models. They obviously realized that they couldn't compete at the frontier so they decided to go for smaller focused models but even those have not been that good.
We moved away from Cursor but I was looking for a model that would help with FIM (fill-in-middle) multiline autocompletion and people were recommending Mistral's Codestral. We gave it a shot and it was lackluster at best.. Even Google's Gemini did a significantly better job than Codestral.
Ultimately Opus-class models got good enough and I don't do much manual coding anymore.
As another user pointed out, it's surprisingly random (task-specific). Llama Scout outperformed Gemini Flash 2.5 on a benchmark I built at the time. I didn't include an OCR models.
Mistral might indeed be the best OCR-specific model for my task, now that you ask. Funny. It's so bad at my work that I didn't register it might be the best in its category. This is just based on vibes from my single scan.
"Winning the race" doesn't give you that in any meaningful capacity. It gives you, at the absolute most, a temporary window where that's the case. See: nuclear weapons.
Europe already has an innovation problem that's already causing structural economic instabilities which Germany has been struggling (and lately failing) to prop up.
As much as it pains me to say this, AI is already a tech revolution and it seems like Europe is just ignoring it. There's more innovation in 3 blocks in downtown San Francisco than the entire continent of Europe.
QED being first didn't grant exclusivity.
The premise wasn't that AGI isn't useful. There are actually layers to the metaphor where first-mover advantage of AGI is even less meaningful than it was for nuclear weapons, but I leave those as an exercise to the reader to discover.
This is the weakest part of the argument. Absolutely no reason to believe the second inventor will be 10x faster.
Does this ever happen? Even in traditional manufacturing, the second “inventor” starts behind and has to improve their own process to surpass the first.
With a lot of Chinese nationals in American companies, and an American corporation owning DeepMind, and a lot of people very upset with all of them at the same time, this is very messy.
Say what you will, it’s been arguably the most successful fundraising campaign in history.
Race dynamics increases p(doom) for everyone.
The non-doom scenarios include "utopia for all", and "power flows to investors, not citizens of whichever nation the winning model's corp. was registered in".
Independently, "oh look all the investors went bankrupt" can happen in both "doom" and "normal technology" timelines.
I’m glad Mistral is working on useful solutions.
OpenAI/Anthropic is like a retarded little sibling chasing “AGI” and giving up on rich media and other modalities.
OpenAI/Anthropic is the worst of the mainstream AI.
It goes:
1. Gemini
2. Vidu
3. Le Chat (Mistral)
4. DeepAI
5. [insert MiniMax provider]
If you’re interested you can find contact to me via this profile.
3.5 usd/1000 pages is just too expensive…
Google documentai costs 1.5$ per 1000 which is probably better in quality and speed.
I’ve been there, implementing a way to linearise text from a document with pages with 1, 2 or 3 columns, some of them in landscape is a nightmare. And that’s not even considering equations.
In the end it’s way easier to use a specialised model, trained by other people to do exactly what I need.
So, maybe it can be tuned for your usecase but with that kind of investment €3.5 for 1000 pages is a bargain...
There's a big gulf between "it's as if a human being had transcribed it and reconstructed the original document" and "so completely broken it can't be used for anything".
Imagine you have a cheap and 99% accurate ocr. The other 1% you can detect and apply more powerful (more accurate but slower and more expensive) ocr method. What would you use? At scale these things add up.
Their hosted, API-based service is something like a third of the cost of this model.
And the deep learning OCR-only models won't censor, but can and do hallucinate. I've yet to see a 'scan with different approaches and reconcile and say you're not sure if they don't agree' system just work for generic complex documents.