Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Mistral has been very good with handwritten ocr, expecting the new models to get better with that across languages.by Utkarsh736
- I've been experimenting with using NuExtract this week on locally OCRing bank statements that don't have a predefined document structure. It's way better than Tesseract or a generic vision-enabled model. It runs great on a single RTX 4090 at the modest throughput I need.
Their hosted, API-based service is something like a third of the cost of this model.
by ad_fontes - Given the latest vibe release's new "follow default" model option, we should see a new coding / general Mistral model, mistral 4, soon.by maelito
- How does this compare to Baidu Unlimited OCR. I've been very impressed with Baidu and it's essentially free to run on a decent computer, other than electricity costs.by Johnny_Bonk
- Where do your documents go?by spiderfarmer
- I won't comment on accuracy, but in internal benchmarks, Mistral OCR is significantly faster than comparable APIs.by ianhawes
- I think people misunderstand the utility of Mistral's OCR. It's not going to beat SOTA models for extraction on edge-case docs, but it's MUCH cheaper and faster and does an excellent job on simple ones. I've been working on converting PDFs to EPUBs and Mistral has been making steady improvements. On a chapter of Bleak House it was able to extract and tag the header, titles, and references at the bottom every time. The only thing it struggled on was line numbers in the right margin which it correctly tagged as "aside text" 3/5 times, but always separated from the core text each time. The important thing to keep in mind is that there's no prompting needed, just upload the PDF and voila!. There's even a batch mode with a 50% discount.by fumeux_fume
- > it's MUCH cheaper and faster and does an excellent job on simple ones.
OCR should be:
1. Privacy-respecting, i.e. running on your own machine without network communications. 2. Fully open-source. 3. Gratis.
The first one is a must, the second is very important for the public interest, and the third one is a nice-to-have.
Mistral does not appear to satisfy even the first-, let alone all three.
by einpoklum - > The important thing to keep in mind is that there's no prompting needed, just upload the PDF and voila!
I'm sorry, noob here. I have a special book that I bought which I can open only inside the Kindle app (Windows/mobile). I have been meaning to screenshot the pages and convert them into a document/PDF. What do I have to do to make it fast? Just upload all the screenshots one by one and tell Mistral "Chat" to OCR them?
by mkbkn - Does anyone know a site that lets you browse examples of input / output pairs?, particularly with layout analysis (bounding boxes of figures, tables, etc).by ks2048
- 1000 Pages / 3.5€ this is expensive as hell. If this is not fastly superior than something like tesseract it is not worth it.by merb
- Even comparing to AWS Textract or Azure Document Intelligence, this is very expensive (more than double)by Oras
- Agreed. Does the GTM team there really sit together like "oh yeah, that sounds reasonable" while being totally beyond typical market prices?by beernet
- The VLM's are so good at complex document understanding now. But you just can't trust them not to invisibly censor sensitive clinical/legal docs, even at the maximally permissive settings.
And the deep learning OCR-only models won't censor, but can and do hallucinate. I've yet to see a 'scan with different approaches and reconcile and say you're not sure if they don't agree' system just work for generic complex documents.
by waldrews - > But you just can't trust them not to invisibly censor sensitive clinical/legal docs
What's an example of this?
by dangoodmanUT - The way my harness set it up is going through 2 or 3 providers, and cross-checking through them, also with plain text extracted if available.
I think we also had a layer that for any quote extracted tested it back if it exists within the original.
If you wanted 100% accuracy, I think it wouldn't be too difficult nowadays to guess the font&size&other text settings, and re render the crucial parts.
by kolinko - For anyone interested, I have an ocr pipeline running on rented GPUs, doing around 1000pages for 0.05-01 usd with around 0.8 seconds per page with full bounding boxes support for grounding.
If you’re interested you can find contact to me via this profile.
3.5 usd/1000 pages is just too expensive…
by piterrro - You should put contact details in your profile :)by x3ro
- Is it European-hosted and fully outside of both CLOUD Act and CCP reach?
Because I'm assuming that's why they get to charge more for the right type of customer.
by vrganj - Can it produce accessible PDF files that will pass accessibility tests? Someone who can do that will make a killing laundering PDFs for academia: by April 26, every PDF, syllabus, and academic document needs to comply with WCAG 2.1 Level AA, which means structural tagging, alt text, and lots of other checklist items that AI could probably generate.by Telemakhos
- Accuracy is truly what people die for in the OCR game. Price isn't the primary function here.. it's an equation of price, accuracy, speed, and in mayn cases regulation.by aliljet
- At this point I lost all hope for Europe playing any significant role in the AI race. If that’s a good or a bad thing I don’t know, but it seems to me like that’s the reality.by king_crimson
- > hope for Europe playing any significant role in the AI race
Ha. It's a large market of the LLMs consumption. So it which will affect the AI race. Just from other perspective than you assumed.
by podgorniy - Even Africa will surpass Europe with its massive data center upstarts breaking ground.by deadbabe
- Mistral is (wisely) are changing their strategy. They switched to hosting open models and they are investing in hosting inference within EU.
Fine tuning an open model to European values is significantly cheaper than making your own model.
by whazor - And US lost the significant role in chip/pc manufacturing. But if a product becomes commodity or utility (which at least for now it seems is the direction), with little lockin, it's not a big deal.
I hope we (EU) don't waste money trying to train local models (which at least some people in Poland try to do), and tries to build our own chips - AI chips have different architecture than regular processor/GPU, and TSMC doesn't need to be winner in this new race.
And if not this, then smaller labs, harnesses and actual application.
by kolinko - Not being a rat in the rat race is the real win.by pbkompasz
- It's not a race. You don't get anything for winning.by kubb
- Much like spaceflight, aerospace or nuclear engineering, you need to retain local talent for national defense purposes. Being 70% as good is still way better than being 100% dependent and heavily leveraged by your opponents.by hadlock
- Commoditization is a beautiful thing. It seems Anthropic and OpenAI are really struggling to maintain much of a moat. Mistral might not be leading but it's not trailing by that much either. And of course the Chinese are doing their own thing quite successfully.
The reality is that the US is betting its economy on data centers at great expense and is exposing its economy to great risk.
Also while geographically a lot of the money and processing power is in the US, the US has been relying on immigration to power its universities and especially AI research has roots all over the globe. India, China, Russia, Europe, etc. AI related know how is finding its way back to all these places.
So, I'm not too worried about the long term here. It will be interesting to see if Anthropic and OpenAI survive their IPOs. Seems like a risky financial bet at this point given the apparent lack of a moat. But if it works out, it will result in a lot of that IPO money being invested in data centers abroad. Including in the EU. Because data residency is a thing here and the EU is too big of a market for companies with that kind of valuation to ignore. We also produce a lot of energy infrastructure (e.g. gas and wind turbines). Those data centers will need lots of power.