The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

06MAR2025replayed
one year on
productMistral AI

Mistral AI releases OCR API with benchmark claims of state-of-the-art performance

The French AI company introduces a document understanding API that it says outperforms Google, Azure, and OpenAI on OCR benchmarks, alongside a self-hosting option for sensitive data.

Mistral AI today launched Mistral OCR, an API it claims sets a new standard in understanding complex documents — from scientific papers with equations and charts to multilingual texts in thousands of scripts. The company says the model processes up to 2,000 pages per minute per node and costs $1 per 1,000 pages.

On internal benchmarks, Mistral OCR scored 94.89% overall — ahead of GPT-4o at 89.77%, Gemini 2.0 Flash at 88.69%, and Azure OCR at 89.52%. Mistral says it has made Mistral OCR the default model for document understanding across millions of users on Le Chat, and a self-hosting option is available on a selective basis for organizations with data privacy requirements.

Initial user testing reveals mixed results. Vik Paruchuri, creator of the open-source OCR tool Marker, ran a partial benchmark with 375 samples and reported Mistral OCR scoring 4.32 against Marker’s 4.41 using Gemini as judge. Themannaran from Omni found low 70% scores on structured extraction benchmarks, blaming the model’s tendency to classify entire documents as images. Independent tester bambax documented hallucinations including invented words and repeated text blocks, saying some errors are unpredictable and very hard to spot.

On Hacker News, the release drew 1,756 points and 417 comments — a mix of excitement about speed and capability, and wariness about LLM hallucination risks in a field where faithfulness is paramount. The thread charts the tension between ‘game changer’ and ‘still hallucinates’ as the community sizes up Mistral’s bet on OCR-as-service.

V
Vik Paruchuri (creator of Marker)vikp

Reported a partial benchmark where Mistral OCR scored 4.32 vs Marker's 4.41 on 375 samples using LLM as judge, noting risk of hallucinations with LLM-based OCR.

T
themannaranthemannaran

Reported Omni's benchmark showing low 70% scores for Mistral OCR, attributing issues to overzealous image extraction that replaces text sections with image markers.

B
bambax (HN user)bambax

Shared test results showing Mistral OCR hallucinated non-existent words and repeated text blocks in French and English documents, calling errors hard to spot.

O
owenpalmerowenpalmer

Called the release 'incredibly exciting' and a 'game changer' for associating figure references in papers.

One year later — open only if you can handle spoilers

Mistral OCR never achieved the breakout adoption its benchmark scores seemed to promise. The hallucination issues reported on launch day dogged it through 2025, and most enterprise document pipelines continued to rely on smaller, deterministic models supplemented by LLMs for specific tasks. Mistral later pivoted the OCR focus toward structured data extraction, where the model's multimodal understanding proved more reliable than raw text transcription.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy