one year on
Ingesting millions of PDFs and why Gemini 2.0 changes everything
A developer's deep dive into using Gemini 2.0 Flash for large-scale PDF ingestion draws Hacker News attention to multimodality and huge context windows.
A Hacker News submission today points to Sergey Filimonov’s January 15 walkthrough on ingesting millions of PDFs using Gemini 2.0 Flash. Commenters say Gemini’s multimodality and large context window make document ingestion simpler and cheaper.
The Hacker News thread draws competing views about software. One fintech engineer reported replacing a dedicated OCR vendor with Gemini, cutting processing time from 12 minutes to 6 seconds while accuracy was like 96% of that of the vendor at significantly lower cost. Another commenter predicting that legacy vendors ‘focusing on a specific type of PDF is going to get obliterated by LLMs.’ Others argued that the non-digital portions of business processes change at people and org speed, not AI speed.
The debate extended to whether LLMs themselves are a commodity or a platform shift. One commenter described the current VC-software-ecosystem as ‘the walking dead,’ while others argued that vendor lock-in would simply shift from OCR companies to LLM providers. The thread sits at 1303 points and 447 comments, suggesting this is a live nerve for developers building data pipelines.
The record
Works in fintech; replaced an OCR vendor with Gemini. Processing time dropped from ~12 minutes to 6 seconds at 96% of that of the vendor and lower cost. Called the developer experience 'stupidly easy'.
Started an LLM doc processing infra company (extend.app); said any legacy vendor focusing on a specific PDF type 'is going to get obliterated by LLMs'. Advised adding chain-of-thought and citations for further improvement.
Argued a smart OCR vendor could shift to using LLMs themselves and sell it with an SLA, but noted the main LLM suppliers are 'purposefully sticking to general-purpose APIs' to avoid starting an us-vs-them war.
One year later — open only if you can handle spoilers
Within months, several OCR vendors announced LLM-enhanced products, while Google released a dedicated Document AI pipeline built on Gemini. The community's debate proved prescient: the shift from bespoke extraction to LLM-based approaches accelerated through 2025, though enterprise adoption remained slower than enthusiasts predicted.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy