one year on
Things we learned about LLMs in 2024
Simon Willison’s annual review captures a year of breakthroughs, price collapses, and unresolved questions about agents and gullibility.
In his annual roundup published on the last day of 2024, Simon Willison catalogued a year of rapid shifts in large language models. The GPT-4 barrier was ‘comprehensively broken’ — 18 organizations now have models ranking above the original GPT-4 on the Chatbot Arena leaderboard, with 70 models total. Prices collapsed: GPT-4o is 12x cheaper than GPT-4, and Google’s Gemini 1.5 Flash 8B costs $0.0375 per million input tokens. Willison calculated that processing all 68,000 photos in his personal library with that model would cost roughly $1.68.
Multimodal vision became common, with nearly every major vendor releasing image, audio, or video models. OpenAI’s Advanced Voice Mode rolled out slowly from August through September, followed by live camera sharing in December; Google shipped a similar feature a day earlier. Willison noted that prompt-driven app generation is now a ‘commodity’ thanks to Claude Artifacts and competing implementations from GitHub, Mistral, and Val Town.
On the Hacker News thread, antirez says most people are exposed to ChatGPT 4o, which is a ‘useless toy’ for programming, while Claude Sonnet 3.5 is helpful if prompted precisely. Another commenter, mvkel, reported building a revenue-generating Swift app using only LLMs despite not being a programmer. The conversation also dug into whether LLMs truly ‘understand’ code or merely pattern-match.
Willison struck a cautious note on agents: he remains skeptical as to their utility based, once again, on the challenge of gullibility. He argued that progress on prompt injection has been minimal since 2022. The post closes with the observation that universal free access to the best models appears over: OpenAI’s $200/month ChatGPT Pro is the only way to access o1 Pro.
The record
Says most people are exposed to ChatGPT 4o which is a 'useless toy' for programming, but Claude Sonnet 3.5 is helpful if prompted precisely.
Claims to have built and shipped a Swift app generating $10,200 MRR exclusively using LLMs, despite not being a programmer.
One year later — open only if you can handle spoilers
Willison’s skepticism about agents proved prescient; by mid-2026, no broadly adopted agent framework had emerged. Meanwhile, the price collapse continued, with inference costs dropping another order of magnitude. The GPT-4 barrier fell so thoroughly that the phrase stopped being used.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy