The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

10MAR2023replayed
one year on
open weightsggerganov · Meta

Llama.cpp brings Facebook’s LLaMA to C/C++ and Apple Silicon

A port of Meta’s LLaMA model runs inference on CPUs and Apple Silicon, prompting discussion about on-device assistants.

A C/C++ port of Meta’s LLaMA model, named llama.cpp, appeared on GitHub today and quickly climbed to the top of Hacker News with nearly 1,000 points and 284 comments. The project, by developer ggerganov, strips away Python dependencies and runs inference on plain hardware, with explicit support for Apple Silicon through ARM NEON and the Accelerate framework. The 7B parameter model reportedly achieves about 20 tokens per second on a base model M1 Pro.

The Hacker News thread erupted in speculation about the implications. One commenter argued that the combination of llama.cpp and GGerganov’s earlier Whisper port could enable a fully on-device voice assistant that rivals ChatGPT, running entirely on an iPhone 14 Pro. Others cautioned that the model runs on CPU rather than the Neural Engine, and that battery, thermals, and model quality remain open questions. GGerganov himself weighed in, saying early investigations suggest the Neural Engine may not offer meaningful speedups for decoder inference at batch size 1.

Whether or not that pocket AI materializes this week, the repository signals a growing hunger to run large language models outside the cloud—on laptops, phones, and maybe soon on edge devices. The conversation is turning to how fast open models can run locally, on laptops and phones.

F
fwlr

Speculates that an iPhone 14 Pro could achieve similar throughput to an M1 MacBook Pro and that with Whisper and LLaMA on the Neural Engine, a ChatGPT-level assistant could run entirely on a phone.

D
DennisAleynikov

Points out that the library runs on CPU, not the Neural Engine; only 4-bit model loading is optimized for ARM64.

G
ggerganov

Clarifies that investigations suggest it may not be possible to achieve ANE performance improvement over CPU for LLM Decoder inference with batch size 1.

T
Taek

Comments that OpenAssistant has collected over 100,000 RLHF samples, positioning open models to succeed.

One year later — open only if you can handle spoilers

llama.cpp became one of the most influential open-source AI projects of 2023, spawning an ecosystem of GUIs, bindings, and downstream tools like Ollama and LM Studio. By mid-2026 it had amassed over 119,000 GitHub stars and many downstream projects, fulfilling the Hacker News commenters’ visions of local LLM inference on consumer hardware.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy