The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

28FEB2024replayed
one year on
researchMicrosoft Research

Researchers unveil 1.58-bit LLM with ternary weights

BitNet b1.58 matches full-precision models in perplexity and end-task performance while using significantly less memory and energy, marking a potential shift in LLM scaling.

Researchers have introduced BitNet b1.58, a 1-bit LLM variant where every parameter is ternary — taking values from the set {-1, 0, 1}. In a paper posted to arXiv on February 27, the team claims the model matches the perplexity and end-task performance of full-precision (FP16 or BF16) Transformer LLMs of the same size and training tokens, while being significantly more cost-effective in latency, memory, throughput, and energy consumption.

The work builds on the earlier BitNet architecture and represents a push toward ultra-low-precision models that defines a new scaling law and recipe for training new generations of LLMs. The authors argue that the ternary approach enables a new computation paradigm and paves the way for specialized hardware optimized for 1-bit LLMs.

Hacker News erupted with discussion, with many commenters drawing connections to earlier work on binary and ternary neural networks from the 2010s, including IBM’s TrueNorth chip and BinaryConnect. Some expressed cautious optimism, noting that the technique’s performance on 3B+ parameter models is promising but that results on larger models like 70B remain unreported. Others debated the theoretical underpinnings, with one commenter suggesting that ternary weights effectively force the network to focus on connectivity rather than precise values, echoing the Lottery Ticket Hypothesis. The thread, still active, reflects a mix of excitement and technical skepticism.

C
cs702

Expressed shock that existing LLM weights can be replaced with ternary values and that dot products become elementwise additions, with significant gains on existing hardware and even greater potential in custom hardware.

P
paul_mk1

Noted that ternary weights were explored in 2016 with BinaryConnect and IBM's TrueNorth chip, and suggested the authors missed prior art; also mentioned that techniques can go down to 0.68 bits per weight.

V
vessenes

Initially skeptical, arguing that early training only coarsely tunes weights, but later corrected himself after re-checking the paper's perplexity tables, which show better results for larger models.

P
pclmulqdq

Questioned why the paper does not report perplexity for a 70B model, suggesting the result may not have been good enough.

One year later — open only if you can handle spoilers

BitNet b1.58 did not immediately displace traditional LLMs, but it spurred further research into extreme quantization. Within months, several follow-up papers from other labs explored ternary and even binary weights for LLMs, though training stability at scale remained a challenge. Custom hardware for 1-bit LLMs never materialized in a commercial product by mid-2026.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy