The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

06MAR2024replayed
one year on
communityAnswer.AI · Tim Dettmers · Hugging Face

Answer.AI releases open-source system to fine-tune 70B language models on consumer GPUs

A new technique combining FSDP and QLoRA enables training of 70-billion-parameter models on two 24GB gaming GPUs, potentially democratizing large model fine-tuning.

Answer.AI, the R&D lab founded by Jeremy Howard and Eric Ries, today released an open-source system that can fine-tune a 70-billion-parameter language model on a desktop computer with two standard gaming GPUs, such as an RTX 3090 or 4090. The system combines Meta’s Fully Sharded Data Parallel (FSDP) with Tim Dettmers’ QLoRA technique, which uses 4-bit quantization and low-rank adapters to dramatically reduce memory requirements.

The project, developed in collaboration with Dettmers at the University of Washington and Hugging Face engineers Titus von Koeller and Sourab Mangrulkar, addresses a key bottleneck: until now, training large models required expensive data-center GPUs like H100s. The team overcame issues including metadata synchronization across shards and memory spikes from quantization by modifying the bitsandbytes library and creating a custom LoRA implementation. They also integrated techniques like gradient checkpointing, CPU offloading, and Flash Attention 2.

The system enables training a 70b model on two 24GB GPUs, a milestone previously requiring far more expensive hardware. The release includes a standalone script that demonstrates the entire fine-tuning pipeline, from loading the quantized model to saving the adapted weights.

J
Jeremy Howard@jph00

Howard, from Answer.AI, said this was the most requested feature from open-source model builders. He defended the project's risk as comparable to releasing 'a new model of pen or pencil', arguing that his response was accurate based on his experience.

T
Teknium

The creator of the OpenHermes models said: "With this capability we can take huge models to new heights locally, and gigantic, hundreds of billions of parameter models are now accessible by small labs."

One year later — open only if you can handle spoilers

The system became a foundation for subsequent fine-tuning tools. By mid-2026, multi-GPU QLoRA was standard in open-source libraries, and the consumer-hardware threshold for 70B models dropped further with quantization advances and memory-efficient kernels.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy