one year on
Answer.AI releases open-source system to fine-tune 70B language models on consumer GPUs
A new technique combining FSDP and QLoRA enables training of 70-billion-parameter models on two 24GB gaming GPUs, potentially democratizing large model fine-tuning.
Answer.AI, the R&D lab founded by Jeremy Howard and Eric Ries, today released an open-source system that can fine-tune a 70-billion-parameter language model on a desktop computer with two standard gaming GPUs, such as an RTX 3090 or 4090. The system combines Meta’s Fully Sharded Data Parallel (FSDP) with Tim Dettmers’ QLoRA technique, which uses 4-bit quantization and low-rank adapters to dramatically reduce memory requirements.
The project, developed in collaboration with Dettmers at the University of Washington and Hugging Face engineers Titus von Koeller and Sourab Mangrulkar, addresses a key bottleneck: until now, training large models required expensive data-center GPUs like H100s. The team overcame issues including metadata synchronization across shards and memory spikes from quantization by modifying the bitsandbytes library and creating a custom LoRA implementation. They also integrated techniques like gradient checkpointing, CPU offloading, and Flash Attention 2.
The system enables training a 70b model on two 24GB GPUs, a milestone previously requiring far more expensive hardware. The release includes a standalone script that demonstrates the entire fine-tuning pipeline, from loading the quantized model to saving the adapted weights.
Howard, from Answer.AI, said this was the most requested feature from open-source model builders. He defended the project's risk as comparable to releasing 'a new model of pen or pencil', arguing that his response was accurate based on his experience.
The creator of the OpenHermes models said: "With this capability we can take huge models to new heights locally, and gigantic, hundreds of billions of parameter models are now accessible by small labs."
One year later — open only if you can handle spoilers
The system became a foundation for subsequent fine-tuning tools. By mid-2026, multi-GPU QLoRA was standard in open-source libraries, and the consumer-hardware threshold for 70B models dropped further with quantization advances and memory-efficient kernels.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy