one year on
Stability AI releases open-source language model StableLM
The company behind Stable Diffusion launches 3B and 7B parameter language models under a CC BY-SA-4.0 license, with larger versions to come.
Stability AI today released StableLM, an open-source language model available in 3 billion and 7 billion parameter versions, with plans for models up to 65 billion parameters. The models are available under a CC BY-SA-4.0 license for commercial and research use, while fine-tuned versions are released under a non-commercial CC BY-NC-SA 4.0 license. Initially, these fine-tuned models will use a combination of five recent open source datasets for conversational agents: Alpaca, GPT4All, Dolly, ShareGPT, and HH.
The company trained StableLM on a new experimental dataset based on The Pile, but three times larger at 1.5 trillion content tokens. The models support a 4096-token context window, longer than the 2048 tokens in LLaMA. However, the release drew criticism for lacking benchmark results. Hacker News users who independently evaluated the model reported underwhelming performance: one found the 3B base model scored 25.6% on MMLU, far below similarly sized models, while another noted the fine-tuned 7B model gave incorrect answers to simple math problems.
Stability AI emphasized transparency and accessibility, saying the models can run on local devices and are designed to support users rather than replace them. The community reaction on Hacker News was mixed, with some praising the open-source effort and others questioning the model’s quality and the restrictive license on fine-tuned versions. The thread reached 1502 points and 497 comments, reflecting strong interest but also skepticism about the model’s readiness.
The record
Criticized Stability AI for not releasing evaluation results, and shared MMLU benchmarks showing StableLM 3B scored 25.6% accuracy, far below flan-t5-xl's 49.3%.
Reported that the fine-tuned 7B model gave incorrect answers to simple math and coding prompts, e.g. saying 12+12 equals 18.
Ran benchmark evaluations and found performance between GPT-2 774M and Pythia 410M, calling it 'yikes' and noting it underperforms other models.
Highlighted pros like 4096 context width and RL-tuned models, but noted the lack of benchmarks and somewhat restrictive license as cons.
One year later — open only if you can handle spoilers
StableLM never became a major open-source contender, overshadowed by LLaMA-based models and later releases like Mistral. The initial alpha versions were quickly superseded, and Stability AI shifted focus to image and video models. The HN skepticism about benchmark-less launches proved prescient as the community increasingly demanded rigorous evaluation.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy