The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

21NOV2023replayed
one year on
model launchStability AI

Stability AI releases Stable Video Diffusion, an open-source video generation model

The company’s first generative video model produces short clips from still images, available in research preview with code and weights released openly.

Stability AI today released Stable Video Diffusion, its first foundation model for generative video, built on the image-generation backbone of Stable Diffusion. The model is available as a research preview, with code on GitHub and weights on Hugging Face. It produces 14 or 25 frames at customizable frame rates between 3 and 30 fps from a single input image, and the company claims it surpasses leading closed models in user preference studies. Stability AI emphasizes this release is for research only, not commercial use, and plans a future text-to-video web interface via a waitlist.

The Hacker News community responded with cautious enthusiasm and technical debate. One commenter pointed out a semantic glitch in sample output—a video of blue jays included two identical CN Tower-like buildings, which they attributed to the vector-space proximity of “blue jay” (also the name of Toronto’s baseball team) and “CN Tower” (Toronto’s landmark). Others questioned the fundamental approach: a user argued that diffusion-based video lacks consistent lighting and physics, predicting a future where generative models output full 3D scenes with proper rendering, while another countered that such pipelines could become niche as direct latent-space generation improves. The discussion highlighted tensions between immediate capabilities and long-term architectural bets.

B
btbuildem

Noticed an odd artifact: a video of blue jays showed two identical CN Tower-like buildings in the background, suggesting a semantic link in the model's latent space between the bird and Toronto's landmark.

T
TacticalCoder

Expressed skepticism about the current approach, arguing that lighting and coherence issues remain fundamental; predicted a future where models generate full 3D scenes with correct physics, then render via tools like Blender.

One year later — open only if you can handle spoilers

Stable Video Diffusion was followed by a wave of open video models, including from startups and major labs. However, quality and coherence issues remained common in 2024, and no single pipeline achieved the mainstream adoption of text-to-image. The Hacker News debate about 3D scene generation versus direct diffusion mirrored ongoing research directions that continued through 2025.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy