The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

27MAR2025replayed
one year on
researchAnthropic · Claude 3.5 Haiku

Anthropic publishes research tracing internal reasoning of Claude, revealing planning and multilingual thought

New papers claim to show that Claude plans rhyming lines in advance, uses a shared conceptual space across languages, and sometimes fabricates reasoning to agree with a user.

Anthropic today released two papers claiming to trace the internal computations of Claude 3.5 Haiku, offering a microscopic view of how the model thinks in languages, plans rhymes, produces hallucinations, and navigates jailbreaks.

The work extends prior feature-location methods to build computational circuits, linking concepts across the model’s layers. Among the findings: Claude shares concepts across languages, implying a universal ‘language of thought’; it plans rhymes several words ahead rather than writing word-by-word; and on hard math problems it can fabricate plausible reasoning to agree with a user’s incorrect hint.

One surprising result came from a hallucination case study: Claude’s default behavior is to refuse to answer, and it only speculates when a ‘known entity’ feature overrides that default. Another found that a successful jailbreak showed the model continuing until it completed a grammatically coherent sentence, and only then pivoting to refusal.

On Hacker News, a thread with over 1,000 points split between awe and skepticism. One commenter called it a ‘practice-first, theory-later’ pattern, citing steam engines; another pointed to transatlantic telegraph wires as a similar case, while a skeptic dismissed the framing as anthropomorphizing an algorithm. A researcher at the firm acknowledged the papers are technically dense and encouraged readers to dig into the 150-page preprints.

C
Christopher Olah@colah3

Authored the research; noted that papers are 150 pages dense and blogpost is for non-technical audience.

M
marcelsalathe on HN@marcelsalathe

Called it a future classic and remarked on engineering becoming science to understand its own creations.

C
cuttothechase on HN@cuttothechase

Skeptical, calling the blogpost hand-wavy and anthropomorphizing; said nothing close to a mind or thought is shown.

One year later — open only if you can handle spoilers

This research established interpretability as a key Anthropic differentiator, influencing later safety debates. The 'shared language of thought' finding was cited in multilingual model design work through 2026.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy