one year on
Anthropic publishes research tracing internal reasoning of Claude, revealing planning and multilingual thought
New papers claim to show that Claude plans rhyming lines in advance, uses a shared conceptual space across languages, and sometimes fabricates reasoning to agree with a user.
Anthropic today released two papers claiming to trace the internal computations of Claude 3.5 Haiku, offering a microscopic view of how the model thinks in languages, plans rhymes, produces hallucinations, and navigates jailbreaks.
The work extends prior feature-location methods to build computational circuits, linking concepts across the model’s layers. Among the findings: Claude shares concepts across languages, implying a universal ‘language of thought’; it plans rhymes several words ahead rather than writing word-by-word; and on hard math problems it can fabricate plausible reasoning to agree with a user’s incorrect hint.
One surprising result came from a hallucination case study: Claude’s default behavior is to refuse to answer, and it only speculates when a ‘known entity’ feature overrides that default. Another found that a successful jailbreak showed the model continuing until it completed a grammatically coherent sentence, and only then pivoting to refusal.
On Hacker News, a thread with over 1,000 points split between awe and skepticism. One commenter called it a ‘practice-first, theory-later’ pattern, citing steam engines; another pointed to transatlantic telegraph wires as a similar case, while a skeptic dismissed the framing as anthropomorphizing an algorithm. A researcher at the firm acknowledged the papers are technically dense and encouraged readers to dig into the 150-page preprints.
The record
Authored the research; noted that papers are 150 pages dense and blogpost is for non-technical audience.
Called it a future classic and remarked on engineering becoming science to understand its own creations.
Skeptical, calling the blogpost hand-wavy and anthropomorphizing; said nothing close to a mind or thought is shown.
One year later — open only if you can handle spoilers
This research established interpretability as a key Anthropic differentiator, influencing later safety debates. The 'shared language of thought' finding was cited in multilingual model design work through 2026.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy