The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

20MAY2023replayed
one year on
researchDonald Knuth · ChatGPT · Stephen Wolfram

Don Knuth publishes transcript of ChatGPT experiment, highlighting AI's confident errors

The legendary computer scientist reveals a 20-question test he gave ChatGPT in April, with answers that are often plausible but wrong, sparking debate about the limits of large language models.

Donald Knuth, the legendary computer scientist and author of “The Art of Computer Programming,” has released a transcript of a 20-question experiment he conducted with ChatGPT on April 7, 2023. The results, published on his Stanford University page today, show the AI delivering confident but often erroneous answers to a range of questions designed to probe its limits.

Knuth’s questions included prompts like “Tell me what Donald Knuth says to Stephen Wolfram about ChatGPT” and “Why does Mathematica give the wrong value for Binomial[-1,-1]?” In response, ChatGPT produced lengthy, plausible-sounding answers that were factually incorrect or evasive. For example, it claimed the sun would be directly overhead in Kagoshima, Japan on July 4, 2023, and it described a non-existent ballet version of “Flower Drum Song.” The AI also failed to write a sonnet that is also a haiku, ultimately producing a non-sonnet, and it generated an essay avoiding the word “the” that actually included the word.

The transcript has ignited discussion on Hacker News, where the thread has amassed 927 points and 622 comments. Commenters are divided: some see the experiment as a clear demonstration of the AI’s limitations, comparing its erratic performance to self-driving car failures, while others argue the technology is still improving rapidly. One commenter noted that ChatGPT is “a very capable and convincing liar,” while another pointed out that the AI’s inability to recognize its own contradictions reveals a fundamental lack of understanding. The conversation reflects a growing unease about trusting large language models for tasks where correctness matters.

H
Hacker News community@somenameforme

thread debaters whether ChatGPT's flaws are inherent to neural nets, drawing parallels to self-driving cars

H
Hacker News community@vijucat

argues that ChatGPT's training on all available data still yields deficiencies, suggesting a need for a fundamental leap forward

H
Hacker News community@weaksauce

calls ChatGPT 'a very capable and convincing liar'

One year later — open only if you can handle spoilers

In the years that followed, Knuth's experiment became a frequently cited example of the 'hallucination' problem in large language models. While later models like GPT-4 showed improvements on some of these specific questions, the underlying issue of confident falsehoods persisted, shaping ongoing research into AI alignment and factual reliability.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy