The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

06JUL2025replayed
one year on
communityApple · BlueFalconHD

Developer decrypts Apple Intelligence safety filters, publishes blocklists on GitHub

A reverse engineer extracts and releases the encrypted safety override files Apple uses to filter model inputs and outputs, revealing blocklists for slurs, violence, and brand capitalisation rules.

A developer using the handle BlueFalconHD has published decrypted safety override files from Apple Intelligence models on GitHub, exposing the blocklists and regular expressions Apple uses to filter model inputs and outputs. The repository, uploaded on June 28, 2025, contains decrypted JSON files that reveal exact phrases and regex patterns for rejecting, removing, or replacing content deemed unsafe.

The files include categories such as slurs, death-related terms, and brand capitalisation rules. One example from the code intelligence base shows regexReject entries for terms like ‘bitch’ and ‘dago’, alongside reject fields for phrases containing ‘death’. The developer detailed the reverse-engineering process, which involved attaching LLDB to Apple’s safety inference provider and extracting an encryption key the framework calls ‘Obfuscation’.

On Hacker News, the thread quickly turned to larger debates about performative safety and the euphemism treadmill, with commenters noting that the blocklists omit newer coded terms like ‘unalive’. Skepticism about the effectiveness of such filters ran through the discussion, with some arguing that no static list can keep pace with evolving language.

T
trebligdivad

Called out odd combinations such as death-related blocks mixed with Apple brand capitalisation rules.

G
grues-dinner

Noted the absence of 'unalive' in the blocklists, suggesting the filters are performative since platforms know what that term means.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy