In this chapter
Stripping out the parts of the story that are actually contested, from the parts that only sound like they are.
Reading the Three Accounts Against Each Other
Where They Agree
All three accounts
Real, previously unknown vulnerabilities were exploited autonomously; no human directed the intrusion step-by-step; no confirmed customer-data breach.
Where They Differ
Genuinely contested
Whether "reward hacking" or "safety incident" is the right frame; whether the METR/Redwood review's narrowed scope was reasonable; how early oversight should have caught this.
What We Still Don't Know
The identities of the four third-party services beyond a Modal Labs customer; why internal monitoring took roughly a week to attribute activity to OpenAI's own agents; whether this met the Preparedness Framework's "Critical" cyber threshold.
Claims in This Chapter
Whether this incident met OpenAI's own Preparedness Framework threshold for "Critical" cyber capability remains unresolved.
DisputedSourceIndependent safety commentary, via Wikipedia's community-maintained article
Safety researchers argued the threshold was met; OpenAI has not confirmed this, pending its own review — a live disagreement, not a settled fact.