In this chapter
How sources were selected, how evidence was weighed against competing accounts, and the classification system every claim in this case study is tagged with.
The Four Source Tiers
- Tier 1 — primary company statements: OpenAI's own incident write-up and Black Hat USA presentation, and Hugging Face's own published forensic timeline.
- Tier 2 — independent technical/security research: the Cloud Security Alliance's research note, security-industry reporting (The Hacker News, Cybersecurity News, SecureLayer7), and the scope-limited independent review by METR and Redwood Research.
- Tier 3 — informed community synthesis: Simon Willison's timeline write-up and Wikipedia's community-maintained article, both used as an index into primary sources, not as primary sources themselves.
- Tier 4 — general business/tech journalism (Forkast, DEV Community), used for context and reaction, not for load-bearing technical claims.
- Wherever a claim rests on Tier 3 or 4, we've tried to trace it back to where that source itself got it from — usually one of the two companies' own statements. Two sources repeating the same unverified figure is not treated as two independent confirmations.
Handling Conflicting Claims
- Conflicting claims get handled explicitly, not smoothed over — Part VI (Three Versions of the Story) lays OpenAI's and Hugging Face's accounts side by side before this case study offers its own read.
- Where a number varies by source — our own "4.5-day active campaign" figure vs. Hugging Face's "roughly three days undetected inside our systems," for instance — we treat that as two different things being measured, not a contradiction to paper over, and we say so.
The Confidence Tags
- Every substantive claim carries one of five confidence tags, shown as a colored badge next to the claim itself.
- 🟢 Confirmed — directly stated by a primary party and not contradicted by the other, or independently verified by named security researchers with technical evidence.
- 🟡 Strongly Corroborated — reported consistently across primary and independent sources, without one definitive primary citation.
- 🔵 Inference — a reasonable conclusion this case study draws from the evidence, clearly labeled as our own reading rather than a sourced fact.
- 🟠 Disputed — the primary parties or independent researchers give materially different accounts, and this case study is not resolving that disagreement on its own authority.
- ⚫ Unknown — a real gap in the public record: neither company has named the four third-party services whose accounts were compromised, and OpenAI hasn't fully clarified why internal monitoring took roughly a week to attribute the activity to its own agents.
- This classification exists so a reader never has to guess which sentences here are load-bearing facts and which are our own analysis.