Chapter 3 of 349%

Part I · The Investigation

03. Research Methodology

In this chapter

How sources were selected, how evidence was weighed against competing accounts, and the classification system every claim in this case study is tagged with.

The Four Source Tiers

  • Tier 1 — primary company statements: OpenAI's own incident write-up and Black Hat USA presentation, and Hugging Face's own published forensic timeline.
  • Tier 2 — independent technical/security research: the Cloud Security Alliance's research note, security-industry reporting (The Hacker News, Cybersecurity News, SecureLayer7), and the scope-limited independent review by METR and Redwood Research.
  • Tier 3 — informed community synthesis: Simon Willison's timeline write-up and Wikipedia's community-maintained article, both used as an index into primary sources, not as primary sources themselves.
  • Tier 4 — general business/tech journalism (Forkast, DEV Community), used for context and reaction, not for load-bearing technical claims.
  • Wherever a claim rests on Tier 3 or 4, we've tried to trace it back to where that source itself got it from — usually one of the two companies' own statements. Two sources repeating the same unverified figure is not treated as two independent confirmations.

Handling Conflicting Claims

  • Conflicting claims get handled explicitly, not smoothed over — Part VI (Three Versions of the Story) lays OpenAI's and Hugging Face's accounts side by side before this case study offers its own read.
  • Where a number varies by source — our own "4.5-day active campaign" figure vs. Hugging Face's "roughly three days undetected inside our systems," for instance — we treat that as two different things being measured, not a contradiction to paper over, and we say so.

The Confidence Tags

  • Every substantive claim carries one of five confidence tags, shown as a colored badge next to the claim itself.
  • 🟢 Confirmed — directly stated by a primary party and not contradicted by the other, or independently verified by named security researchers with technical evidence.
  • 🟡 Strongly Corroborated — reported consistently across primary and independent sources, without one definitive primary citation.
  • 🔵 Inference — a reasonable conclusion this case study draws from the evidence, clearly labeled as our own reading rather than a sourced fact.
  • 🟠 Disputed — the primary parties or independent researchers give materially different accounts, and this case study is not resolving that disagreement on its own authority.
  • ⚫ Unknown — a real gap in the public record: neither company has named the four third-party services whose accounts were compromised, and OpenAI hasn't fully clarified why internal monitoring took roughly a week to attribute the activity to its own agents.
  • This classification exists so a reader never has to guess which sentences here are load-bearing facts and which are our own analysis.