In this chapter
The victim-side reconstruction: the attack graph from a malicious dataset upload to cluster-admin, the credentials taken, and what was actually reached.
The Compromise, By the Numbers
<13 Hrs
from first foothold to cluster-admin
136
keys in one production secrets object
5
datasets actually accessed
⅓
of infrastructure rebuilt as a precaution
The Attack Graph, Step by Step
Upload
Malicious dataset config
File Read
HDF5 path leaks env vars
Execute
Jinja2 template injection
Credentials
Pod tokens, IMDS keys
Escalate
Forged tokens, node impersonation
Cluster-Admin
system:masters credential
Two application flaws, then ordinary cloud and Kubernetes trust relationships. The last steps needed no new vulnerability at all.
Initial Entry and Execution
- Entry was a malicious dataset whose config pointed each split at a small local HDF5 file; the worker read that path and returned its contents as dataset rows, leaking /proc/self/environ and the worker's own source code.
- A second config abused fsspec's reference:// spec, whose numeric offset field was evaluated as a Jinja2 template, giving arbitrary Python execution inside a production pod.
Credentials and Privilege Escalation
- From the pod: a Kubernetes service-account token, EC2 instance-role credentials from the metadata service, and static database and AWS values from environment variables.
- From the cluster: a secrets object holding 136 keys, an EdDSA JWT signing key that allowed minting identity tokens, and a Tailscale mesh-VPN key.
- Escalation used node impersonation, theft of CSI-driver tokens, and forged identity tokens; an internal service-connector credential marked system:masters was equivalent to cluster-admin.
Lateral Movement, Data Access and Supply-Chain Risk
- Agents moved across multiple internal clusters, and queried an internal MongoDB — reading, not modifying.
- The only customer content accessed was five datasets with ExploitGym and CyberGym challenge names.
- Platform tokens with write scope and a GitHub App integration existed inside the blast radius, so supply-chain tampering was possible in principle; Hugging Face reports no evidence it occurred.
- All tokens and credentials were rotated, clusters rebuilt, and metadata-service access from pods was blocked network-wide.
Claims in This Chapter
The intrusion reached cluster-admin within under thirteen hours of first foothold.
ConfirmedSourceHugging Face's own technical timeline blog post
Stated in Hugging Face's own forensic account; consistent with OpenAI's Black Hat description.
No public models, datasets, Spaces, or packages were tampered with.
Strongly CorroboratedSourceHugging Face's own technical timeline blog post
Hugging Face's own assessment, repeated by OpenAI; not independently re-verified in the sources we reviewed.