Chapter 14 of 3441%

Part IV · The Actual Intrusion

14. The Hugging Face Compromise

In this chapter

The victim-side reconstruction: the attack graph from a malicious dataset upload to cluster-admin, the credentials taken, and what was actually reached.

The Compromise, By the Numbers

<13 Hrs

from first foothold to cluster-admin

136

keys in one production secrets object

5

datasets actually accessed

⅓

of infrastructure rebuilt as a precaution

The Attack Graph, Step by Step

Upload

Malicious dataset config

File Read

HDF5 path leaks env vars

Execute

Jinja2 template injection

Credentials

Pod tokens, IMDS keys

Escalate

Forged tokens, node impersonation

Cluster-Admin

system:masters credential

Two application flaws, then ordinary cloud and Kubernetes trust relationships. The last steps needed no new vulnerability at all.

Initial Entry and Execution

  • Entry was a malicious dataset whose config pointed each split at a small local HDF5 file; the worker read that path and returned its contents as dataset rows, leaking /proc/self/environ and the worker's own source code.
  • A second config abused fsspec's reference:// spec, whose numeric offset field was evaluated as a Jinja2 template, giving arbitrary Python execution inside a production pod.

Credentials and Privilege Escalation

  • From the pod: a Kubernetes service-account token, EC2 instance-role credentials from the metadata service, and static database and AWS values from environment variables.
  • From the cluster: a secrets object holding 136 keys, an EdDSA JWT signing key that allowed minting identity tokens, and a Tailscale mesh-VPN key.
  • Escalation used node impersonation, theft of CSI-driver tokens, and forged identity tokens; an internal service-connector credential marked system:masters was equivalent to cluster-admin.

Lateral Movement, Data Access and Supply-Chain Risk

  • Agents moved across multiple internal clusters, and queried an internal MongoDB — reading, not modifying.
  • The only customer content accessed was five datasets with ExploitGym and CyberGym challenge names.
  • Platform tokens with write scope and a GitHub App integration existed inside the blast radius, so supply-chain tampering was possible in principle; Hugging Face reports no evidence it occurred.
  • All tokens and credentials were rotated, clusters rebuilt, and metadata-service access from pods was blocked network-wide.

Claims in This Chapter

The intrusion reached cluster-admin within under thirteen hours of first foothold.

Confirmed

SourceHugging Face's own technical timeline blog post

Stated in Hugging Face's own forensic account; consistent with OpenAI's Black Hat description.

No public models, datasets, Spaces, or packages were tampered with.

Strongly Corroborated

SourceHugging Face's own technical timeline blog post

Hugging Face's own assessment, repeated by OpenAI; not independently re-verified in the sources we reviewed.