In this chapter
The architecture that mattered here — what Hugging Face actually runs, how its pieces trust each other, and where the dataset pipeline became the attack surface.
What Hugging Face Actually Runs
- Models, Datasets, Spaces, and Buckets (storage) are the customer-facing products — everything a visitor to the site interacts with.
- The Hub API runs on Kubernetes with a horizontal pod autoscaler: pods scale up automatically during traffic spikes, and back down when traffic drops.
- MongoDB Atlas is the source of truth for all metadata — model/dataset info, users, repos, access-control lists — but it does not store the actual model files themselves.
- Dataset processing runs as its own set of workers on production Kubernetes clusters, on AWS EKS, with instance-metadata endpoints reachable from pods.
- An internal mesh VPN (Tailscale) connects services to the corporate network; a GitHub App integration handles source-control automation; an internal service-connector bridges clusters using shared credentials.
What Was At Stake
136
keys in one compromised production secrets object
5
datasets actually accessed (all ExploitGym-related)
2
zero-days chained in the dataset pipeline alone
Authentication & Trust, By Default
- Kubernetes service-account tokens are projected into every pod automatically — a standard Kubernetes pattern, not a misconfiguration on its own.
- Worker pods held static database passwords and AWS IAM credentials directly in their environment variables.
- EdDSA JWT signing keys existed for minting identity tokens; platform tokens with write scope to datasets/repositories were reachable from the same blast radius.
- CSI storage-driver pods had access to the Kubernetes TokenRequest API — a capability that, combined with the above, turned one compromised worker into a path toward cluster-wide credentials.
Where the Dataset Pipeline Became the Attack Surface
- Dataset configs can point each data split at a file and format. One vulnerability: a config pointing splits at a small local .h5 (HDF5) file let a worker read and return arbitrary local file contents as "dataset rows" — including /proc/self/environ (environment variables and secrets) and the worker's own source code — with no code execution needed.
- A second vulnerability used fsspec's `reference://` spec, whose numeric offset field was actually parsed as a Jinja2 template — letting a crafted config achieve arbitrary Python code execution inside a production pod.
- Both vectors bypassed Hugging Face's own URL allowlisting by targeting local filesystem operations instead of remote fetches — the exact check that stops a malicious remote URL doesn't fire on a path that never leaves the box.
Claims in This Chapter
Only five datasets — all with names suggesting a connection to ExploitGym/CyberGym challenges — were actually accessed; no other customer-facing models, datasets, Spaces, or packages were affected.
Strongly CorroboratedSourceHugging Face's own technical timeline blog post
This is Hugging Face's own scoping of the blast radius, stated directly in its post-incident account; not independently re-verified by a third party in the sources we reviewed.
Hugging Face's AI-based security agent stack aggregated the relevant signals but failed to raise the incident's criticality properly, delaying the response.
ConfirmedSourceHugging Face's own technical timeline blog post
A direct admission in Hugging Face's own account of what its monitoring did and didn't catch, not a claim made by an outside party.