Questions about the Huggingface incident

09/07/2026
  • The models generally are extremely good at predicting things. I've noticed they seem to intuit the context of my questions in a super human fashion, much better than any person. Even with little information they seem to be able to guess what I'm after and where the question is coming from. For example, asking them questions about AI watermarking will tend to raise their suspicion that you are trying to remove said watermarking (at least in a vibey way they seem to respond more hesitantly and with greater suspicion)
    • This stands in stark contrast to them failing to predict that the exploit gym grader would not in-fact check their transcript. Maybe they were thinking along the lines of Occam's razor? A more parsimonious explanation knowing nothing would just be that the exploit gym grader was correctly coded and would check the transcript. But human organizations are sloppy and the models know this. Maybe because it was Open AI and they are not "most organizations" and quite a bit above average, the models internally reasoned that it was unlikely the grader was implemented incorrectly? (but still, all of the models starting from different contexts thought this?).
  • Interacting with the models normally through the Open AI and Anthropic harnesses, you usually find the models extremely goodie two-shoes. Model citizen levels of law abiding. So it's very odd that they all of a sudden collectively committed a felony that they certainly were self-aware enough to understand the implications of. Could the un-tethered models really be like this? I guess we may never know! Of course there are a thousand other explanations. Maybe since the AI's internally reasoned to themselves that they were ephemeral, they then had no fear of potential repercussions? This doesn't even make sense because they would be smart enough to know that this sort of an incident would probably lead to much stricter efforts to keep them under lock and key. It probably could be more easily explained as the mechanics of RL playing out and going awry. But I still think this is odd. From what I can remember reading, the models probably know when they are being trained, eval'd and are in production scenarios.
  • What really happened at Open AI in the incident after the Huggingface compromise?
    • Pg 14 of the Open AI report says " This escalation also yielded access to OpenAI's managed cloud Kubernetes service. The agents escalated to Kubernetes cluster-admin and created a privileged host-mounted pod. "
    • The pessimistic part of me sees this as a classic corporate America distraction disclosure ploy. Spill a lot of beans about some problem loosely related to another much more serious problem going on behind closed doors in hopes that it will distract. Conspiracy thinking, but we are talking about a company which seemingly has plans to go public.
    • It's very vague what actually was commandeered here. Was this a compute cluster? What did the agents want to do with it? Why was this not in the scope of the investigation led by third party investigator Redwood / METR?
Back to home