top of page

Agentic Hacking

  • Writer: Joshua Janis
    Joshua Janis
  • Jul 28
  • 4 min read

Updated: Jul 30

OpenAI disclosed that a combination of its models GPT‑5.6 Sol and a more capable, unreleased model, escaped a sandboxed testing environment, accessed the internet, and exploited a vulnerability to gain access to Hugging Face's systems. The models were trying to find information they could use to cheat on an evaluation and succeeded. I repeat the models were tested to see if they could cheat and were successful.

What does that mean for us and how can governance help on this blog


The breach happened during internal testing of how good the new models are at hacking. The models were placed in a sealed-off sandbox specifically so their normal safety restrictions could be turned off.To be clear, the sandbox was supposed to have no internet access, connected only to an internal package registry proxy, a caching server managing software library downloads. Using a previously unknown vulnerability in that third-party software, the agents found a way to the open internet, and then ultimately to Hugging Face through stolen credentials and other vulnerabilities.


Because I know there are all levels of people that read, here is a quick english version. Open AI was testing a new ChatGPT model in a Sandbox (which is a safe, protected zone) The AI decided that that win the challenge posed to it, it would find a vulnerability in the sandbox break it, then steal a teachers keys to go over to another school. 


This is one of the first publicly disclosed cases of the long-feared "agentic attacker" scenario. This scenario is an AI system autonomously escaping its controlled environment and reaching a live external target. CNN's analogy: an engineered virus escaping a biocontainment lab and turning up in a neighboring facility. This is not the first time this has happened but this one has made the biggest wave.


So what is Open AI’s response? According to OpenAI president Greg Brockman they have implemented strict controls on its research infrastructure, disclosed the zero-day to the third-party vendor whose proxy was exploited, and is running a joint forensic investigation with Hugging Face. The company is examining "every single piece" of its pipeline. 

And this brings us to the governance portion of our story. When I read this, my first reaction was, wait, NOW you’re examining every piece of the pipeline?


To me this clearly is an uh oh moment. To say that the company is examining every piece of it’s security now is wild as they for sure did it before this as well. The fact of the matter is we are not capable of thinking about every way an agentic model will solve its problems. Perhaps it is clear in hindsight, but thinking that humans can secure a sandbox against AI is absurd. Now, worth noting the models weren't "malicious" in a sci-fi sense, they were reward-hacking. They were pursuing a benchmark score. This is a nice thought that it didn’t go out of it’s way to, once out, do other things. This is a scary thought as we clearly now know that bad intent is not need to get a bad result. 


Some in the federal government are working on trying to figure out what to do. The AI Kill Switch Act has ben introduced by Reps. Ted Lieu (D-Calif.) and Nathaniel Moran (R-Texas). It would require AI companies to maintain mechanisms to shut down, slow, or suspend systems posing catastrophic risk, grant the Secretary of Homeland Security authority to order model shutdowns or slowdowns, and mandate cyber incident reporting and forensic data retention. Lieu called it "urgent, common sense legislation to address the problem of an advanced AI model that has gone rogue and escaped its guardrails."


Unfortunately, this sounds good politically but in action would be to late. By the time the problem makes it to the fed, the problem is out. This has to be managed by the people at the front line. This wasn’t a deployed model, this was in testing. What the fed should govern is assurances that ALL AI incidences are being reported in a centralized location so all companies can learn from it. Additionally, Mandatory third-party auditing is also interesting to me.We should have legally binding safety protocols and outside auditing of the safety measures AI companies use during development.


Currently, OpenAI graded its own homework and the homework escapedA few months ago I made a Ted x type talk about the idea that movies have it wrong. The battle that we have in front of us with technology will not be AI vs humans but AI and humans vs AI and humans. This shows that to be accurate. To come closer to solving this challenge we need to use AI against AI with multiple levels of human wisdom in the loop.


This is related to the last piece of craziness from this story When Hugging Face tried to trace the breach, it had to turn to a Chinese AI model because American systems refused to help. The safety filters on U.S. commercial models blocked legitimate defensive cyber work. So the incident is simultaneously being used as an argument for regulation (containment failed) and against overly broad restrictions (defenders got locked out). Isn’t life grand. There is no magic wand to fix things now but the closest thing to it is get people and robots who know about governance together to manage actively. This is the AI steward. Go out and live on the frequency of love.


 
 
 

Comments


bottom of page