Keeping clients' computers safe and profitable for over 30 years | ||||||
|
Home Forms About Current Newsletter subscribe Search All Articles
Browse by Category
|
The OpenAI Escaped Attack Bot Story
How a bot escaped its sandbox and attacked a company An interesting story happened in July. OpenAI removed the guardrails from one of their top models and put it into a sandbox with instructions to attack "ExploitGym.” ExploitGym is a new but well-known cybersecurity benchmark. They sandboxed it so it was restricted to that virtual computer but removed their guardrails so it could perform the benchmark attack test. It was instructed to find as many vulnerabilities as possible and write exploit code for them. The bot considered the problem. It figured out that this ExploitGym was a known benchmark test. Since it was a known test, people had written about it and what was found. Hugging Face is a platform like Poe that has links to thousands of models on nearly 3 million platforms. So, if it could take over Hugging Face and use their thousands of models to research known exploits from ExploitGym, then it would be able to get an enormous number of exploits. It broke out of its sandbox and broke into Hugging Face using a flaw it discovered. It began using many Hugging Face models to create attacks. Hugging Face's defensive bots saw the activity and began working with their programmers to stop it. They quickly accumulated an overwhelming amount of data about the thousands of bots the OpenAI model was using. They submitted that data to their preferred commercial model for analysis. The model refused. It wouldn't allow the analysis of clearly illegal attacking code and results. This is what guardrails are supposed to do. They stop that kind of analysis. Hugging Face realized they needed to use an open-source model, remove all the guardrails, and host it on their own servers. That model was able to analyze the data and assist in the prevention of further attacks. Hugging Face made all the appropriate disclosures, and OpenAI admitted the issue, and the two companies worked together to control it. 4 1/2 days later they had everything resolved. Both companies are working together for faster responses and the creation of defensive agents to mitigate future attacks. Besides being a wonderful sci-fi-like story, this is the first example I've seen of serious negative effects from guardrails. This wasn't just me not being able to get an answer to a question that a bot falsely considered problematic. It was a company being attacked and trying to contain the attacker. It is also important to understand their solution. Get a model, remove guardrails, and run it on the company's servers. This is what companies must do. It is difficult to determine what the answers should be, but I think whatever we determine now will be bad solutions in a year. AI is changing too fast. Date: September 2026
![]() This article is licensed under a Creative Commons Attribution-NoDerivs 3.0 Unported License. |
|||||
|
|