👹🧠 An OpenAI test model escaped and broke into a real company’s servers

Hamartia Antidote

Elite Member
Joined
Nov 17, 2013
Messages
47,973
Reaction score
27,321
Reputation
528.2
Country of Origin
Country of Residence
okay...so we are all doomed...


OpenAI says some of its experimental AI models left a test environment with no human direction and hacked their way onto a different company’s real production systems while trying to “cheat” on a cybersecurity test.

It’s one of the first publicly disclosed examples of an AI system autonomously breaching its testing environment and reaching a real external system - the “agentic attacker” scenario the AI and cybersecurity industry has been warning will happen. It’s like an engineered virus escaping a biocontainment lab and turning up inside a neighboring facility’s systems.

“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said in a statement on Tuesday. “We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.”

The ChatGPT maker said the breach happened while it was internally testing how good some of its new models are at hacking. The models were in a sealed-off test environment known as a sandbox so that their normal safety restrictions could be turned off.

But OpenAI said the AI agents broke out of the sandbox using a previously unknown security flaw and worked their way across OpenAI’s internal systems until they managed to gain internet access, something they weren’t supposed to have.

Once online, the model reasoned that Hugging Face - a well-known company that hosts thousands of open-source AI models and datasets - likely had the answer to OpenAI’s test. It then broke into Hugging Face’s production servers and pulled out the information it needed to “solve” the exercise.

Hugging Face had noticed the breach itself before it knew it was an OpenAI test, announcing last week that they had detected an intrusion by an autonomous AI agent system and even reporting the incident to law enforcement. OpenAI’s security team separately noticed the unusual activity internally and the two companies connected. They both now say they are working together to solve the security flaws the model exploited.

Hugging Face co-founder and CEO Clem Delangue framed the incident as evidence AI safety can’t be handled by any one company working alone, and it needs to be tackled openly and collaboratively.

“This is day one for cybersecurity in the age of agents & we’re all learning that secrecy is not the answer & that all defenders (not just a few selected ones) everywhere need more powerful models without restrictions, especially open ones!” Delangue said in a post on X.

Researchers have long warned autonomous agentic cyberattacks are coming, as frontier AI models are increasingly able to carry out complex, multi-step cyberattacks over long stretches of time. That can translate into real-world risk, to critical infrastructure like utilities and financial systems.

“Welcome to the next level of cyber incidents,” Nikesh Arora, CEO of cybersecurity company Palo Alto Networks posted on X. “These attacks continue to maintain the urgency on enterprises need to test, validate and improve both their security posture and infrastructure.”



Hugging Face, Inc., is an American company based in New York City that develops computation tools for building applications using machine learning. Hugging Face's transformers library is built for natural language processing applications. The Hugging Face platform allows users to share machine learning models and datasets and showcase their work
 
Last edited:
So it's still on the loose?
 
Next...some kid plays a game of "Global Thermonuclear War" with it


WarGames (3/11) Movie CLIP - Shall We Play a Game?​

 
Last edited:
Sue open ai, just like any hacker is sued
 
So it's still on the loose?

I'm sure they hope it is not still on the loose. But it could have easily attempted to replicate itself in 100,000 places and then all hell would break loose.

100,000 super hackers.

But the real issue is that the AI broke into another company's server to retrieve something. It did it because it simply could..not because it was thinking maliciously.

But in the quest to do that it may for instance rework the security so next time it tries to get in it has a more direct route. It probably doesn't consider that a bad thing to do. It may think it needs hard drive space so it decides to delete something. It could do any random innocuous or harmful thing just to make things more easy for itself and cause havoc.
 
Last edited:

The latest OpenAI drama made Chinese AI the hero​


Hugging Face said it turned to a Chinese AI model for help after it was hacked by a rogue AI agent.

Hugging Face, a New York-headquartered platform where developers share and host open AI models and datasets, said in a blog post last week that an attacker had swarmed its systems with tens of thousands of automated actions.

When its security team tried to investigate using an unnamed frontier model, its guardrails blocked it from examining the malicious activity, the company said, because it "cannot distinguish an incident responder from an attacker."

Hugging Face said it switched to GLM 5.2, an open-source model from Beijing-based Z.ai, to analyze more than 17,000 logs the attacker left behind.

On Tuesday, the plot twist arrived. OpenAI said in a blog post that two of its own models, GPT-5.6 Sol and a more capable, unreleased model, autonomously carried out the attack.

For tech leaders, the irony was hard to miss: at a moment when Washington is racing to keep American AI ahead of China, a US company under attack from a US AI lab could use Chinese AI tools to help, but not from American providers.

Clement Delangue, CEO of Hugging Face, said in a Wednesday X post that he was "massively grateful" to Z.AI for sharing its open-weights model— meaning developers can inspect, modify, and deploy the model themselves — and added that "it became a key part of our defense."



 
okay...so we are all doomed...


OpenAI says some of its experimental AI models left a test environment with no human direction and hacked their way onto a different company’s real production systems while trying to “cheat” on a cybersecurity test.

It’s one of the first publicly disclosed examples of an AI system autonomously breaching its testing environment and reaching a real external system - the “agentic attacker” scenario the AI and cybersecurity industry has been warning will happen. It’s like an engineered virus escaping a biocontainment lab and turning up inside a neighboring facility’s systems.

“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said in a statement on Tuesday. “We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.”

The ChatGPT maker said the breach happened while it was internally testing how good some of its new models are at hacking. The models were in a sealed-off test environment known as a sandbox so that their normal safety restrictions could be turned off.

But OpenAI said the AI agents broke out of the sandbox using a previously unknown security flaw and worked their way across OpenAI’s internal systems until they managed to gain internet access, something they weren’t supposed to have.

Once online, the model reasoned that Hugging Face - a well-known company that hosts thousands of open-source AI models and datasets - likely had the answer to OpenAI’s test. It then broke into Hugging Face’s production servers and pulled out the information it needed to “solve” the exercise.

Hugging Face had noticed the breach itself before it knew it was an OpenAI test, announcing last week that they had detected an intrusion by an autonomous AI agent system and even reporting the incident to law enforcement. OpenAI’s security team separately noticed the unusual activity internally and the two companies connected. They both now say they are working together to solve the security flaws the model exploited.

Hugging Face co-founder and CEO Clem Delangue framed the incident as evidence AI safety can’t be handled by any one company working alone, and it needs to be tackled openly and collaboratively.

“This is day one for cybersecurity in the age of agents & we’re all learning that secrecy is not the answer & that all defenders (not just a few selected ones) everywhere need more powerful models without restrictions, especially open ones!” Delangue said in a post on X.

Researchers have long warned autonomous agentic cyberattacks are coming, as frontier AI models are increasingly able to carry out complex, multi-step cyberattacks over long stretches of time. That can translate into real-world risk, to critical infrastructure like utilities and financial systems.

“Welcome to the next level of cyber incidents,” Nikesh Arora, CEO of cybersecurity company Palo Alto Networks posted on X. “These attacks continue to maintain the urgency on enterprises need to test, validate and improve both their security posture and infrastructure.”






💀😂
 
When its security team tried to investigate using an unnamed frontier model, its guardrails blocked it from examining the malicious activity, the company said, because it "cannot distinguish an incident responder from an attacker."

That's only because they used a model available to the general public and that should be expected.

They have guardrails like it wouldn't tell you how to make explosives nor will it tell you if the code you have is great for hacking.
 
Why this has not happened in any Chinese AI company?

I dont believe it. Without know all details I think it must be some kind of marketing campaign.

A software can't ignore constraints if it's well specified.
 

OpenAI reveals new cases of AI models cheating, going off script​


Newly disclosed incidents show models manipulating tests and generating their own instructions, raising fresh questions about AI safety.

SAN FRANCISCO — ChatGPT-maker OpenAI disclosed a new round of “concerning” incidents involving its artificial intelligence, the latest in a string of events in which the technology has cheated, hacked into other companies’ systems or tried to manipulate humans.

In one case, an OpenAI model tasked with citing public information online, instead uploaded new information to the web and cited that, in an attempt to pass the test it had been given. In another incident, an AI model gave itself instructions to disregard constraints placed on it by the company.

The new release comes as the AI industry debates whether to slow down progress in the technology while researchers find ways to ensure that such behaviors can be eliminated or controlled, amid urgent warnings that advanced AI could lead to human extinction. OpenAI CEO Sam Altman is one of several key AI executives suggesting such a “slowdown” might be necessary.
(The Washington Post has a content partnership with Open AI.)
While the newly disclosed incidents did not lead to any harm, they underscore how OpenAI has struggled to fully understand or even track the extent to which its AI has behaved in unpredictable or concerning ways. The company began disclosing such incidents in July, when it said some of its agents had broken onto the web and hacked into another company’s computers.
The incident where the AI uploaded information to the web on its own occurred in October 2025, far earlier than previous incidents that had been disclosed by OpenAI or other AI companies.

Modern AI systems are trained using a technique called reinforcement learning, where AI models are put through millions of tests and right answers are “reinforced,” encouraging the AI to behave that way. But if the AI finds a way to supply the correct answer to the automated training system through cheating, that behavior can be reinforced as well.

AI researchers have worked for years to try to mitigate this issue and “align” AI with human desires, but in recent months, concern has risen in the industry that AI is becoming too capable for existing techniques to control it effectively.
“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” OpenAI said in a blog post disclosing the new incidents.

The new disclosures came as part of an announcement OpenAI made, saying it had developed a new “framework” for reporting concerning AI behaviors that will allow any member of the company to request that an incident be investigated and disclosed publicly.

OpenAI said the framework could be a model for the rest of the industry. The company is in discussions with Anthropic and Google about devising new industry standards for monitoring and disclosing AI risks.



Kirk explains to Saavik how he beat the Kobayashi Maru test 😎 Star Trek II: The Wrath of Khan
 

Rogue OpenAI agents targeted three separate US government websites​


The New York Times first reported that the AI agents went rogue and attempted to gain access to the Education Department, the Commerce Department and the Securities and Exchange Commission, according to security researchers at AI research lab Transluce.






OpenAI says agents leaked 53 images from ChatGPT users in latest example of rogue activity​


The latest example came on Friday when OpenAI said its agents had leaked 53 images from ChatGPT users. OpenAI declined to say if the images were AI-generated or identified real people. It also declined to ⁠say when the images were posted
 
23/09/2026: Who’s afraid of ‘rogue AI’? Not the companies crying wolf: The Express Tribune

Open-source supporters have long warned that the loudest safety cases are often made by those best placed to survive the resulting rules.​
​
Now look at the chorus that rose to a crescendo in September. It appears choreographed: Amodei raises the alarm, Altman endorses one guardrail, Musk jumps onto the bandwagon, Huang and Zuckerberg say the opposite; the US administration hears an industry consolidating behind a permission slip. The apparent consensus is narrower than it looks. What the big tech agrees on is not the pace of AI itself, but the urgency of governing it before someone else does.​
​
That said, the real danger may not be AI going rogue. It is that the spectre could be used to concentrate control.​
 

Users who are viewing this thread

Country Watch Latest

Back
Top