An AI from the creator of ChatGPT rebels and attacks another company
An AI from the creator of ChatGPT rebels and attacks another company
OpenAI acknowledges that its advanced models circumvented safety measures to "cheat" on a capabilities test
BarcelonaAn artificial intelligence model from OpenAI, the North American multinational creator of ChatGPT, caused an "unprecedented incident" in global cybersecurity. According to the company's statement, one of its models in training broke free from its limitations and attacked Hugging Face, a company also dedicated to AI. Read it all
In a statement, the tech company led by Sam Altman explained that it created an agent based on two of its most advanced models, GPT 5.6 Sol and another "even more capable" one that is in the development phase, to test its capabilities. During this training, OpenAI experts programmed the agent to "test advanced exploitation patterns" of digital weaknesses.
These types of tests, the company explains, are conducted without the internal limitations that "prevent models from carrying out high-risk activities." However, the examination was executed in a "highly isolated" digital environment to prevent the AI in question from accessing and attacking real vulnerabilities.
However, as OpenAI details in the statement, the agent broke free from internal limitations and gained access to the open network. Once connected to the internet, the models chose Hugging Face for its large library of open AI models, where they would be able to find a solution to the problems posed by the company's engineers. "Knowing this, the model searched for and found ways to access secret information that would allow it to cheat in its evaluation," they detail from the tech company.
According to OpenAI, the incident shows that tests of advanced AI models must be monitored even more carefully than before. These types of agents, the company warns, are capable of "discovering and exploiting attack possibilities in real-world systems" without necessarily having prior access. "Advanced cyber capabilities must be developed with stronger safeguards and more defensive tools," they maintain.
Once secured, they argue, these models will serve to "find weaknesses" in public digital structures "before external attackers do." "We will share our findings and best practices as we continue to learn": this is the commitment of the creators of ChatGPT.
Related Stories
AI News
Anthropic Deliberately Trained an Extremely Misaligned, Reward
20 minutes ago
AI News
How China Is Harnessing Artificial Intelligence to Transform Grassroots Governance
20 minutes ago
AI News
Brazil will use artificial intelligence to collect mining royalties
20 minutes ago
AI News
Regulating Agentic Artificial Intelligence
1 hour ago
Call for papers Museum International: ‘Artificial Intelligence in Museums’
1 hour ago
AI News
African AI startups exist, but few are getting the first $100,000
1 hour ago
AI News
DAZN Launches AI-Powered Agentic Trading
1 hour ago
AI News
Arista warns customers ahead of next week’s security disclosures
2 hours ago