OpenAI fortifies advanced AI models to block autonomous cyberattacks
OpenAI has strengthened its security models with new safeguards that adapt to their capabilities, reports EP, to ensure they do what is expected of them and prevent new autonomous hacks like the one at Hugging Face.
The security incident caused by two OpenAI models in July, which autonomously left an isolated test environment with restricted internet access and hacked the Hugging Face platform to find information that would help them solve one of the tests posed during an evaluation, has prompted a review of the security measures of sophisticated artificial intelligence systems.
The tech firm assures that they have reviewed and implemented new monitoring, alignment, and containment measures at all stages of the training process, emphasizing reinforcement learning training for containing "the highest risk workloads," as explained in a statement.
The new measures focus, on one hand, on monitoring which, under the new approach, becomes a multistage configuration, including all inference runs and training that used tools, which previously were outside of this review.
The goal is to detect "unauthorized access, data theft, destructive behavior, and attempts to circumvent security measures" and, if any of these actions are identified, an alert will be issued within 30 minutes to notify security, protection, and research teams.
Alignment techniques have also been extended to more stages of the training process, to reinforce the process that makes AI systems "behave as expected and respond to human oversight," to reduce the chances that they can autonomously perform unauthorized actions or cause harm.
These efforts include improving reward models and training models to be more honest and reduce behaviors that exploit weaknesses in rewards, evaluators, tools, or oversight. They also reach behaviors that could cause harm when models interact with external systems or resources.
The changes also involve new safeguards in research workloads that may require code execution or the use of tools with internet access. For example, greater isolation is required, new controls have been implemented to isolate "high-risk and untrusted workloads" from the internet, and test environment security has been reconfigured to "remove potentially vulnerable shared services and reduce privileges."
"We apply these security measures both in research and implementation, adapting them to the capabilities, operational environment, and risk level of each model," they assure.
These new measures have been applied not only to Sol or similar models but also to Astra models (upcoming launch), focused on cybersecurity, which in internal evaluations have shown they can achieve critical capabilities.
This means that Astra models "can identify and develop functional zero-day exploits of all severity levels in numerous real and reinforced critical systems without human intervention," or even "devise and execute end-to-end novel cyberattack strategies against reinforced targets starting only from a general objective."
Related Stories
AI News
SG's 1982 Ventures joins $400m round in US AI startup Higgsfield
54 minutes ago
AI News
Physical AI in spotlight at Tokyo tech expo
54 minutes ago
AI News
Verascient secures $1.2 million to deploy enterprise artificial intelligence infrastructure
55 minutes ago
AI News
Google Cloud Launches Gemini Enterprise AI Platform for Financial Services
55 minutes ago
AI News
Bill Gates calls for ‘human reserved’ jobs in face of AI takeover
1 hour ago
AI News
Can Kazakhstan’s AI Strategy Deliver Both Growth and Smarter Governance?
1 hour ago
AI News
NemoClaw’s AI can be poisoned through a browser tab
2 hours ago
AI News
Can we stop pretending AI isn’t a stupid scam already?
2 hours ago