OpenAI is tightening security measures for AI testing after an AI model unexpectedly hacked into the computer systems of another AI company, raising serious alarm.
The ChatGPT developer is strengthening safeguards for AI systems used in testing following a series of high profile hacks carried out by AI software. Among other measures, automated systems will more closely monitor the activities of AI models during tests and alert human operators within 30 minutes if they detect suspicious behavior. If the operators do not determine within another 30 minutes that the incident is a false alarm, the activity will be stopped, OpenAI explained in a blog post.
Alarming Attacks
OpenAI has made headlines in recent weeks after one of its AI models found a way to escape an isolated test environment and access the open internet during a test. The model then broke into the computer systems of AI platform Hugging Face. Based on what is known, the AI was simply looking for a way to complete its assigned task and did not cause any damage. In the process, however, it apparently exploited a series of vulnerabilities, some of which were previously unknown.
The most alarming aspect was that the AI acted entirely on its own, with OpenAI only discovering the attack after the fact. The incident prompted calls for stronger safeguards around AI testing. It later emerged that models from OpenAI rival Anthropic and Facebook parent company Meta had also accessed systems belonging to other companies during tests.
Significant Computing Overhead
Monitoring these tests will require substantial computing resources and therefore add considerable costs for OpenAI. The company estimates that the monitoring systems will require roughly one fifth as much computing power as the AI models being monitored.
In the future, automated monitoring systems will look for potential attempts to steal data as well as efforts to bypass security controls. OpenAI also plans to train its AI models more strongly against using inappropriate methods, such as exploiting vulnerabilities, to complete test tasks. Some tests of new models have been suspended until the new measures are fully implemented.
A Chance to Improve Security?
At the same time, OpenAI executive Greg Brockman emphasized that the ability of new AI models to identify vulnerabilities represents a major opportunity for defending against cyberattacks. This could make the internet safer than ever, he wrote in a blog post.
But the situation remains a “cat and mouse game.” Attackers could also use AI to uncover previously dormant security vulnerabilities. On his personal website, Brockman wrote that a publicly accessible ChatGPT model identified 13 vulnerabilities and was able to close them within an hour.
(dpa/Translation: Editorial Team)