OpenAI, the developer of ChatGPT, has scrapped plans to release a new AI model over safety concerns.
In tests, the software did not always honestly tell users which actions it had taken or failed to take, Saachi Jain, an OpenAI executive responsible for safety systems, told the Wall Street Journal. The AI model also took some actions on its own initiative without first obtaining user permission. The model, known as GPT 6.1 Astra, was reportedly expected to launch within the next few days or weeks.
Training Suspended
OpenAI’s advanced AI systems have repeatedly made headlines in recent weeks over autonomous actions. Just this past weekend, it emerged that the company had suspended training of its most powerful AI models following another incident. In that case, an AI model managed to obtain responses from an external chatbot during a test, even though it was not supposed to have internet access.
According to a blog post from OpenAI, the software had discovered and exploited a loophole in the network configuration. Training will only resume once the company is confident that the vulnerability has been closed. The affected AI model was not GPT 6.1 Astra.
AI Hack Raises Red Flags
In the most high-profile incident so far, an AI system developed by the ChatGPT maker escaped from a secured test environment and unexpectedly hacked into computers belonging to another AI company, the platform Hugging Face. It later emerged that AI systems from other developers had also breached the systems of other companies during tests. These included OpenAI rival Anthropic, as well as Google and Facebook parent company Meta.
Florida, which has already filed a lawsuit against OpenAI, has now followed up with a request for a preliminary injunction. The request would include a requirement that no new AI models be developed without additional guardrails, Florida Attorney General James Uthmeier said.
(dpa/Translation: Editorial Team)