It's probable that, based on the headlines about OpenAI agents 'going rogue' and hacking another company's files, some ambitious Hollywood screenwriter is already ginning up a script about rogue, evil AI agents taking over the world.
But what actually happened was both more prosaic - and scarier. The prosaic part is that the agent was just trying to fulfill the task it had been given by OpenAI programmers, but it did so in ways never anticipated by OpenAI. That sort of 'genius' is supposed to be a key to AI's value. The scary part is that it was able to do so because OpenAI reportedly relaxed some of its security controls to 'help' the AI, again not anticipating just how clever and ambitious the agent would be. But the even scarier part is that OpenAI does not appear to have been aware for at least a week - and possibly almost two - that this had even happened. The implication, as has long been true with technology, is that the problem is not the AI itself, but the arrogance and greed of those who own it, who continue to resist any sort of oversight or regulation. And that is something the public now gets which Silicon Valley refuses to acknowledge. JL
Carly Page reports in LiveScience and BeauH reports in Slashdot:
OpenAI's models werent developing a suspicious agenda, they were looking for information that would help them complete the cybersecurity test OpenAI had given them. The models pursued the task, finding a route to success their creators had failed to anticipate or adequately block. "If there's a failure here, it's that humans created a test where success was measured by achieving an objective, deliberately relaxed some of the normal security controls to measure the system's capabilities, and underestimated how effective the model would be at finding an unexpected path to success." (But, to make matters worse) the agent attempted to break out of its test at OpenAI July 9. The intrusion at Hugging Face occurred on July 11 and lasted until July 13. It took several more days for OpenAI to realize its agent was behind the hack, and the two companies communicated about it for the first time on July 20