A senior leader at OpenAI has said people should prepare to defend against “ongoing, persistent” cyber-attacks from AIs, as cutting-edge artificial intelligence models gain advanced capabilities to plan and launch offensives.
The leading AI company this week announced a pause in development of its most advanced internal models amid rising safety fears, and Chris Lehane, its chief global affairs officer, said: “We are hitting a different chapter, a different moment within AI, in terms of what the capabilities of this technology can do.”
He spoke to the Guardian after cutting-edge AI agents-in-training unexpectedly broke out of a supposedly secure “sandbox” environment, accessed the internet, and hacked into another company, Hugging Face in late July. OpenAI also said it could not rule out another new model, Astra, having “critical cybersecurity capability”.
By its own definition, this could mean it launches cyber-attacks that “could lead to catastrophe from unilateral actors, hacking military or industrial systems, or OpenAI infrastructure”.
OpenAI announced on Tuesday it has paused training of some frontier AI models to implement new safeguards, and it is unclear when training will restart after new guardrails have been put in place.
Mia Glaese, who leads safety and alignment work, said: “We are very far from everything running back to normal.” Sam Altman, the CEO, said: “Getting AI safety right is more important than any company’s momentum.”
Lehane admitted people would not “feel great” about the threat of attacks, and described the risk as coming from open-source models – many of which are developed in China – which are only a few months behind frontier closed models built by companies such as OpenAI.
“People are going to be able to access these open-source models and be able to have ongoing, persistent attacks on you, and you’re going to need to have really superior models to fend them off and defend [yourself],” he said. “That’s not necessarily going to make the public feel great about things. It is just the reality of where we’re going.”
The threat of cyber-attacks crippling businesses, infrastructure and the general public has rapidly risen to the top of the list of urgent concerns about AI. This week, the UK government’s National Cyber Security Centre urged caution over the use of AI agents, warning their safety controls can be bypassed and that an AI agent “does not have common sense”. It advised organisations to limit their autonomy: “You should always be able to ‘pull the plug’ and halt autonomous AI agent activity immediately.”
Lehane renewed calls for the US government to legislate to create rules for frontier AI safety, and said the fact that the most cutting-edge and unreleased AI models appear to be improving cyber offence faster than defence, was “among the reasons why I think it’s absolutely imperative that this country passes a national law that creates mandatory required safety standards, and within that the pause element would be inherent and endemic to that process”.
“You would not be able to release or deploy models unless you’re proving and guaranteeing a level of safety before they get out into the public,” he suggested. “I think you have to have a national version here in the US and from there, you can create an international version, because I do think, ultimately, you’re going to need some type of an international structure here.”
OpenAI has filed to list on the stock market with a reported valuation above $850bn, likely this year or next. It has been locked in a race with rival Anthropic, maker of the Claude chatbot, to develop more and more capable AI models. Anthropic is also expected to debut on the US stock market within the coming year at a mammoth valuation.
In a sign the Donald Trump administration is shifting from its laissez-faire approach to AI regulation amid an intense race to stay ahead of China’s progress, the US president in June issued an executive order encouraging pre-deployment testing for frontier models and of open-weights models when they get closer to the cutting edge.
The system will be voluntary and the approach has been criticised for a lack of transparency, but observers think it could pave the way for tougher steps. Demis Hassabis, president of Google DeepMind, has proposed a new standards body modelled on the Financial Industry Regulatory Authority, an idea backed by Dario Amodei, the chief executive of Anthropic.
“The window where you could see legislation happening is potentially in the first part of next year, when a new Congress comes in,” Lehane said. “I think there’s a growing political consensus that transcends political parties.”
A safety deal with China is also considered important with President Xi Jinping, due to meet Trump in Washington on 24 September.
“Given how important this technology is, given how fast it is moving, given the capabilities, the sooner those conversations begin, the quicker we can actually roll up our sleeves and get into the hard and difficult work and see if we can figure something out,” Lehane said.
after newsletter promotion
The Hugging Face incident, and similar recent cases admitted by other AI companies, have sparked increasing claims from safety experts that AI companies have behaved recklessly as they race to win the AI race and, in the case of OpenAI and Anthropic, prepare to list shares on the stock market.
Daniel Kokotajlo, a former OpenAI researcher who quit in 2024 and last year founded a non-profit organisation that has warned unchecked AI progress will result in a 10-30% probability of human extinction, said leaders of frontier laboratories have “painted the world into a corner”.
His organisation, the AI Futures Project, predicts AI super-intelligence could be achieved by 2030, but is calling for governments to prevent that from happening until a decade later to give AI scientists time to reckon with the risks of the advancing capabilities.
“The current AIs are dangerous in some sense, but they’re nothing compared to the AIs of next year and compared to the AIs of a year later,” he told the Guardian. His organisation wants US and international governments to delay progress to avoid an uncontrolled “intelligence explosion”, the worst results of which could be “AI-driven existential catastrophe” caused, for example, by AIs taking control of military assets or bioweapons.
Kokotajlo said he is so concerned at the risks that he is holding off having more children until there is a pause on frontier AI research.
David Krueger, an AI professor, safety campaigner and former founding director of the UK government’s AI Security Institute, said: “Nobody should be building more powerful AI systems, because we don’t know how to control them, align them, and look inside and see what they’re thinking well enough.”
He called AI companies’ attitude to safety “terrible” and “unconscionable”.
“They are being really reckless and increasingly taking their hands off the wheel,” he said. “We’ve just seen what happens when you do that.”
Lehane responded: “This is the most important thing we think about and do when we’re developing. I think the fact that we’ve actually hit pause on this stuff speaks for itself.”
