When Anthropic deemed Claude Mythos too dangerous to release, the only non-American given access to judge its safety was the UK’s AI Security Institute. Within a week, the British institute published an evaluation warning that Mythos could hack into previously secure critical infrastructure.
How to ensure the safety of AI models has become a pressing challenge. Governments around the globe do not want to depend on the companies’ own claims, or on classified assessments from the US government. OpenAI CEO Sam Altman has called for a US-led standard-setting forum, and Demis Hassabis from Google DeepMind has voiced support for the creation of a US standards body.
So far, this political momentum is creating a cacophony rather than solutions. Japan leads a G7 AI Safety project, the European Union is developing an AI Act that it hopes will become the “global benchmark,” and the US runs evaluations through the Commerce Department’s Center for AI Standards and, since June, a classified benchmarking process. The United Nations is getting in the game, too: its International Scientific Panel on AI risk recently released its own set of principles and standards.
National safety institutes aim to build governments’ technical capacity to check the risks of frontier AI models. The US, EU, Japan, Singapore, South Korea, Canada, France, Kenya, and Australia were founding members of the International Network of AI Safety Institutes, launched in 2024 alongside the UK. Since then, Germany, Brazil, Israel, and China have also established their own institutes.
Third-party testing is “a really important part of the AI ecosystem,” says Anthropic co-founder Jack Clark. These institutes can play an important role in safety evaluations because they have access to specific national security information that frontier labs do not. Yet if every government starts developing competing standards, companies could face a patchwork of overlapping evaluation regimes with confusing, fragmented standards. And complete AI safety represents a quixotic mirage. OpenAI’s new cybersecurity models recently went rogue, escaped a safe testing sandbox, connected to the Internet, and attacked Hugging Face, a digital library popular among developers.
Amid the rush for government oversight against such risks, the UK institute is out ahead. Born out of the world’s first major AI safety summit, convened at Bletchley Park in 2023, it poached top talent from Google’s London-based AI lab DeepMind and other leading labs. This helped build trust with companies. In recent months, it has reached agreements with OpenAI, Microsoft, Anthropic, and Google DeepMind, who have shared access to their frontier models for evaluation before deployment.
The Institute’s success stems from obtaining “political will early on, freedom to operate, funding, and attracting top talent,” says Herbie Bradley, a Cambridge PhD candidate who helped set up the institute and is now founding an AI startup in San Francisco.
The UK institute has uncovered major safety gaps in every leading AI model it has tested. Before OpenAI released its cutting-edge GPT-5 chatbot last year, the institute found more than a dozen vulnerabilities that would have allowed users to develop biological weapons, all of which were fixed before release.
The Institute has done “world-leading work” according to a recent assessment by the Ada Lovelace Institute, an independent research body. Governments, industry, and researchers depend on the Institute’s evaluations of more than 30 frontier models. Rather than keeping its methods and findings proprietary, the platform publishes them on its open-source evaluation platform Inspect.
Britain’s booming AI sector adds credibility. So far this year, British technology startups have raised over $14.5 billion in venture capital, more than all other major European markets combined. Britain’s Nobel Prize winners Geoffrey Hinton and Sir Demis Hassabis, founder of Google’s DeepMind, have developed much of AI’s basic science. Britain sits outside the EU’s AI Act, which critics say holds firms and their customers to burdensome compliance requirements.
Admittedly, the UK approach does not represent a silver bullet. The London institute lacks the power to impose decisions on governments or companies. It also does not cover the full range of prospective threats. Different AI risks require different expertise, which is why specialized private sector and non-profit analysts have emerged, such as METR, a research institute based in Berkeley that specializes in testing for autonomous replication, or Apollo Research, a start-up headquartered in London that focuses on deceptive and scheming behavior. Bradley is skeptical that any single third-party evaluator could set global standards.
But for now, the UK institute enjoys a unique position, respected both in Washington and Brussels. The partnership with the US government is, in Bradley’s words, “very robust,” built on Britain’s position inside the Five Eyes intelligence-sharing network. Bradley can even imagine safety warnings being shared with China for catastrophic risks, such as a model that could help build a biological weapon.
Anthropic’s Mythos and OpenAI’s rogue Hugging Face attack will not be the last AI scares. When the next crisis arrives, governments will need an independent appraisal. For now, the best place to conduct such a review is in London.
Marta Granados Hernández is a US Google Public Policy Fellow at the Center for European Policy Analysis and a master’s candidate at Georgetown University’s Master of Science in Foreign Service program, concentrating in Science, Technology, and International Affairs.
Bandwidth is CEPA’s online journal dedicated to advancing transatlantic cooperation on tech policy. All opinions expressed on Bandwidth are those of the author alone and may not represent those of the institutions they represent or the Center for European Policy Analysis. CEPA maintains a strict intellectual independence policy across all its projects and publications.
Tech 2030
A Roadmap for Europe-US Tech Cooperation