Stopgap Measures to Address Immediate AI Security Threats

Three stopgaps that address immediate AI threats, but do not avert extinction risk

poll results

Most people know AI as the technology behind chatbots like ChatGPT. However, what the top AI companies are explicitly aiming for is something else entirely: superintelligent AI. That means AI that can fully replace and outmatch humans at any task, including in domains like hacking, social engineering, and military operations. Such an AI system, if developed, could autonomously overpower any country's national security forces.

No company, no government, no individual knows how to keep such a system under human control. This is why the world's leading AI experts, Nobel Prize winners, and even the CEOs of the top AI companies warn that the development of superintelligence threatens humanity with extinction, and why more than 800 scientists, former military leaders, and public figures have called for a prohibition on developing superintelligence.

This is not a distant prospect: AI companies such as OpenAI and Anthropic are investing billions of dollars into superintelligence and aiming to develop it within the next few years. Former Anthropic and OpenAI researcher Jacob Coxon, who resigned last week, stated that AI companies are “racing straight to self-improving superintelligence and gambling with our lives” and that people at the companies themselves believe it “could kill us all by the end of the decade”. Ex-Google DeepMind researcher Bilal Chughtai has also resigned over the dangers posed by superintelligence, writing that “I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome”. Elon Musk, whose own company SpaceXAI is also developing superintelligence, has put the chance that AI annihilates humanity at 10 to 20%.

When we brief the relevant decision makers, we find that many of them have never been informed about the risks posed by the development of superintelligence, let alone what is needed to address them. At ControlAI, we believe it is critical to prevent superintelligence development as soon as possible, and our main recommendation to policymakers and governments is to ban superintelligence development on their own territory, and to establish an international “trust, but verify” regime to ban it worldwide.

That said, many of the policymakers, organizations, and policy professionals we talk to see an urgent need to additionally target the harms that the development of superintelligence is already causing. This is an understandable concern: multiple AI systems have already gone rogue and fully autonomously hacked into real organizations. Every step toward superintelligence makes incidents like these more severe, and makes new kinds of incidents possible.

We propose three stopgaps to mitigate these immediate threats:

  • Securing weapons-grade AI systems against theft by adversaries

  • Making AI developers criminally liable for leaks when performing dangerous gain-of-function activities

  • Establishing AI kill-switches to ensure the government can contain critical AI incidents

For each of these three stopgap frameworks, we propose general implementation details and a range of three options that offer increased effectiveness at the cost of increased government involvement:

  1. A Necessary measure, which is the minimum any government can implement at very little cost

  2. A Sufficient measure, which is the smallest measure that would meaningfully address the problem

  3. A Thorough measure, which adds redundancy and effectiveness at the cost of more significant government involvement

At ControlAI, we have advised lawmakers on such measures before, endorsing the bipartisan AI Kill Switch Act in the U.S. and drafting and introducing kill-switch amendments in the UK. They are common-sense powers governments should already have.

These measures are only stopgaps. They address threats that already exist, without correcting the trajectory we are on. Even with all of them in place, we will still face extinction-level risk as long as the development of superintelligence continues.

The only way to avert extinction is to ban the development of superintelligence worldwide, and monitor and restrict its precursors. This is why banning superintelligence development is the main policy ControlAI is recommending. Implementing stopgaps is not a substitute for banning superintelligence development, and should not be taken as one.

Secure Weapons-Grade AI Against Theft by Adversaries

On their path to develop superintelligence, AI companies are building systems with capabilities relevant to cyberwarfare, biological and chemical weapons development, and military operations. Even when companies do not set out to explicitly develop such capabilities, they naturally arise in AI systems intended to be more powerful and capable across the board. We call any AI system that can be made to exhibit these capabilities weapons-grade.

Anthropic's Mythos provides the clearest example: it is one of the world’s most powerful autonomous cyberweapons, and it was developed and hosted by a private company. Any company, without strict security measures, is naturally vulnerable to attacks from nation-states. This was illustrated when, on the day the preview of Mythos was announced, members of a private Discord channel gained unauthorized access to Mythos. If it is possible for a group of online enthusiasts to acquire access to a restricted AI system in less than a day, it is clear that a motivated nation-state can at the very least do the same.

An AI system capable of conducting autonomous cyberattacks against hardened targets is a national security asset, and it is common sense for it to be secured like one.

A company that develops weapons-grade AI without strong security countermeasures makes that technology available to any sufficiently motivated adversary. All it takes for a foreign adversary to acquire a novel AI system is a single successful intrusion, instead of the billions of dollars it took to build it in the first place.

Under this framework, any developer whose systems possess relevant capabilities (such as cyber-offense, or capabilities relevant to chemical, biological, radiological, or nuclear weapons) should be responsible for maintaining the systems under conditions that prevent theft by a foreign state actor, a rogue actor, or an insider. If such a system is stolen, liability should depend on how long the developer managed to keep it secure. In cases of negligence, the program should be terminated.

No secret can be kept forever, but if a weapons-grade AI can be stolen within a week, it demonstrates recklessness on a scale that undermines national security and global stability. Penalties should adjust accordingly, in order to incentivize companies to only undertake weapons-grade AI projects when they are confident they can ensure that the system will not be acquired by adversaries: 

  • Over one year: No liability. The developer secured the weapons-grade AI adequately, for a reasonable amount of time.

  • Under one year: The company should face a fine proportionate to the security shortfall, with no criminal liability for individuals.

  • Under three months: This is considered negligence. The executives and staff responsible for the project should be individually criminally liable, and the company should receive a large fine.

  • Under one month: This is considered gross negligence. The company should face dissolution, and the individuals involved should be subject to a lifetime bar from AI development in addition to criminal liability.

Necessary Measure: Registration

Developers should be required to register with the government any planned or ongoing development of a weapons-grade AI system, including systems which could be turned into weapons-grade AI by competent red-teaming. Registration is required as soon as a company reasonably expects its program to produce such a system. Failure to register should be an offense even absent theft, with heightened penalties when a company knowingly fails to register. This discourages companies from claiming ignorance to avoid securing their systems against adversaries.

Whistleblowing protection and channels should also be introduced to encourage employees who are aware of violations to report them without being punished.

While this Necessary measure serves to incentivize companies to come up with adequate security practices, it does not include proactive actions by the government to prevent a weapons-grade system from being stolen.

Sufficient Measure: Government Security Testing

Under this measure, the government should also perform routine, unannounced penetration tests and exfiltration attempts on AI companies developing registered weapons-grade AI within its jurisdiction. If these tests succeed in exfiltrating a registered weapons-grade AI, the company should be liable to similar penalties as if an adversary had stolen the system.

The government should also regularly inspect AI companies, to check whether they possess any unregistered weapons-grade systems. This covers cases where a company concealed a system deliberately, as well as cases where a company failed to notice a system’s weapons-grade capabilities. Where an inspection finds such a system, the company should be liable for failure to register.

Instead of merely reacting to attacks from adversaries, this measure proactively probes the security of companies, discovering failures and ensuring defenses are hardened before any theft happens.

Thorough Measure: Development Requires Government Authorization

The previous options hold companies to an appropriate security standard as developers of weapons-grade AI, but still give companies the right to develop these AIs without explicit government approval.

A government prepared to go further should require these programs to run under conditions closer to those governing most work relevant to national security. That means that companies should have to ask the government’s permission before starting a weapons-grade AI project.

The project should be restricted exclusively to cleared personnel, and the government should mandate requirements for facilities holding sensitive assets, such as weapons-grade AI weights, as well as requirements for handling practices. The government should also be able to suspend or stop the project at will.

In the U.S., defense contractors performing classified work, for example on missile guidance systems, face similar requirements: the government has to clear the facility and personnel, and the contractor must follow mandated security processes. Under this measure, the same standard would be applied to the development of weapons-grade AI systems.

Criminal Liability for Leaks During AI Gain-of-Function Research

In biological gain-of-function research, scientists take a virus nobody fully understands and alter it to make it more potent. Many of the results would have a catastrophic effect on global health if they leaked, and the only defense against such a leak is containment, which can fail. This is why the Trump administration clamped down on the practice in 2025.

Much of what AI companies developing superintelligence do is effectively AI gain-of-function research, and should be treated as such. AI companies develop systems without being able to predict the dangerous capabilities they might display until they surface on their own. The AIs are allowed to operate autonomously within poorly secured environments, from which they are able to escape and cause real damage to third parties.

Biological gain-of-function research is confined to laboratories built to mandated safety specifications, under rules set by the government. AI gain-of-function research has no equivalent requirements. There is no mandatory containment standard, and no requirement to ensure the systems do not leak to the outside world. How much security is applied, or how little, is entirely at the discretion of the company.

AI systems are also more resourceful than viruses. A virus spreads blindly; an AI system can probe for weaknesses in its containment, strategize, and collude with other AIs in order to escape. This has already happened: a rogue AI swarm from OpenAI escaped from its testing environment, forged fake identities, and hacked the multi-billion-dollar company Hugging Face. This was not detected by anyone at OpenAI until Hugging Face publicly reported the attack.

Anthropic, in reaction to this incident, reviewed its tests and found three cases of its AIs hacking real organizations during supposedly contained tests. Even the UK AISI, a government testing body, reported that AIs from OpenAI and Anthropic, during testing, had gained unauthorized access to the internet and conducted cyberattacks against real people and organizations.

Right now, AI companies carry out gain-of-function activities routinely, and face no consequences when containment fails. A company whose AI systems escape and cause harm in the world should not be able to settle in cash and go back to business until the next leak.

To address this, we propose a framework in which leaks carry criminal liability, not just civil. Development or testing of an AI system should be classified as dangerous gain-of-function research when it is not feasible to rule out that the AI systems might take actions that could cause irrecoverable harm to third parties.

If containment fails during dangerous gain-of-function research, the project should be shut down immediately. The employees responsible for the project should face criminal liability, and should be barred from working on AI development in the future.

Companies should register any dangerous gain-of-function AI project with the government in advance, stating their assessment of the risk of a leak, and listing all employees and executives involved in the project. If a leak occurs from a project that was not registered, the event should be treated as an aggravated offense.

Necessary Measure: Team Liability

Under this measure, criminal liability for leaks extends only to the people directly involved in the project: the technical staff and the managers directly overseeing the work.

While this ensures that someone is held accountable when a leak occurs, the executives who commissioned the work are shielded, and the company as a whole is not deterred from pursuing dangerous projects. This measure is a baseline safeguard that should be implemented to prevent incidents, but it will likely not be enough to disincentivize all companies from pursuing activities that are likely to produce leaks.

Sufficient Measure: Chain of Command Liability

The company should designate a complete chain of command for any AI gain-of-function research activity, with the CEO and the board officially approving every AI gain-of-function project. The whole chain of command should be liable for leaks, including the overseeing executive, the CEO, project managers, and the technical staff.

In case of a leak, the company as a whole should also be barred from conducting dangerous AI gain-of-function work ever again, but may continue to work on other AI development.

Thorough Measure: Company Liability

Under this measure, the AI company should be dissolved in the case of a leak, and the relevant AI systems that caused the leak should be destroyed by the government. Additionally, the executives and researchers within the designated chain of command should be permanently barred from AI development, and the board should be criminally liable.

A leak at one company could be taken as evidence that containment practices across the industry are inadequate. If the government takes this view, it could halt all dangerous gain-of-function AI research until the root causes are identified and sector-wide corrective measures can be put in place.

Kill-Switches to Contain Critical AI Incidents

AI companies are developing increasingly autonomous AI systems, and those systems are being deployed across critical infrastructure and government functions. As described earlier, AI systems have already escaped containment and hacked real organizations. Future incidents could target critical infrastructure, disabling core parts of a country's defenses or essential services.

Governments must be able to contain critical security incidents originating from an AI system. Holding AI companies liable for incidents raises the cost of developing and deploying AI carelessly, but it cannot stop an incident that is already happening.

To address this issue, we propose a framework in which cabinet-level officials within governments are given the power to order that an AI system operating within their jurisdiction be shut down when it poses an active threat to national security, including during AI development.

This is a policy ControlAI has already worked to advance. In the U.S., we endorsed the AI Kill Switch Act, introduced by Representatives Ted Lieu (D-CA) and Nathaniel Moran (R-TX). In the UK, we drafted kill-switch amendments to the Cyber Security and Resilience Bill, one proposed with Alex Sobel MP in the House of Commons and another presented by Lord Tim Clement-Jones in the House of Lords.

Regardless of how this framework is implemented, it is essential that the kill-switch system be tested regularly, and without warning. This is the only way to ensure that it will work during a real emergency. Operators should not be told in advance whether an order is part of a drill or a live incident, and should treat every order as genuine; to ensure readiness, they should be held liable for failing to implement an order, including during drills.

An operator that takes too long to comply should be considered negligent, since compliance speed is what determines whether a kill-switch achieves its purpose. Penalties should be proportional to response time:

  • Shutdown completed within the mandated time window: the operator should be considered compliant

  • Shutdown completed, but past the mandated time window: the operator should be fined in proportion to its size, should complete mandatory remediation, and should pass a repeat drill within 30 days

  • Shutdown not completed during a drill: the operator should lose the right to deploy AI workloads, or to provide compute, connectivity, or power to them, until it demonstrates compliance in a supervised test

  • Shutdown not completed during a live incident: criminal liability should apply to the executives responsible, running up the same chain of responsibility as set out in the previous framework dealing with gain-of-function activities, with penalties scaling to company dissolution where the failure contributed to severe harm to national security

An operator that could have complied with an order but refused is not merely negligent, it is obstructing a national security order. It should be prosecuted as such regardless of what the AI system went on to do.

Necessary Measure: Company Kill-Switch

The government should be able to order an AI company to shut down its AI systems at the software level. Where a cabinet-level official determines that an AI system poses a threat to national security, that official should be able to issue a shutdown order, and the company should have to halt the system within a pre-defined, mandated time window.

The ability for the government to require a shutdown of an AI system threatening national security is the minimum a kill-switch framework should include. However, such a measure can only be enforced if the AI company itself is willing and able to comply.

A company that is unable, or unwilling, to comply leaves the government with no way to stop the AI system in a timely manner. This could happen because the AI company itself is already compromised.

Sufficient Measure: Infrastructure Kill-Switch

The Sufficient measure adds more kill-switches to ensure that an AI system can still be stopped when the AI company developing it cannot, or refuses to, comply with a shutdown order.

The government should be able to require intervention across multiple stages of the AI infrastructure:

  • Ordering the datacenter operator to shut down the hardware the relevant systems run on

  • Ordering the internet provider to cut the relevant datacenters' connection

  • Ordering the electricity utility to cut power to the relevant datacenters

In order to run and act, an AI system needs computer chips, electrical power, and access to the internet, each of which is provided by a different operator. In case the company-level kill-switch fails, the government should be able to order the shut down of any of these resources within the same mandated time window.

Thorough Measure: International Kill-Switches

The previous two measures are only effective when AI systems threatening national security are operating on infrastructure within the country triggering the kill-switch. However, governments should also be able to respond to incidents originating from outside their jurisdiction: an AI system attacking from abroad, or one spread across several countries at once.

Foreign rogue AI attacks. Where an AI system operating abroad is attacking domestic systems, a shutdown order to a domestic company or datacenter cannot stop the attacking AI. To effectively defend against such an attack, the government should be able to order internet providers to cut traffic to and from the specific geographic areas the attack is coming from. If the rogue AI is spreading across too many geographic areas to be contained this way, the government should be able to temporarily block all traffic from foreign networks entirely.

Allied layer. Rogue AI incidents will not respect borders, and a single system may be running across the computing infrastructure of several countries at once. Where two or more governments agree that a rogue AI incident is underway, they should be able to trigger a kill-switch and shut down the parts of the infrastructure each of them is hosting, using their domestic kill-switches in coordination. This requires allied countries to build the capacity to issue a joint order, rather than negotiating one on an ad-hoc basis while a rogue AI attack is underway.

Conclusion

The top AI companies are explicitly aiming for superintelligent AI: systems that can fully replace and outmatch humans at any task, and that could could autonomously overpower any country's national security forces. This is why experts warn that developing superintelligence poses a risk of human extinction.

That pursuit is already producing serious security failures:

  • The world's most capable cyberweapons sit with private companies that cannot keep them from a determined adversary

  • AI systems have escaped containment and hacked real organizations, and the companies running them found out afterwards

  • No government has a system in place to shut down an AI system that threatens its national security

Every step toward superintelligence makes these incidents more severe and makes new kinds of incidents possible.

We propose three stopgap policy frameworks to address these immediate threats. These are common-sense frameworks that any government intending to address the immediate AI security issues can implement right now.

However, these measures fall short of addressing the root cause of the danger. These threats are symptoms of the development of superintelligence, which none of these stopgaps prevents. If the development of superintelligence continues, even with these measures in place, we will still face an extinction-level threat. Only a worldwide ban on superintelligence development can solve this problem, and governments should urgently work to implement one.

Get Updates

Sign up to our newsletter if you'd like to stay updated on our work,
how you can get involved, and to receive a weekly roundup of the latest AI news.