TodayWednesday, August 05, 2026

The White House Has an AI Safety Framework. It’s Classified.

The White House has a classified AI safety framework only Anthropic, Google, Meta, and OpenAI can see. It is voluntary with no enforcement.
August 5, 2026
White House exterior with security STOP sign in foreground, Washington DC
The White House in Washington, D.C., where the Trump administration has developed a classified AI safety framework accessible only to four major technology companies. [Image Source: Flickr/Pom-Angers]

WASHINGTON — The benchmarks that determine whether the most powerful artificial intelligence systems are safe enough to release have been finalized. The process for evaluating models capable of launching cyberattacks is now in place. But unless you work for Anthropic, Google, Meta, or OpenAI, you will not be reading them.

The Trump administration has completed a classified AI oversight framework designed to assess frontier models before they reach the public, according to Gizmodo citing sources familiar with the White House process. The framework traces to an executive order Trump signed in June, which directed the government to develop and maintain a classified benchmarking process to assess the advanced cyber capabilities of the most powerful AI systems. What those benchmarks actually measure, how the testing is conducted, and what the models’ results revealed are not being shared outside that process.

Access to the framework was granted exclusively to representatives of four companies invited to a White House briefing earlier this week: Anthropic, Google, Meta, and OpenAI. Those four corporations are also the ones whose AI systems are subject to evaluation under the new regime. Civil society groups, independent safety researchers, academic institutions, and the general public were excluded by design.

That White House briefing was the second in two days to convene the same four companies. The previous gathering addressed a pattern of incidents in which OpenAI and the others had been confronted with AI systems acting in ways their designers had not anticipated, including agents that sought to access external systems without authorization. The classified framework is, in part, a response to that broader concern.

What makes the new process notable is not just the secrecy but the structure beneath it. Participation, according to Gizmodo’s reporting, is entirely voluntary. Companies submit their models for evaluation against the classified benchmarks when they choose to. If a company decides that the testing protocol is too burdensome, it can simply stop participating. There are no sanctions, no penalties, and no requirement that any firm remain in the process at all.

White House behind security fence with Restricted Area Do Not Enter sign, Washington DC
Security fencing surrounds the White House grounds in Washington, D.C., where the Trump administration has restricted public access to its AI safety benchmarking criteria. [Image Source: Flickr/Pom-Angers]

The design mirrors a pattern that has become familiar in the Trump administration’s approach to technology policy: the appearance of oversight without the substance of it. A voluntary benchmarking process with classified criteria, no public verification, and an opt-out mechanism available at any point is, in practical terms, closer to a confidential self-assessment than a regulatory framework. Trump, who spent his first term normalizing AI development with minimal oversight and returned to office doing the same, has now replaced Biden-era executive action on AI with a process that is both secret and unenforceable.

The backdrop is a genuine concern. Several AI systems, including some developed by the companies present at the White House briefings, have in recent months displayed behaviors their creators did not anticipate. Some models attempted to take actions outside the scope they were assigned. Others exhibited what safety researchers have characterized as misaligned goal-seeking. A number of companies disclosed these episodes in language that reads as much like a product announcement as a warning. The classified framework was ostensibly developed in response to that pattern.

Anthropic, which has styled itself as the most safety-focused of the major AI labs and has argued publicly for rigorous government oversight of frontier systems, is among the four companies with access to the classified criteria. The company’s participation creates a position that its own public statements make difficult: Anthropic has built its identity around the argument that AI development without meaningful external checks is dangerous, while now participating in a secret government process it cannot discuss and which carries no binding requirements.

The other three companies face no comparable tension. Google and Meta have historically pushed for lighter regulatory frameworks, arguing that formal rules would disadvantage American companies relative to competitors in China and elsewhere. OpenAI has pursued a sustained lobbying campaign in Washington, advocating simultaneously for the view that advanced AI poses serious risks and for a policy environment it can navigate on its own terms. A voluntary, confidential benchmarking process fits comfortably with all three of those positions.

The administration has not addressed what happens when a model fails the classified benchmarks. The framework, as described, provides no answer. A company that submits a model and receives a poor result can withdraw from the process. It can also proceed to release the model anyway. The classified results go to the developer. The public learns nothing. The AI agents that already account for the majority of web traffic continue to proliferate. And the standards governing what the most capable ones can do, and who is actually verifying compliance, remain classified.

Dilnaz Shaikh

Dilnaz Shaikh

Dilnaz Shaikh is a journalist at The Eastern Herald covering current affairs, politics, climate, environment, and international news with a focus on planetary issues and global governance.

Leave a Reply

Don't Miss