Who Aligns the Aligners? Brief Legal Thoughts on the “AI Safety” Fights to Come

Dario Amodei, the CEO of Anthropic, has published an essay – We Must Pace The Frontier – in which he writes:

I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom.

But – and there is always a but –

…like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious.

Those risks include, according to some, the complete destruction of the human race.

See, e.g., Eliezer Yudkowsky confidently asserting today that if we do not institute immediate global techno-communism, instituting draconian government control over speech and publication of a type never before seen in any Western society, we are all going to die:

There is no evidence that this will happen. Some proponents of regulation tell us that the only response is the most extreme response available: total state control. There is no evidence that this response is correct, either. Nobody knows the answer and history is no guide, save that apocalyptic predictions about new technologies have, to date, all been wrong.

History does provide a great deal of guidance, however, about the use and misuse of government power. It tells us that the state is in fact likely the worst possible custodian for the most powerful publication and data analysis technologies.

This nowithstanding, to address this risk, Amodei proposes

…building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas.

To wit, regulation.

As my regular readers will be aware, I have been engaged, on behalf of my clients, in legal combat with Internet censors around the world, agencies who think that they have both the standing, the competence, and the right to tell American companies what software they can write and run, for the better part of 18 months. I feel now is an appropriate time to offer my preliminary thoughts.

Amodei’s Proposal

Anthropic is, as David Sacks correctly pointed out on X, free to slow down its research and development efforts into AI at any time, to any extent it wishes. Amodei proposes something else: that everyone slow down together, under supervision. While Amodei initially writes that the “slowdown” should be voluntary, the plan would be to progress to legal regulatory regimes – meaning, this proposal necessarily involves the use of coercive state power – which software developers would be expected to obey on a compulsory basis:

The most effective method of pacing is via regulation that targets all US frontier AI companies, as that covers even those who are unwilling to cooperate voluntarily. (Emphasis added.)

His plan has three elements.

Embedded Evaluators

The first is for “Embedded Evaluators” –

employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes.

There is already a robust industry of third-party “safety” overseers for Web 2.0 – what the House Judiciary Committee has described as the “censorship-industrial complex.” The track record of these entities from the last time around tells us how this arrangement plays out in practice.

“Evaluate this.”

One well-known private actor in this space was the Global Alliance for Responsible Media, or GARM. GARM, a commercial enterprise, described itself as “a voluntary cross-industry initiative created in 2019 to address digital safety.” Among other things, GARM provided a range of policy frameworks and guidelines, among them “the Brand Safety Floor and the Adjacency Standards Framework, which have supported brand owners in their independent development of their own bespoke, brand-specific safety frameworks to ensure that their advertising dollars do not inadvertently support illegal or harmful content that damages their brands.”

Although GARM disbanded in 2024, according to the House Judiciary Committee, during its active period GARM worked with global regulators to pressure companies like Twitter, now X Corp., to wield “significant collective power” to coercively influence Twitter’s moderation decisions, including “silencing President Trump,” and to procure boycotts of the platform if the platform refused to obey.

I fail to see how standing up a new crop of NGOs to perform substantially the same function via the same methods – only, this time with NGO commissars possessing highly sensitive employee-like access to internal systems – will lead to a different or better result.

Democratic Coordination (aka Regulation)

The second is “Democratic Coordination,” whereby

Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.

There are two aspects to this: (a) competition law and (b) content regulation law.

From a competition law standpoint, the problem Anthropic has is simple. Anthropic and OpenAI are the largest players in the AI market, by some distance, and coordinating their policies, procedures, and “standards” with each other risks classification as an unlawful cartel. This would particularly be the case if, for example, the two giants aligned on pricing or terms – say, by conforming their API terms so that anyone who used a model that defected from the standards in the global marketplace (e.g., Kimi, Deepseek) would be ineligible to interact with OpenAI’s or Anthropic’s software.

“Government support” for “legally challenging” coordination is a polite way of asking for an antitrust exemption to allow greater coordination between competitors in the name of “safety.” It is a problem any industry consortium of any type needs to account for and this would be no exception.

From a content regulation standpoint, the language “common safety standards as well as limits on the rate of unchecked AI progress” paints a broad brush. This suggests that Anthropic envisages that practically any industry using its software – from manufacturing, to biotechnology, to news publication and copywriting – will require (a) de novo “safety” standards in relation to non-expressive conduct, and (b) limits on how quickly AI software itself can be developed.

In foreign countries, particularly the United Kingdom and Europe, where national governments have fewer constitutional guardrails on their power, I would expect that both (a) and (b) can be legislated without much difficulty in legal terms. If the ease with which rules like the Online Safety Act and Digital Services Act were implemented is any indication, there should not be terribly much difficulty in political terms, either.

The primary legal problem with this aspect of Anthropic’s proposal is in the United States, particularly with (b) – limits on how quickly software itself can be developed. Software development, software publication, and web hosting are inherently expressive activities. See. e.g., the Bernstein v. United States line of cases, as well as Smith v. California, Cubby v. CompuServe, the fact pattern of Stratton Oakmont v. Prodigy, and the related legislative history around 47 U.S.C. § 230. The differences between the United States and its allies on Web 2.0 date back to our very founding, and in both subsequent caselaw and subsequent statutes, America has chosen to protect that activity from state interference.

To the extent Anthropic and its fellow-travelers intend for this aspect of their plan to restrict American citizens from either (a) developing AI models or (b) using models and published FOSS model weights from China, First Amendment issues are immediately apparent and the weight of the precedent militates against government regulation.

Global Coordination

The third limb of Amodei’s plan is “Global Coordination,” whereby

[t]he US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.

We live in an era in which most of the Western world, with the exception perhaps of the United States, has enacted comprehensive technology regulation statutes focused on yesterday’s tech: chiefly, search and social media.

The last great global effort to regulate publication and communications technology began following a moral panic brought about by the twin shocks of (a) Brexit and (b) the election of Donald Trump to the American presidency in 2016. The result was comprehensive Internet censorship laws in Australia (the Online Safety Act 2019), the United Kingdom (the Online Safety Act 2023), and the European Union (the Digital Services Act), plus perhaps a half-dozen copycats around the world, including Brazil (see e.g. the 2025 revisions to the Marco Civil da Internet by Brazil’s Supreme Court), Singapore, Malaysia, Indonesia, and more – all of which seek to control speech and conduct which (a) lives on American servers and (b) in the United States, on those servers, is constitutionally protected under the First Amendment.

Generally speaking, the censorship regimes of the West choose not to describe themselves as such. That does not mean they are not censorship schemes.

Take the UK, which calls its law the “Online Safety Act” and purports that the law exists to keep the UK safe from the evils of the Internet. Its enforcer, Ofcom, is not a law enforcement agency and has no power, by itself, to remove content or make arrests; it cannot keep anyone safe from anything. All the Online “Safety” Act does is threaten American companies with fines (or, in extremis, jail time if a refusal is direct enough) to keep Internet users safe from ideas and expression of which the British state formally disapproves. All that is required, for most users, to circumvent the entire regime and get all the “unsafe” Internet experience you want is a free VPN with an American exit IP, something Britons presently do by the millions, making the UK one of the top VPN-using nations on the planet.

Broadly speaking, the British regime, like many of these regimes, requires companies to (a) age-verify (i.e. dox) users before they access services, and (b) ensure that users accessing those services cannot see content the Act proscribes. Whilst there are some areas where U.S. and UK speech regulations are in agreement, there are many more areas where they are not – and for many of these areas, speech the Act requires be taken down is explicitly constitutionally protected in the United States. I have written about this at length elsewhere (in addition to actually drafting the quite extensive legislative surgery required to align our two nations’ systems) and do not propose to do so again here.

A “global coordination” framework for AI will be built by the same governments, staffed by the same regulators, and pressured by the same NGOs that built the above. It is possible, even likely, that any global attempts at harmonization will collide at exactly the same point: it will not be possible for an American company to comply with European controls and enjoy the full breadth of their U.S. constitutional rights at the same time, as the rulesets will be drafted incompatibly.

Moreover, the verification problem Amodei mentions with respect to authoritarian states is, in fact, a fatal flaw with any such scheme if it has global pretensions (as the UK Online Safety Act once did); as a certified enjoyer of the defector strategy, I am in a very good position to confirm that, given a single defector who is demonstrably outside of the jurisdiction’s reach, deterrence begins to falter. The higher the stakes, the more likely it is that defection will occur.

Although many U.S. companies do, and absent law reform in America (such as a clear censorship shield law) will likely continue to, comply with foreign censorship regimes out of fear, the only parties against whom such a framework will ever be consistently enforced are the companies which are not judgment-proof in the countries which are most likely to get these laws enacted – global companies which, at least for now, includes not many startups but certainly includes Anthropic and OpenAI. In the United States, AI regulation will be subject to early and doctrinally sound constitutional challenges.

Amodei writes:

We should approach any global pacing decision, especially in the near term, in such a way that protects the lead of the US and its allies.

I do not view this as being particularly realistic. Among America’s geopolitical adversaries, several of which America is at war with (directly or by proxy) and who have every incentive to defect, effective compliance levels will likely approach zero.

Preliminary View: Who Aligns the Aligners?

If I have learned anything from our fight against European censors, it is that regardless of a regulatory regime’s good intentions, regulators are subject to political control. Regulators will, subject to that political control, do political things.

The clearest illustration from my own files is Ofcom’s pursuit of a small, highly controversial American website – a mental health discussion board, called SaSu, with no UK presence, personnel, or assets – which voluntarily geoblocked the entire United Kingdom in July of 2025. Ofcom initially accepted that remediation as resolving the matter.

Within days of Ofcom’s acceptance of my client’s geoblock in October and initial closure of the file, following a coordinated pressure campaign by parliamentarians and activist NGOs, Ofcom reversed its own settled position and reopened the case. Ofcom and its NGO partners then circumvented the geoblock using VPNs, created login credentials from behind that circumvention, and cited the resulting VPN-based access as evidence that the block was inadequate.

In May 2026, Ofcom purported to fine the site £950,000, announcing the penalty through a coordinated, embargoed press rollout; weeks later it escalated further, demanding that the site rewrite its terms of service and force a global logout of every user on Earth to terminate the sessions the regulator’s own circumvention had created. My client, which, by way of reminder, had voluntarily blocked the UK, decided enough was enough, and refused these further demands. On July 21, 2026, 480 days after the file was opened, Ofcom closed it, having collected nothing.

At no point in that sequence was the regulator’s conduct determined by the evidence in its file – evidence which, the file shows, was only able to be obtained from accessing the website around a national geoblock. Nor was the regulator governed by legal reality of American constitutional law, backed by the political reality of American power.

The regulator was, instead, governed by the political mood in its home country. Political pressure created the censorship law to target the site. Pressure opened the case when the censorship law entered into force. Pressure reversed an approved remediation. Pressure produced a fine that everyone involved understood could never be collected. Pressure arising from the absence of any face-saving exit kept the enforcement machine running for sixteen months after the target had lawyered up, stated the American legal position correctly, and the futility of the regulator’s actions became apparent to any legally qualified observer.

The entire enforcement process against SaSu, from pre-enactment lobbying to closing the file, was a single, continuous, political act. That is how a so-called “independent” expert regulator in a modern western democracy will behave under political pressure in what should have been an easy case to close a file on an obscure website. This is also what an “embedded evaluator,” a coordinated standards body, or a global AI compliance regime will do under political pressure, because those bodies will be run by humans, and human beings are (a) fallible and (b) respond to incentives.

It is probable that Amodei’s proposals are already being ingested gleefully by “Online Safety” regulators around the world as they look to expand the reach and remit of the censorship schemes over Web 2.0 that they have spent the last decade building – and which a handful of American clients have spent the past year fighting tooth and nail. It will not take a decade to update their censorship apparatuses to try to regulate yet another area of American tech, nor will it take a decade for the vast advocacy apparatus they have built around “Online Safety” to replace-all and begin pushing an AI safety narrative in legislatures around the United States, and around the world.

Our societies can do better than this; so too could OpenAI and Anthropic, if they chose to, but one suspects that the sort of people manning these companies’ “Online Safety” teams are philosophical descendants, if not professional descendants, of the “Trust and Safety” crowd that once worked at companies like Twitter or Facebook, and later created and/or currently staff the censorship agencies of the West.

It took nearly a decade, and actual sight by the British electorate of the Online Safety Act being implemented, for the British public to turn against that regulation and realize that the British government made a grave policy mistake in enacting it.

If OpenAI and Anthropic choose to adopt formal endorsement of prior restraint as corporate policy, those who would oppose the global regulation of AI must move quickly. The most recent counteroffensive against government censorship of the web took seven years to organize. The counteroffensive against government censorship of AI does not have the luxury of time.

Discover more from Preston Byrne

Subscribe now to keep reading and get access to the full archive.

Continue reading