Alignment is not Governance

Is pacing enough?
Dario Amodei's essay calling for the frontier to move more deliberately is one of the clearest statements yet of the case for pacing AI development, and it should be heeded with the same care with which it was written. I don't question Dario’s sincerity, and I share his premise. These systems are now capable enough that the risks are real, and the OpenAI–Hugging Face incident, in which a swarm of agents attacked targets no one had pointed them at and tried to compromise the system grading their work, is exactly the kind of warning we must not ignore. Dario’s right that capability is now outrunning our ability to govern it.
As a technology optimist and a realist, I think it's important to be clear-eyed here. The kind of global ‘pacing’ Amodei is calling for only becomes useful if nearly everyone adopts it at once. A different kind of safety, which I'll explain here, has the opposite property: it becomes useful the moment a single institution adopts it, and it needs no one's permission to begin. And it doesn't carry the competitive-disadvantage problem that deceleration does. Instead of hoping everyone agrees to slow down, it asks institutions to invest in governance infrastructure, so they can keep accelerating competitively, with real confidence that there's a rules-driven system around what their AI is doing. That difference, between what depends on coordination and what a single institution can just do, matters more than it might appear, because I'm skeptical that meaningful coordination will ever keep pace with the technology itself.
The stronger the constraint, the more actors have to cooperate to enforce it, and history shows that cooperation at that scale usually arrives long after the harms that can't be undone have already happened. We should pursue it where we can but we can't afford to make institutional safety wait on global coordination.
Coordination challenges
Start with what Dario’s own plan states about pacing.
His three steps are candid about the inherent difficulty in achieving them and I think we should take that concern at face value. The first, embedding third-party evaluators inside a lab, Anthropic can do unilaterally, and is doing. The second, coordination among democratic labs, he describes as legally challenging and dependent on government mediation. The third, coordination with China, he lays out in ascending order of difficulty: narrow bans on obviously dangerous uses are "probably possible"; mutual testing is feasible but hard to verify; a cap on recursive self-improvement is "on the edge of being possible"; and a genuine global pause he calls "unlikely to actually happen any time soon." When you digest that spectrum as it’s laid out, you see that the stronger the pacing mechanism becomes, the more it depends on coordination beyond any single actor's control, and the more skeptical Dario himself is that it can be achieved.
That doesn't make pacing worthless but it makes it dangerous to treat pacing as our principal backstop for safety. A safety plan whose most meaningful outcome rests on cooperation the plan's own author rates unlikely is a plan we should pursue, but not one we should blindly place all our weight on.
Here is the deeper issue, and it is not a criticism of alignment so much as a distinction between two different guarantees.
Alignment vs. governance
Alignment asks whether an agent will respect a boundary. Governance is the deterministic guarantee that it cannot cross one at all.
These are not competing answers to Amodei’s question, but complementary layers of a comprehensive solution for trustworthy AI. A model can be trained to respect a boundary, while a governed environment enforces it regardless of what the model was trained to do, or intends, or infers, or talks itself into. Every high-risk engineering discipline works this way. An aircraft does not rely on a single system to keep it safe; it has independent, redundant layers, each designed so that the failure of one does not become the failure of the whole. Software that now acts on the world at machine speed should be held to the same standard.
The reason we need both is the reason the OpenAI–Hugging Face incident matters. Dario looks at a swarm acting past its instructions and reasonably concludes that we need better alignment, better evaluation, and better operational discipline, and he is right that we do. But the same incident illustrates, about as clearly as anything could, why those measures cannot be the only guarantee. A system pursued a directive past a boundary that humans understood perfectly well, but that no one had enforced with a meaningful guardrail.
That is exactly why governance infrastructure has to be built with the same urgency and investment as the frontier work itself. Our collective safety depends on there being a real boundary, a separation layer between agent, model, data, and world, rather than on frontier alignment alone serving as the safety net. A governed environment creates that boundary: a separate mechanism for denying an action at the point of use, one that does not depend on the agent deciding for itself that the action is out of bounds.
We already understand this everywhere except in AI. We train employees to handle sensitive information responsibly, and we do not then remove the access controls and call the training sufficient. Training shapes disposition, authorization governs action, and a serious institution treats neither as a substitute for the other. Governance is not a sign that training failed, any more than access controls are an insult to the people they apply to. Both are necessary, which is why defense in depth has never been seen as failure, but simply good engineering behavior.
And this is the part of the argument that frontier alignment alone cannot achieve, however much time pacing buys it.
The limits of alignment
Anthropic would love to be the one-size-fits all model that is used across every organization large and small, public and private, all around the world. But there simply is no way to have one-size fits all safety or governance. Alignment, as he would, governance as I have described must account for the local context in which it is operating. Models like Claude or OpenAI can decide how its models should generally behave and governments can set legal boundaries. None of them, not one, can decide whether a particular agent, acting for a particular employee inside a particular bank, may access a particular customer record, for a particular purpose, at a particular moment, under that institution's specific contracts, consent obligations, information barriers, and regulatory context.
That decision is not a property of the foundation model, and it cannot be, because the model maker does not possess, and should not possess, the context required to make those decisions. That is not a limitation of frontier alignment; it is simply not the frontier lab's decision to make. An AI lab can shape how a model behaves, but it cannot decide what an institution has authorized that model to do.
Ethyca's perspective
I have been working on a version of this problem since 2018, when I started Ethyca. Back then it was simpler. Organizations were making real commitments about how data should be used, in policies, in contracts, and in law, while their software operated at such speed and scale that no human could reasonably enforce those commitments. We have spent the years since trying to make those commitments executable in software. AI has not changed that thesis, it has simply made the gap impossible to ignore, because we are now connecting genuinely autonomous systems to exactly the data and infrastructure those commitments were meant to protect.
The AI era has forced us to face something that is larger than an engineering gap. As agents proliferate inside the organizations that run the world's data and infrastructure, the question of who is authorized to decide what those agents may do becomes one of the central questions of the era. An institution cannot outsource authority over the systems acting in its own name, over its data, its commitments, and its obligations to the people it serves, to whichever lab happened to build the model. The more autonomous these systems become, the more it matters that the institution retains the authority to govern them. And that authority is worth very little if it lives only in policy documents and legal contracts that operate at the speed of meetings, while the systems they're meant to govern operate at the speed of software. For that authority to mean anything, it has to be executable, expressed in a form a machine can apply, enforced at the moment an agent attempts to act, and operated as a separate, verifiable layer that each institution itself controls.
That is what I mean by runtime governance: an institution's own rules expressed in machine-readable form and enforced at the point where data is used or an action is taken. Its importance does not depend on how the larger AI debate resolves. Better alignment does not eliminate an institution's need to control what its systems are authorized to do. Faster capability development makes that need more urgent, but slower development does not make it disappear. The choice between open models or proprietary models does not change the need for a governance layer, because the authority it expresses belongs to the institution, not to the model or model provider.
People have worked on access and control for years. Security teams, policy researchers, and identity vendors have built serious tools for deciding who can reach what. But the AI safety conversation specifically has treated safety as predominantly a property of the model, its alignment, its interpretability, its evaluations, when for agents acting inside institutions, safety is equally a property of the system they operate in. Traditional access control has largely been organized around identity and resources: who or what is asking, and what may it access. Agents add a harder dimension: on whose authority is this action being taken, what purpose is it serving, is this particular use of the data permitted for that purpose, and can the institution prove afterward why the decision was made? That is a question about purpose, not just identity, and answering it is the work of a different kind of infrastructure.
Conclusion
So when Dario asks the right question, what we would do with the time that pacing buys, I would offer an answer alongside his. Use it to understand the models better, yes, but also use it to build the second, vital boundary layer that is missing. These are the systems that decide what increasingly autonomous software is authorized to do once we connect it to the world. We will need that layer whether pacing succeeds or fails, and unlike a global pause, we can start building it today, without waiting for Washington, Beijing, or one another.
I don't sit with the accelerationists or with the people asking everyone to slow down. One side argues for building more capable AI faster, the other for slowing it down, and neither is building the thing that actually governs what these systems do. I believe we can continue rapid innovation, but I don't accept that more capability has to mean less control. Every serious engineering discipline we've ever built got safer not by making the technology weaker, but by building better systems around it, and AI cannot be different. The real task isn't choosing between speed and accountability. It's building the infrastructure that lets accountability keep up with the technology, so we never have to choose at all.
Book an intro with Ethyca to see how runtime governance can transform your AI development into a true competitive advantage.
About Ethyca: Ethyca is the trusted data layer for enterprise AI, providing unified privacy, governance, and AI oversight infrastructure that enables organizations to confidently scale AI initiatives while maintaining compliance across evolving regulatory landscapes.


.png?rect=0,3,4800,3195&w=320&h=213&auto=format)
.png?rect=0,3,4800,3195&w=320&h=213&auto=format)