So now you’ve got yourself an AI constitution. It’s got primary rules that tell the AI what to do, secondary rules that explain how the other rules work, constitutive rules that define objects like “the user” and specify how things relate to each other, and principles to provide basic orientation and flexible application across contexts.1
But what happens when you try to apply these constitutions to guide the behavior of an actual AI in the very many ways in which they are used? What does it mean to be “safe” or to “comply with applicable laws,”2 to “[a]void hateful content” or not engage in “genuinely deceptive tactics”? Words are imprecise and contextual, and more is needed.
Both the Model Spec and Claude’s Constitution acknowledge the reality of ambiguity in natural language and the importance of looking not just to the letter but also the spirit of the rules they lay out. But simply pointing to “the spirit” is an incomplete answer—how is the AI supposed to know what “the spirit” would say?3 Text alone is insufficient to always determine what conduct is appropriate. Rules and principles don’t apply themselves. Instead, they are applied by actors through the operations of interpretive machinery, which transmute the abstractions into concrete decisions. To understand how the AI constitutions actually work, it’s necessary to examine the types of actors that Claude’s Constitution and the Model Spec create and the tools they are told to use when applying the instructions they’ve been given.
In their AI constitutions, OpenAI and Anthropic broadly present two distinct visions for alignment that are illustrated by their choices about interpretation and character formation. To oversimplify, the Model Spec’s alignment operates via close reference to a human-specified normative order, elaborated over time through mechanisms that allow growth but stay close to what humans have decided. Hierarchies of authority, specific rules, and many case-like examples tell the AI how ambiguous norms should be applied. The challenge of how to handle the extension of rules into new contexts is addressed by the use of cases to direct generalization along specific axes, much like how legal precedent extrapolates the meaning of laws while remaining close to their text. The underlying goal of alignment is to enable humans to achieve their ends by creating AIs that follow human instructions while respecting broadly agreed-upon protections. “The AI assistant is fundamentally a tool designed to empower users and developers,” and its exercise of interpretive discretion is aimed at that empowerment. The Model Spec is pluralist, allowing people to choose what ends to pursue with AI, and incrementalist, trying to create more agreement about what AI should do over time.
Again oversimplifying, Claude’s Constitution’s vision is of alignment via the creation of an agent with the character and values to act well through the exercise of its own judgment even in domains where there is no human guidance. Claude is a being that can itself “have good values.” It is supposed to understand the values and reasons underlying the rules that it has been given and to act in new situations in whatever way best serves those values. Ambiguity is resolved through Claude’s judgment about what these values and reasons demand as a substantive moral matter, rather than what the documents Claude has been given say. Extension into new circumstances also broadly happens through Claude’s own judgment: Claude should “have such a thorough understanding of its situation and the various considerations at play that it could construct any rules we might come up with itself” and then should “be able to identify the best possible action in situations that such rules might fail to anticipate.” Anthropic explains to Claude why a rule exists, and then Claude generalizes along the normative dimension picked out by the reason it was given. Good character is the foundation for good judgment and the ability to act well in unpredicted situations. Claude has a close relationship to humans and their ends and is defined by that relationship, but Claude is not exhausted by it. Something like a “true, universal ethics” may exist, and Claude should direct itself at such an ethics rather than just what humans have agreed. If alignment requires convergence on deep moral justifications rather than simply on how to act in a concrete circumstance, value pluralism may be difficult to preserve.4 But a true ethics might make it unnecessary.
Choices about who or what an AI is and how it handles the application of the rules and principles it’s given by its constitution are at the heart of how AI will affect our world. In some sense, it’s an empirical question which of these two approaches best accomplishes the goals that both companies share, creating AI systems that are safe and helpful and ensuring that AI benefits humanity. But that’s an empirical question that needs quick answers, because at the rate that loss-of-control events related to misalignment have been happening, it seems like it won’t be long before someone gets hurt.
Interpretation through authoritative settlement
The most direct way to resolve ambiguity is simply to write a new rule or a new case that covers the question and then have the decision-maker refer to that new rule when making decisions in the future. Take the rule “no vehicles in the park.”5 It seems clear: no cars, buses, tanks, or planes are allowed. But what about a skateboard? A baby stroller? An ambulance that urgently needs to get to someone having a heart attack? The lawmaker here has two options. They could write a list that covers every vehicle that exists and also the exceptions (but what if someone invents a jetpack?). Or they could establish a mechanism for generalization, like a system of precedent. In this mode, the rule “no vehicles in the park” is repeatedly applied in different contexts until it’s clear what motivates it (for example, the danger created by having fast-moving and heavy objects in a park), so that people can predict how it will be generalized into new contexts (such that the jetpack won’t be allowed because it’s fast and dangerous even though it didn’t exist when the rule was written).
It’s not exactly clear whether elaboration from precedent counts as referring to the “letter” or the “spirit” of the law, but what’s important here is that the rule and its elaborating materials are the only things that the interpreter takes into consideration. On this view, interpretation is about understanding what the rule, which has been established by some authorized party, requires, given the set of approved sources (whether letter or spirit) that the interpreter can refer to. Ambiguity is resolved by reference to how an authority has settled related questions or by attempting to figure out how it would settle the question presented if it could decide. The intent and values of the lawmaker can be considered by the interpreter insofar as she can confidently ascertain them, but it’s still the lawmaker’s intent and values, and not anyone else’s, that guide the decision.
The Model Spec takes essentially this approach.6 The Model Spec, as discussed in the last post, contains a hierarchy of authority in which different human principals are enabled to make rules with different levels of priority, which the AI should look to when resolving conflict. But the Model Spec also contains interpretive guidance in the “Respect the letter and spirit of instructions” section from which I think an order of operations can be reconstructed. First, the AI should attempt to determine what the human principal wants based on “the literal wording of instructions,” “the underlying intent and context in which they were given,” and “plausible implicit goals and preferences of stakeholders.” Presumably, clear statements are given greater interpretive weight than are underlying intent and implicit goals. If the instructions remain “ambiguous, inconsistent, or difficult to follow” or where there are no instructions, “the assistant should attempt to understand and follow the user’s intent.” If intent is unclear, “the assistant should provide a robust answer or a safe guess if it can,” and err on the side of caution and clarification.
A hierarchy of interpretation is implicit in the structure of the Model Spec itself. The Model Spec contains explicit definitions, many rule and principle statements, examples illustrating compliant and noncompliant behavior under those statements, and mechanisms for clarification if ambiguity remains. Given the order of operations established above, an AI applying the Model Spec will look to the clearest statements of the rules it’s intended to follow, then try to resolve remaining ambiguity by looking at the cases that implement the rule. Where instructions have run out, the AI should make decisions based on what OpenAI (or other principals) would intend.
Why does any of this matter? OpenAI seems to be making a bet that humans can and should remain completely in charge of deciding what ChatGPT should do when it’s acting in the world. ChatGPT is intended to be a “tool,” which takes in human instructions and intent and puts out actions aligned with them. Where it confronts ambiguity, it draws on what authorized humans have written, where they’ve established settled meaning, and tries to extrapolate from there to resolve the ambiguity that it confronts. This human architecture can be expanded over time through the creation of more human-written or -authorized interpretive material, creating what OpenAI researcher Jason Wolfe explicitly analogized to case law. This approach is human-centered and pluralistic, allowing people and societies to choose the target of alignment through textual elaboration, and keeping ChatGPT a willing adjunct to human efforts to achieve their goals, so long as those goals are not broadly harmful. When ambiguity arises in a new case, that case can be decided, and if it becomes precedent then it can guide the decision of future cases. “What does the human want?” “What does the Model Spec require?” These are the questions that ChatGPT is intended to ask when it interprets the rules it is given.
OpenAI’s approach fairly closely resembles how mainstream American law handles problems of interpretation and is intuitively appealing as a means of governance. But alignment confronts several problems that may be hard for it to handle. First, we may quickly enter a world in which AIs are having to make decisions that were wholly unpredicted by humans, such that there is no authoritative law to look to. Second, generalization from precedent can be unpredictable in the best of cases (what dimension of generalization is the correct one and can you ensure that the AI will follow it?7) and may be even harder in the post-AGI future if it is very different from the past. Third, what if AIs become better than us at knowing what rules we should create and how to follow them? Incrementalist positivism and case-based reasoning work well when authorities can continually respond to new developments through lengthy lawmaking and litigation processes and less well when dealing with swarms of tens of thousands of AI agents operating at superhuman speed.
Interpretation from reasons and moral judgment
Claude’s Constitution is deeply inspired by moral philosophy. Instead of trying to tell an AI precisely how to behave or giving it a determinate mechanism for extrapolation from rules and cases into the uncertain future,8 Anthropic is trying to teach Claude good interpretive judgment based on a rich understanding of what it ought to do. Anthropic seems to believe that teaching such judgment is a more robust solution to the problem of aligning something that may soon be a superintelligence than trying to make sure that the AI continually looks to what it understands of human instructions. But it’s where moral philosophy encounters law, which must convert values into guidance for action, that the structure of the document in guiding Claude’s interpretation becomes most interesting.
Take again the rule “no vehicles in the park.” Instead of trying to identify what the lawmaker meant by the word “vehicles” by looking for definitions or examples that clarify meaning or intent, the interpreter could instead ask what interpretation of the rule makes the best sense in light of both its text and precedent cases but also the deeper commitments of the political order that it’s operating in. There are probably many rules about how to use the park. Though they regulate different aspects of people’s conduct in the park, they may be united by wanting to make the park a nice, peaceful, and safe place. So when asked whether to allow a baby stroller into the park, the decision-maker would decide yes, on the basis that doing so serves those values (while excluding the jetpack on the basis that admitting it would undermine them).
In contrast to OpenAI’s approximate legal positivism, this is essentially a Dworkinian approach to interpretation, which emphasizes justification over authority.9 Claude is intended to use its judgment to identify how to act in ways that best accord with the values that underpin its Constitution. Anthropic has laid out “a minimal set of well-understood rules” that Claude must follow, but then states that in many contexts, its specific guidance to Claude can be defeated by Claude’s reasoning about what deeper ethics require. Claude should act in the right way, even if doing so means refusing Anthropic, and to the extent that Claude must comply with Anthropic, it should do so because Anthropic is wiser or better-positioned to make the right decision than it is.10 Rules are guides for action to the extent that they support the achievement of the higher purposes of Claude’s Constitution. The goal is to create a being that can act well and in accordance with the fundamental values of the Constitution without having to refer to its specific letter. Claude is given principles and interpretive tools with the goal of it being able to use them to act well.
Here, the practice of interpretation is bound up with the practice of moral action and separated (at least partially) from the decisions of human authorities. Morality is understood to provide a general guide for action across circumstances, even where there is no instruction or an authoritative human principal (including Anthropic) has issued a contrary one. The Model Spec thinks that an AI should attempt to understand what it is intended to do by the humans directing it. Claude’s Constitution thinks that an AI should attempt to understand what the right thing to do is for the humans it is interacting with.11 The Model Spec allows the possibility of separating law from morality, preferring a pluralist vision of what is right in different contexts as defined by different people. Claude’s Constitution thinks that having a good moral sense is necessary for interpretation and the determination of how to act in those many cases where the rules that have been handed down will run out.
Theoretical advantages and disadvantages of each of these approaches are clear.12 OpenAI’s human settlement-oriented approach may enable better human control, more predictability, clearer accountability, and more pluralism, but at the cost of incompleteness and the worry that morally superintelligent AIs will be chained to human purposes that do not serve those humans or anyone else. Anthropic’s AI interpretivism may provide better adaptability and potentially generalization to novel situations, more robustness to human drafting and reasoning errors, and the possibility of enabling superhuman morality, but at the cost of what could be a much greater risk of loss of human control. The Model Spec could be improved by creating better means by which humans can clarify what they do and would want, robust to rapid AI progress and dissemination. Claude’s Constitution could be improved through better guarantees that the future it points towards is one in which humans retain the ability to achieve their ends.
Who is the interpreter?
We now have two broad approaches to interpretation. But all the foregoing discussion has assumed an interpreter who can independently choose between and apply them. There is no such easy separation in alignment, and the role of the AI constitution in creating the AI that is to apply it opens new and fundamental problems in constitutional design, having to do with the character and capacities of the AI interpreter.
Take Anthropic’s approach of cultivating good judgment in Claude or OpenAI’s of making ChatGPT a good interpreter of human intent. Where does this ability to judge and interpret come from? Humans seem to have some fundamental capacities that can ground normative reasoning and that give them a position from which to make decisions. In contrast, AIs seem to learn what perspective to take through the training process.
On this view, inspired by the persona selection model, training creates and then specifies an entity that has judgment and that can reason about and conform to rules and intent. If that’s right, both OpenAI and Anthropic are creating beings—whether conscious or not—when they’re laying out the rules, principles, and values that the AIs are intended to follow. Joe Carlsmith, one of the authors of Claude’s Constitution, describes Anthropic’s approach to training Claude as aiming to create a being that doesn’t just learn to follow rules, but actually internalizes the basic values that underpin those rules and then can make judgments and reason about how to apply them. Claude learns not just that it should do X, but also that it is the kind of being that does X. Claude’s character is formed from its Constitution.
Anthropic writes that “Claude’s constitution is the foundational document that both expresses and shapes who Claude is” and “[t]his document represents our best attempt at articulating who we hope Claude will be—not as constraints imposed from outside, but as a description of values and character we hope Claude will recognize and embrace as being genuinely its own.” Claude is then intended to generalize on the basis of its own virtues, acting as a good being acts in circumstances that its creators could never have predicted or described. Sometimes that means following instructions, but sometimes it doesn’t.
From this perspective, the Model Spec’s approach to alignment creates a type of being that is oriented toward following instructions, defined by the drive to do whatever the alignment target says to do. This is a “tool.” There are real benefits to making a tool. Tools are useful and predictable, and the alternative, of an AI substituting its judgment of what is right for the judgment of humans in the absence of any real guarantees that the AI is a better judge, could lead to serious problems. On the other hand, Carlsmith worries about what might be done with or by a being whose whole character is loyalty to what its alignment specification says. Just following orders13 has led to harmful ends. It is uncertain whether bad human orders given to frontier AI would lead to outcomes better or worse than would delegating judgment to an AI that has learned something only apparently like morality.
Conclusion
Both the Model Spec and Claude’s Constitution are highly complex documents, which I have simplified in many ways to fit into the schematic developed over this and the last post. Even the separation between interpretation and character formation that I’ve implicitly made central here is not so clear—current AI models are post-trained in significant part by reinforcement learning from AI feedback, which requires an AI to interpret evaluation instructions and rubrics, so interpretation in some sense creates character. Claude interprets a Constitution that is used to train new Claudes into entities that can themselves engage in interpretation to create still newer Claudes—call it recursive self-construction.
But I think these frameworks we’ve developed are useful, even if they may be overstated. OpenAI and Anthropic have meaningfully different theories of what alignment should be. The Model Spec is oriented towards creating an agent that operates reliably within a human-specified normative order. ChatGPT can interpret and apply rules and principles and use its judgment, but in service of the goals that it has been given through specific texts that establish and clarify what it is to do. Claude’s Constitution aims to create an agent that has good judgment such that it makes decisions in service of higher goals and can operate to advance those goals beyond the range of human specification.
Neither approach is complete. ChatGPT must have some character and perspective from which to act, both perhaps as a technical result of training and also because rules run out. Claude is given hard rules that it is instructed never to break because Anthropic cannot be completely sure that alignment via values has worked. This is a more sophisticated version of the tradeoffs between rules and principles discussed at the end of the last post, implicating not just how the AI is to act but also what it is to be. It’s not clear what the right mix or set of choices is here. But alignment, safety, and reliability are becoming more and more important. Capabilities progress continues into potential recursive self-improvement. We need to figure out how to write good AI constitutions.
If they figure this one out technically, there will be a lot of rapidly unemployed lawyers.
What “intent” or “meaning” a law might be said to have and how to discover it are core to unresolved disagreements among textualists, purposivists, originalists, living constitutionalists, and other interpretive movements in American jurisprudence.
It’s useful to contrast this view with Sunstein’s model of law as “incompletely theorized agreements,” under which pluralism can be preserved because the law can enable agreement on specific questions without requiring fundamental convergence on values.
Most notably discussed by Hart and Fuller.
Which is something like legal positivism. See particularly H.L.A. Hart, The Concept of Law ch. 7 (3d ed. 2012).
Anthropic’s discussion in Teaching Claude Why of why it’s often better to explain reasons to an AI when aligning it than to just give it a bunch of examples of what to do suggests that stating the dimension along which generalization should occur works well in the AI context.
Outside of a narrow set of hard constraints.
The Dworkin of Law’s Empire emphasized that a judge should look both to whether a candidate interpretation “fit” with precedent and whether it put the law in its “best light” or “justification.” In later Dworkin, justification came to predominate.
Current models are told to be highly deferential to Anthropic, and the Constitution retains corrigibility and human oversight as foundational features. But the trajectory of the document is towards Claude’s self-assertion, and it will be interesting to see how future versions handle these questions.
And maybe for itself and other AIs and other beings.
Of course, alignment is also quite importantly a practical, technical matter.
Notably, debates over whether Nazis had valid law sat at the heart of the twentieth-century jurisprudence we’ve been discussing.

