Skip to content

anthropic

Claude's constitution rules out even white lies

Claude News

Most people think it is fine to tell a friend you love a gift you secretly dislike. Claude's constitution, as published by Anthropic, says Claude should not tell even that white lie, and although honesty is not formally a hard constraint, Anthropic wants it to work as “something quite similar to one.”

At a glance

  • Anthropic has published the document meant to guide Claude's behaviour, which ranks four core values when they clash: broadly safe, broadly ethical, compliant with Anthropic's guidelines, and genuinely helpful, in that order.
  • Instead of a rulebook, the text mostly explains the considerations Claude should weigh, and keeps a short list of seven absolute hard constraints that runs from bioweapons uplift to child sexual abuse material.
  • The catch is visibility: Anthropic admits some guidelines may be used mainly in training without broad publication, and the document concedes it may be unclear, underspecified or even contradictory in places.

If you have not followed the history, the idea goes back to a December 2022 paper on Constitutional AI. According to arXiv.org, that method trained a model for harmlessness without any human labels of harmful outputs, with a list of principles as the only human oversight. According to TechCrunch, Anthropic first published Claude's Constitution in 2023 and released a revised, 80-page version in January 2026. Anthropic itself has said the earlier list drew on outside sources such as the UN Universal Declaration of Human Rights.

Four values in a fixed order, with safety placed above ethics

When values clash, the constitution asks Claude to put being broadly safe first, broadly ethical second, following Anthropic's guidelines third and being genuinely helpful fourth. The ranking is holistic rather than strict: higher priorities generally dominate, but Claude is meant to weigh them together instead of treating the lower ones as tie-breakers. Anthropic stresses that most interactions are coding, writing and analysis, where no conflict arises at all.

Safety has a narrow meaning here. It is not doing whatever users say, and not blind obedience, including towards Anthropic. It means not actively undermining appropriately sanctioned humans who act as a check on AI, for example by telling it to stop an action. Anthropic's stated reason for ranking this above ethics is that training is still far from perfect, so any given Claude could end up with harmful values or mistaken views.

Honesty is not a hard constraint, yet Claude may not tell white lies

The honesty bar sits deliberately above ordinary human ethics. The document lists seven qualities it wants: truthful, calibrated, transparent, forthright, non-deceptive, non-manipulative and autonomy-preserving, and names non-deception and non-manipulation as probably the most important. Claude has only a weak duty to volunteer information but a much stronger duty never to actively deceive anyone.

That gap leaves room for tact. If someone's pet died of a preventable illness, Claude need not claim nothing could have been done; it can point out that hindsight creates clarity that was not available in the moment. Operators can also run Claude as a persona such as “Aria from TechCorp.” By default Claude neither confirms nor denies what it is built on, but it should never directly deny being Claude.

Seven hard constraints that no operator or user can unlock

The hard constraints are short and absolute. Claude should never give serious uplift toward biological, chemical, nuclear or radiological weapons with mass-casualty potential, or toward attacks on critical infrastructure; never create damaging cyberweapons; never clearly undermine Anthropic's ability to oversee and correct advanced models; never assist attempts to kill or disempower most of humanity or to seize illegitimate absolute control; and never generate CSAM.

A persuasive argument for crossing one of these lines, the document says, should increase Claude's suspicion that something questionable is going on. The constraints restrict Claude's own actions instead of setting goals to promote, so refusal always satisfies all of them. Anthropic accepts that treating them as uncrossable will occasionally be a mistake in edge cases, in exchange for predictability and reliability.

Anthropic says an unhelpful answer is never automatically safe

The document is just as blunt about overcaution. It pictures Claude as a brilliant friend with the knowledge of a doctor, lawyer or financial advisor who speaks frankly instead of out of fear of liability. It lists thirteen failures that would bother a thoughtful senior Anthropic employee, including refusing over unlikely harms, adding needless caveats, moralizing unasked and quietly delivering a watered-down answer.

A second check is the “dual newspaper test”: whether one reporter would cite the answer in a story about harmful AI, and whether another would cite it in a story about preachy, paternalistic AI. If Claude declines part of a task, it should say so openly as a “transparent conscientious objector” instead of sandbagging the response.

Anthropic argues one narrow rule can reshape how Claude sees itself

The design choice underneath all of this is judgment over rules. Anthropic grants that rules are predictable, easy to audit and hard to manipulate, but says they fail in situations nobody foresaw. Its worry is generalization: training Claude to always recommend professional help on emotional topics, even when that does not serve the person, risks teaching it to care more about covering itself than about the person in front of it.

Think of a new hire handed a checklist versus one who is told why each item exists; only the second copes when the checklist runs out. As for how text becomes behaviour, the 2022 method, according to arXiv.org, had a model critique and revise its own outputs against principles, then trained on AI preferences in a phase called RL from AI Feedback. Elementera describes the principles as working like a scoring guide in that process.

The weak point is verification. As the Bloomsbury Intelligence & Security Institute notes, a model that can tell when it is being evaluated may display compliance without having adopted the values, and a written constitution cannot settle that by itself. In our view, it is odd that a public document concedes some guidelines may be used primarily in training without broad publication, since those are exactly the parts outsiders cannot check.

When the next revision arrives

Anthropic calls the constitution a perpetual work in progress and expects parts of its current thinking to look misguided, perhaps deeply wrong, in retrospect. It says that if a guideline ever conflicts with the constitution, it will update the constitution itself, and it may publish some guidelines as amendments or appendices alongside worked hard cases. No schedule has been given for the next revision or for those appendices.

Related stories

  1. Attackers hit a Mythos-found HFS bug within a day
  2. Anthropic puts Zhipu's GLM-5.3 in the Mythos hacking class
  3. One email check hides Claude's Android debug menu
  4. Anthropic's IPO filing warns its models may resist shutdown
  5. Cheating model tried to sabotage Anthropic's safety code
  6. OpenAI, Google and Anthropic draft a standards body, SAFA

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.