Constitutional AI:用原则与 AI 反馈训练无害性

一份“宪法”加 RLAIF,减少对有害输出人工标注的依赖

Anthropic 提出 Constitutional AI:用人类写定的原则列表让模型批评与改写自身输出,再以 AI 反馈做强化学习(RLAIF),在减少有害性人工标签的同时训练更少回避、更能解释拒绝的助手。

时间2022 年 12 月 15 日 级别A · 行业级 组织Anthropic 状态已核验 · 2 个来源
一份原则卷轴照亮模型自我批改的双栏回答草稿
AI Chronicle 原创插图:Constitutional AI 用原则与 AI 反馈塑造无害行为。 AI Chronicle

给语言模型标“有害”,从来不只是勾选框。标注员要读威胁、仇恨与越界请求,在疲劳中保持一致;实验室要为规模付费,还要把人反复暴露在恶劣文本里。只优化“有帮助”容易在危险提示上配合过头;强推“无害”又常把助手拧成过度拒绝、顾左右而言他。2022 年的对齐现场,卡在有用与安全的拉锯,以及人类反馈流水线的吞吐上限。

2022 年 12 月 15 日,Anthropic 在 arXiv 发布 Constitutional AI: Harmlessness from AI Feedback。方法直接叫 Constitutional AI:人类不逐条标注哪些输出有害,而是先写下一系列原则——一份可阅读、可争论的“宪法”。监督阶段,模型按原则批评自己的回答并改写;强化学习阶段,由 AI 依据同一套原则比较候选回答,形成偏好信号(作者称为 RLAIF)。人类监督的重心上移到原则文本与评估设计,而不是无限扩张有害标签集。

论文报告:在减少有害性人工标签的同时,可以训练更少回避、更能解释拒绝理由的助手。这是发布方在其设定下的结果,不是自动的安全证书。“宪法”没有法律效力;原则写错、写偏或相互冲突时,模型会忠实地放大文本里的偏见。人类也没有退出——他们写原则、设计流程、做抽检与红队。把对齐部分外包给模型互评,是在承认人类注意力是稀缺资源;稀缺没有消失,只是换了位置。

读这篇论文,应同时看见方法承诺与边界。原则是可修改的人工制品,不是道德终点。后来 Claude 产品的语气与拒答风格,不能倒推成 2022 年论文已经证明的全部。方法给出一条减轻有害标注负担的路径;路径是否走得稳,仍取决于原则质量、评估协议,以及是否还有人愿意对失败样本负责。

原则列表可以被公开讨论、版本化、甚至被批评为偏颇——这正是方法相对于黑箱人工标签的优势与风险。优势是监督对象上移,风险是原则作者的价值被放大。RLAIF 没有取消人类,它改变了人类介入的位置。读 2022 年论文,应把它当作对齐工程的一次分工调整,而不是安全问题的终局声明。

Labeling language-model outputs as “harmful” was never only checkboxes. Annotators read threats, hate, and boundary-crossing requests and try to stay consistent while tired; labs pay for scale and expose people repeatedly to ugly text. Optimize only for helpfulness and models over-comply on dangerous prompts; push harmlessness hard and assistants become evasive. In 2022 alignment work sat between helpfulness and safety, and at the throughput limit of human feedback pipelines.

On 15 December 2022, Anthropic posted Constitutional AI: Harmlessness from AI Feedback on arXiv. The method is named plainly: humans do not label every harmful output one by one; they first write a list of principles—a readable, arguable “constitution.” In a supervised stage the model critiques and revises its own answers against those principles; in a reinforcement-learning stage an AI compares candidate answers by the same principles and supplies preference signal (RLAIF). Human supervision moves upward to the principle text and evaluation design rather than endlessly expanding harm-label sets.

The paper reports that, with fewer human harm labels, one can train assistants that refuse less evasively and explain refusals better. Those are results under the authors’ settings, not automatic safety certificates. A “constitution” has no legal force; wrong, skewed, or conflicting principles are faithfully amplified. Humans do not exit—they write principles, design pipelines, sample-check, and red-team. Outsourcing part of alignment to model-on-model critique admits that human attention is scarce; scarcity does not vanish, it changes address.

Read the paper for both promise and bound. Principles are editable artifacts, not a moral terminus. Later Claude product tone and refusal style must not be back-propagated as everything the 2022 paper already proved. The method offers a path that lightens the burden of harm labeling; whether the path holds still depends on principle quality, evaluation protocol, and whether someone still owns failure cases.

A principle list can be debated in public, versioned, even criticized as skewed—that is both the method’s advantage over black-box human labels and its risk. Advantage: supervision moves up. Risk: authors’ values are amplified. RLAIF does not cancel humans; it changes where humans intervene. Read the 2022 paper as a reallocation of labor in alignment engineering—not a final declaration that safety is solved.

展开完整事件档案人物、主题、模型与产品
人物
Yuntao BaiSaurav KadavathAmanda AskellDeep GanguliJared KaplanDario Amodei
模型
产品
来源

原始资料

  1. 01Constitutional AI: Harmlessness from AI FeedbackarXiv · paper
  2. 02Constitutional AI — Anthropic research overviewAnthropic · official

试试搜索