OpenAI o1 发布

测试时计算成为新的能力杠杆

OpenAI 发布 o1-preview,把更多计算放到回答前的推理过程,在数学、科学和编程任务上显著提高复杂问题求解能力。

时间2024 年 9 月 12 日 级别A · 行业级 组织OpenAI 状态已核验 · 2 个来源
OpenAI o1 系统卡首页,正文介绍以强化学习训练推理过程及其安全评估
OpenAI 于 2024 年 12 月发布的 o1 系统卡首页;它记录的是完整 o1 与 o1-mini 的后续安全评估,而非 9 月预览版的开放界面。 OpenAI

计算机通常为等待道歉。进度条、旋转图标和“正在加载”,都在解释结果为什么还没来。即时响应被当作美德;慢,几乎总是故障。

2024 年 9 月 12 日,OpenAI 发布 o1-preview 与 o1-mini。等待第一次被前沿厂商当作一种能力呈现:有些问题值得在回答之前多花一点计算。OpenAI 描述的训练路线是通过强化学习形成更长的内部推理——尝试、检查、修正;测试时投入的计算增加,部分复杂任务表现继续提高。扩展不再只发生在训练完成以前。模型上线以后,每一道题仍可以获得不同的思考预算。

官方在 AIME、Codeforces、GPQA 等设置上相对 GPT-4o 报告显著提升——厂商评测,不是通用智力证书。竞赛题有可验证答案;开放世界的事实、含糊需求与价值冲突,并不会因为模型多算一会儿就变成数学题。产品账本同样要写清。o1-preview 以预览进入 ChatGPT 与 API,速率限制严格;o1-mini 并行出现,从一开始就按体量与成本拆分。首发缺少网页浏览与文件上传等已有能力,把推理模型关在较封闭的输入里。完整思维链不向用户展开:慢下来不等于可审计证明,听来周密的解释也不一定是真实路径。

旋转图标仍在。除了慢,它开始表示一笔有意支出的计算——在 2024 年秋天,这笔支出仍带着预览、限额与功能裁剪的标签。模型选择开始带上时间价格:即时问答仍适合许多任务;难题可以调用更昂贵的推理。测试时计算没有取代数据、参数和训练算力,它只在旧杠杆旁边添了一根新的刻度。

等待被标价之后,产品开始区分“该快”与“该慢”。客服式短答不应烧掉推理预算;竞赛式难题则可能值得多转几圈。隐藏思维链保护了训练细节,也挡住了审计。用户得到的是更慢的答案与更强的竞赛分数叙事,不是一份可逐步复核的证明。旋转图标的新含义,正是这种交换的可见符号。

安全材料讨论推理模型可能用更长内部过程试探拒绝边界,需要匹配的训练与监测。等待增加成功机会,也增加信任诱惑——人们更愿相信慢下来之后的答案。产品必须抵抗把延迟直接翻译成正确性的冲动。

Computers usually apologize for waiting. Progress bars, spinners, and “loading” explain why a result has not arrived. Immediate response is treated as virtue; slowness is almost always a fault.

On 12 September 2024, OpenAI released o1-preview and o1-mini. Waiting was presented by a frontier lab as a capability: some problems deserve more computation before an answer. OpenAI described training that develops longer internal reasoning through reinforcement learning—trying, checking, revising—with partial gains on complex tasks as test-time compute rises. Scaling no longer happens only before training ends. After deployment, each problem can still receive a different thinking budget.

Official reports showed large gains over GPT-4o on AIME, Codeforces, and GPQA under OpenAI’s settings—vendor evaluations, not a general certificate of intelligence. Contest items have checkable answers; open-world facts, ambiguous demands, and value conflicts do not become math problems merely because the model spends longer. Access belongs in the same ledger. o1-preview entered ChatGPT and the API as a preview with strict rate limits; o1-mini shipped beside it, splitting products by size and cost from day one. At launch it lacked features such as web browsing and file uploads already present elsewhere, keeping first reasoning models inside a more closed input surface. The full chain of thought was not shown: slower is not an auditable proof, and a polished explanation is not necessarily the path that produced the answer.

The spinner remains. Besides slowness, it marks intentional expenditure—and in autumn 2024 that expenditure still carried preview, rate-limit, and feature-gap labels. Model choice acquired a price in time: instant answers still suit many tasks; hard problems can call costlier reasoning. Test-time compute does not replace data, parameters, or training scale; it adds another dial beside those older levers.

Once waiting is priced, products begin to separate “should be fast” from “should be slow.” Short support-style answers should not burn reasoning budget; contest-hard problems may deserve extra turns. A hidden chain of thought protects training detail and blocks audit. Users receive slower answers and a stronger contest-score story—not a step-by-step checkable proof. The spinner’s new meaning is the visible sign of that trade.

Safety materials discuss how reasoning models may probe refusal boundaries with longer internal processes and need matching training and monitoring. Waiting raises success chances and a trust temptation—people prefer answers that arrived slowly. Products must resist translating latency directly into correctness.

展开完整事件档案人物、主题、模型与产品
人物
模型
o1-preview
产品
来源

原始资料

  1. 01Introducing OpenAI o1-previewOpenAI · official
  2. 02OpenAI o1 System CardOpenAI · paper

试试搜索