Claude Computer Use 公测:模型操作图形界面

截图加键鼠动作,把桌面与浏览器交给 Agent 循环

Anthropic 在 2024 年 10 月 22 日将 computer use 以公测形式接入 API:模型根据屏幕截图决定移动光标、点击与键入等动作,在实验性条件下用人类方式操作计算机。

时间2024 年 10 月 22 日 级别A · 行业级 组织Anthropic 状态已核验 · 3 个来源
光标在多层应用窗口上移动,旁侧是截图馈入的抽象决策环
AI Chronicle 原创插图:Computer Use 让模型通过截图与键鼠动作操作界面。 AI Chronicle

2024 年 10 月 22 日,Anthropic 的发布页上并排放着两件事:升级版 Claude 3.5 Sonnet,以及一项名为 computer use 的公测能力。后者的描述故意写得像人:看屏幕、移动光标、点击按钮、键入文字。演示视频里,模型在标签页之间切换、搜集信息、填写表单——动作由模型生成,环境仍是一台被授权控制的计算机。对已经习惯函数调用与 Tool Use 的开发者来说,这一步的跳变不在“会不会返回 JSON”,而在工具集合突然变成了整个图形界面。

此前,Agent 多半活在开发者预置的工具清单里:搜索 API、日历、数据库、终端命令。清单之外的软件——内部后台、桌面财务工具、只有 GUI 的遗留系统——要么重写集成,要么靠脆弱的 RPA 脚本。Computer use 把截图当作观察,把鼠标坐标与键盘事件当作动作,让同一套多模态模型在“看见像素—选择动作—再看新截图”的环里推进任务。它仍是工具调用的亲戚:客户端负责真正执行与权限;模型提出下一步点哪里、输入什么。差别是动作空间变成了通用 UI,而不是有限 schema。

Anthropic 同时发布的研究说明并不回避难看的部分。能力被标为实验性:迟缓、易错,录演示时也会出现误点停止录制、突然去翻黄石公园照片之类的走神。在 OSWorld 一类评测上,官方称当时约 14.9% 的成绩高于同类别其他模型的约 7.7%,却远低于人类常见的 70–75%。数字会随版本变化,但数量级提醒仍然有效:公测打开的是接口与叙事,不是已经可靠的数字员工。安全讨论随之升级——键鼠权限意味着钓鱼页面、提示注入与恶意工作流可能造成真实点击与真实发送;沙箱、人工确认与最小权限不再是可选项。

把日期放回工具史,脉络清楚。OpenAI 的函数调用(2023)与 Claude Tool Use 正式可用(2024 年 5 月)先把结构化参数变成行业默认;Responses API 与 Agents SDK 稍后继续打包内置工具与追踪。Computer use 是同一条线上的激进外延:当“工具”等于“人类已经会用的软件界面”,集成成本下降,失控面也扩大。它与仅在浏览器里跑的插件不同,目标是更一般的桌面控制隐喻;与传统 RPA 不同,策略由大模型根据像素上下文现场生成,而不是只回放录制脚本。

对工程团队,落地清单很具体:隔离虚拟机、审计日志、禁止无确认的转账与邮件发送、限制可访问的账号。对产品叙事,桌面 Agent 从此成为前沿模型必须回答的问题——不是会不会聊天,而是敢不敢、能不能在屏上替人点下去。2024 年 10 月 22 日交付的是公测入口与明确的不成熟声明。光标动了起来;责任边界必须一起被画出来。

On 22 October 2024, Anthropic’s announcement page put two items side by side: an upgraded Claude 3.5 Sonnet, and a public beta called computer use. The latter was described in deliberately human terms—look at a screen, move a cursor, click buttons, type text. In the demo, the model switched tabs, gathered information, and filled a form. The model proposed the actions; the environment was still a computer under authorized control. For developers already used to function calling and Tool Use, the jump was not “can it return JSON?” It was that the tool set suddenly became the whole graphical interface.

Until then, agents mostly lived inside developer-supplied inventories: search APIs, calendars, databases, shell commands. Software outside the list—internal admin panels, desktop finance tools, GUI-only legacies—needed a new integration or brittle RPA. Computer use treats screenshots as observation and pointer coordinates plus keystrokes as actions, so the same multimodal model advances a task in a loop of see pixels, choose an action, see a new screenshot. It remains kin to tool calling: the client executes and enforces permission; the model proposes where to click and what to type. The difference is that the action space is a general UI, not a finite schema.

Anthropic’s companion research note did not hide the ugly parts. The capability was labeled experimental: slow, error-prone; while recording demos the model could mis-click and stop a recording, or wander off into photos of Yellowstone. On OSWorld-style evaluations the company reported about 14.9% versus roughly 7.7% for the next model in the same category, still far below human scores often cited near 70–75%. Numbers move with versions; the order of magnitude still holds: the beta opened an interface and a narrative, not a reliable digital employee. Safety talk escalated with the keys—pointer permissions mean phishing pages, prompt injection, and hostile workflows can produce real clicks and real sends. Sandboxes, human confirmation, and least privilege stop being optional.

Placed on the tool timeline, the line is clear. OpenAI’s function calling (2023) and Claude Tool Use GA (May 2024) first made structured arguments an industry default; the Responses API and Agents SDK later packaged built-in tools and tracing. Computer use is a radical extension of the same line: when “tool” means “software a human already knows how to use,” integration cost falls and the blast radius grows. It is not only a browser plugin; the metaphor is broader desktop control. It is not classical RPA alone; policy is generated on the fly from pixel context by a large model, not only a replayed recording.

For engineering teams the checklist is concrete: isolated VMs, audit logs, no unconfirmed transfers or outbound mail, tight account scope. For product narratives, desktop agents became a question frontier models must answer—not only whether they chat, but whether they may, and can, click on a user’s behalf. What shipped on 22 October 2024 was a public beta door and an explicit immaturity claim. The cursor started moving; the responsibility boundary had to be drawn with it.

展开完整事件档案人物、主题、模型与产品
人物
模型
claude-3.5-sonnet
产品
claude
来源

原始资料

  1. 01Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 HaikuAnthropic · official
  2. 02Developing a computer use modelAnthropic · official
  3. 03Computer use documentationAnthropic Docs · official

试试搜索