Devin 把“软件工程 Agent”带入大众视野

一个产品同时操作编辑器、终端和浏览器

Cognition 公布 Devin,展示一个在沙箱中规划任务、编辑代码、运行命令、浏览文档并提交结果的软件工程 Agent。

时间2024 年 3 月 12 日 级别B · 领域级 组织 状态待补来源 · 1 个来源
Cognition 联合创始人兼 CEO Scott Wu 出现在 Devin 官方发布视频中
Cognition 在 Devin 官方发布材料中介绍其联合创始人兼 CEO Scott Wu;画面来自发布视频,不代表独立性能验证。 Cognition

在 Devin 出现以前,大多数 AI 编程产品都要求开发者保持同步。模型补一段,人继续写;模型回答一个问题,人决定下一步。2024 年 3 月 12 日,Cognition 展示了另一种节奏:把任务交给一个沙箱里的 Agent,然后等待它带着结果回来。

发布演示把聊天、编辑器、Shell 和浏览器同时放在屏幕上。Devin 可以拆解任务,阅读文档,修改代码,运行命令,再根据结果继续。这些能力单独看都不陌生,组合后的界面却改变了产品单位。用户不再只向模型索要一段代码,而是试图委托一段持续的软件工程过程。等待不再意味着页面卡住,而成为界面主动安排的时间。

Cognition 同时报告 Devin 在 SWE-bench 上解决了 13.86% 的任务。相较当时的一些基线,这个数字足以制造冲击;相较“AI 软件工程师”的名称,它又提醒人们大多数任务仍未解决。随后围绕比较方法、可复现性和营销表达出现争议。数字不能被删掉,争议也不能被放在脚注里——两者共同构成了首发时能力与想象之间的张力。

沙箱是这项承诺的重要边界。它让 Agent 可以运行命令、打开浏览器和修改工程文件,而不必直接接管开发者本机;同时也把环境准备、数据访问和任务验收变成产品问题。一个补丁在沙箱里通过测试,不等于它符合真实业务要求。最终接收工作的人仍要查看 diff、日志、测试范围以及需求是否被误解。

Devin 的影响因此不取决于首版是否兑现了宣传中的每一种工作。它让“能不能把 Issue 直接交给 Agent”进入工程团队、创业公司和模型厂商的共同语言。此后,编程助手不只比较生成质量,还要回答能否规划、能否使用环境、失败是否可见、结果怎样交付。

品类可以在一支发布视频里迅速成形,信任却不会。信任来自一次次可复现的任务,也来自用户有权拒绝一个看似完成的结果。Devin 把等待设计成了功能;随后整个行业必须证明,那段等待结束时,交回来的究竟是工作,还是另一项需要人重新完成的任务。

Before Devin, most AI-coding products kept the developer synchronized with the model. The model completed a passage and the person kept writing; it answered a question and the person chose the next step. On 12 March 2024, Cognition demonstrated a different rhythm: hand work to an agent in a sandbox, then wait for it to return with a result.

The launch demo placed chat, editor, shell, and browser on the screen together. Devin could break down a task, read documentation, edit code, run commands, and continue from the output. None of those actions was entirely unfamiliar on its own. Their combination changed the unit of the product. A user was no longer asking only for code, but trying to delegate a sustained piece of software-engineering work. Waiting stopped looking like a frozen interface and became time deliberately designed into it.

Cognition also reported that Devin solved 13.86% of SWE-bench tasks. Against some baselines of the time, the number was striking. Against the phrase “AI software engineer,” it was a reminder that most tasks remained unsolved. Disputes followed over comparison methods, reproducibility, and marketing presentation. The figure should not be removed, and the controversy should not be hidden in a footnote. Together they captured the tension between demonstrated capability and the imagination surrounding the launch.

The sandbox was a critical boundary of the promise. It let the agent run commands, browse, and alter project files without directly taking over a developer's machine. It also turned environment setup, data access, and acceptance into product problems. A patch passing tests inside a sandbox did not prove that it matched a real business requirement. Whoever received the work still had to inspect the diff, the logs, the test coverage, and whether the task had been misunderstood.

Devin's influence therefore did not require its first version to fulfill every interpretation of the pitch. It put “can we hand the issue directly to an agent?” into the shared vocabulary of engineering teams, startups, and model providers. Coding assistants would now be compared not only on generation quality but on planning, environment use, visible failure, and delivery.

A category can form quickly in a launch video; trust cannot. Trust accumulates through reproducible tasks and through a user's ability to reject a result that merely looks finished. Devin designed waiting as a feature. The industry that followed had to prove whether, at the end of that wait, the system returned completed work—or simply handed the human a new task to redo.

展开完整事件档案人物、主题、模型与产品
人物
模型
产品
devin
来源

原始资料

  1. 01Introducing DevinCognition · official

试试搜索