DeepMind 发布 Gato 通才智能体
一个模型同时玩电子游戏、控制机械臂、写文字
DeepMind 发布 Gato,一个用单一 Transformer 权重同时完成 604 项任务的多模态智能体,从电子游戏、机械臂控制到对话与图像描述。它被视为「通才智能体」路线的标志性实验。
2022 年 5 月,DeepMind 发布了 Gato,一个试图「什么都会一点」的智能体。论文标题直白地写着「A Generalist Agent」,它用单一 Transformer 的权重,同时处理 604 项任务——从玩 Atari 游戏、给图像写描述,到控制真实的机械臂完成实体操作。在当时的语境里,这几乎是一个宣言:也许我们不需要为每个任务造一个模型。
Gato 的技术路径并不玄妙,甚至可以说朴素:把文本、图像、动作这些不同模态统一成 token 序列,用 Transformer 一起训练。它没有在每个单项任务上碾压专用系统,Atari 玩得不如专门的强化学习智能体,机械臂控制也说不上精湛。但它的意义恰恰在于「广度」——一个模型,多种能力,这本身就构成了一种路线证明。
在 Gato 之前,视觉、语言、控制大体上分属不同的模型、不同的研究社区。Gato 是「通才智能体」路线一次高调的公开演示:把感知、语言与行动统一进一个模型,是可行的、值得认真对待的方向。论文发布后,关于「这是不是走向 AGI 的路线」的讨论迅速蔓延,虽然 Gato 离通用智能还很远,但它让这个话题从口号变成了可以讨论的技术路线。
后来的一切证明这条路线没有消失,而是不断进化。GPT-4o 把视觉、听觉、文本装进一个模型,SeedRealtime 把音视频对话做成全双工,Gemini Robotics 把多模态模型接进机器人身体——它们与 Gato 一脉相承,只是规模与技术都换了几代。
回看 Gato,它像一次「全都要」的实验:不追求单项第一,而是证明统一的模型可以同时触碰很多世界。当多模态大模型与具身智能在 2026 年已经成为主流叙事时,2022 年那个什么都会一点的智能体,正是这条长路的早期路标。
In May 2022 DeepMind released Gato, an agent that tried to be "a bit good at everything." The paper's title said it plainly—"A Generalist Agent"—and it used the weights of a single Transformer to handle 604 tasks, from playing Atari games and captioning images to controlling a real robot arm for physical manipulation. In that era's context it was nearly a declaration: perhaps we do not need a separate model for every task.
Gato's technical path was not mysterious; if anything it was plain. Unify text, images, and actions into token sequences and train a Transformer on all of them together. It did not dominate any single task—Atari play fell short of dedicated reinforcement-learning agents, and the robot control was hardly refined. But its meaning lay precisely in that breadth: one model, many abilities, a route demonstrated rather than a record broken.
Before Gato, vision, language, and control largely lived in separate models and separate research communities. Gato was a high-profile public demonstration of the generalist route: unifying perception, language, and action in one model is feasible and worth taking seriously. After the paper, debate about "is this the road to AGI" spread quickly—even though Gato was far from general intelligence, it turned that topic from a slogan into a discussable technical route.
Everything that followed shows the route never vanished; it kept evolving. GPT-4o packed vision, audio, and text into one model; SeedRealtime made audio-visual dialogue full-duplex; Gemini Robotics wired multimodal models into robot bodies. They are all descendants of Gato's lineage, with scale and technique several generations ahead.
Looking back, Gato reads like a "have it all" experiment: not chasing first place in any single task, but proving a unified model can touch many worlds at once. By 2026, when multimodal models and embodied AI have become the mainstream narrative, that 2022 agent that was a bit good at everything stands as an early landmark on the long road.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- —
- 产品
- —