AlphaStar 击败职业星际玩家
深度强化学习在即时战略游戏上的胜利
DeepMind 公布 AlphaStar,一个在《星际争霸 II》中击败顶级职业玩家的 AI,展示了深度强化学习在不完美信息、长时决策的实时战略游戏中的能力。
2019 年 1 月,DeepMind 和暴雪公布了一场测试:AlphaStar 以 10:1 的总比分击败两位欧洲职业《星际争霸 II》选手。消息一出,游戏圈和 AI 圈都沸腾了。因为星际争霸和围棋完全不同——围棋信息完全、落子有序,而星际争霸是实时战略游戏:地图有战争迷雾、对手行动不确定、单位以毫秒计移动、一局动辄几千次操作。这是 AI 第一次在如此复杂的即时战略游戏里正面击败人类顶尖选手。
AlphaStar 是怎么做到的?它融合了多种技术:从人类对局中模仿学习获得基础操作,再用强化学习在大量自我对局中不断进化,配合精心设计的策略模型与多智能体训练。它学会了侦查、运营、出兵时机、兵种克制,甚至发展出了人类选手不常用的战术。在公布的比赛中,AlphaStar 的操作速度和 APM(每分钟操作数)远超人类,一度引发「它是不是作弊」的质疑。
质疑并非没有道理。早期版本 AlphaStar 拥有完美视野——它能同时看到整张地图,而人类选手必须依靠侦查。为了公平,DeepMind 后来给 AlphaStar 加了视角限制,并限制它的 APM 到接近人类水平,让它「像人一样」玩游戏。即使在这样的约束下,它依然保持了顶级职业水准。这种「主动给自己设限」的测试方式,本身就是研究严谨性的体现。
AlphaStar 的意义远远超出游戏。它把强化学习从完全信息的棋盘推向了不完美信息、实时决策的真实复杂场景——这种能力与自动驾驶、游戏对战、军事模拟等现实问题高度相关。多智能体训练、模仿学习、实时策略决策,这些方法都在 AlphaStar 中得到了验证。
回看 AlphaStar,它像 AlphaGo 在围棋之后扔出的第二块石头。围棋证明 AI 能赢完美的棋局,星际证明 AI 能在混乱、模糊、实时的世界里赢下战争。当今天的大模型在多智能体协作、Agent 规划上不断进步时,AlphaStar 留下的不仅是几场精彩对局,更是「AI 如何在复杂动态环境里做决策」的早期答案。
In January 2019 DeepMind and Blizzard revealed a test: AlphaStar beat two European pro StarCraft II players 10:1. Game and AI communities erupted. Because StarCraft is nothing like Go—Go is perfect information with orderly moves, while StarCraft is real-time strategy: fog of war, uncertain opponent actions, units moving in milliseconds, and thousands of actions per game. This was the first time AI defeated top human players head-on in such a complex real-time strategy game.
How did AlphaStar do it? It combined multiple techniques: imitation learning from human replays for basic play, then reinforcement learning through massive self-play to evolve, plus carefully designed strategy models and multi-agent training. It learned scouting, economy, build timing, unit counters, and even tactics human players rarely used. In the published matches, AlphaStar's actions per minute far exceeded humans', prompting cries of "is it cheating?"
The criticism was not baseless. The early version had perfect vision—it could see the whole map at once while humans had to scout. To be fair, DeepMind later gave AlphaStar a camera restriction and limited its APM to near-human levels, making it "play like a human." Even under those constraints it stayed at top-pro level. This willingness to handicap itself was itself a sign of research rigor.
AlphaStar's significance goes far beyond gaming. It pushed reinforcement learning from perfect-information board games into imperfect-information, real-time decision-making in genuinely complex scenarios—capabilities highly relevant to autonomous driving, game combat, and military simulation. Multi-agent training, imitation learning, and real-time strategy decision-making were all validated within it.
Looking back, AlphaStar is the second stone DeepMind threw after Go. Go proved AI could win perfect games; StarCraft proved AI could win amid chaos, ambiguity, and real time. Today, as large models advance in multi-agent collaboration and agent planning, what AlphaStar left behind is not just a few spectacular matches but an early answer to "how does AI make decisions in complex dynamic environments."
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- —
- 产品
- —