Google 发布 TPU

为深度学习定制的第一代专用芯片

Google 在 I/O 2016 公布自研 Tensor Processing Unit,这是一款针对神经网络推理优化的专用芯片,已在 AlphaGo 与搜索排序等场景使用。TPU 标志着大公司开始为深度学习定制硬件。

时间2016 年 5 月 18 日 级别B · 领域级 组织Google 状态已核验 · 1 个来源
张量核心发光的定制芯片插画
TPU 是谷歌为深度学习量身定制的芯片,从此巨头开始自研 AI 算力。 AI Chronicle

2016 年 5 月,Google I/O 大会。台上出现了新东西:Tensor Processing Unit。它不是一颗面向消费者的芯片,而是 Google 为神经网络推理定制的专用处理器。对台下观众来说,这只是发布会的一个环节;对算力行业来说,这是一个信号——最大的互联网公司开始自己造 AI 芯片了。

Google 为什么要造自己的芯片?答案藏在数据中心的账单里。2015 年前后,深度学习已经从论文走向生产:语音识别、搜索排序、街景处理,都在吃 GPU 算力。通用 GPU 功耗高、成本贵,而神经网络推理的运算模式其实非常固定——Google 完全可以设计一颗只做这一件事的芯片。TPU 就是这种「为特定工作流定制硅片」思路的产物。

TPU 也不是凭空出现的。它已经在 Google 内部跑了很久:AlphaGo 与李世石那场著名的对弈,背后的模拟就有一部分跑在 TPU 上;搜索排序、街景文字识别也早已在用。到 I/O 公布时,它不是实验室原型,而是一颗经过生产检验的芯片。Google 宣称它的推理性能相比同期 GPU 有数量级提升,并且功耗更低。

TPU 的意义远远超出 Google 一家公司。它向整个行业证明:为 AI 造专用芯片是可行且划算的路线。此后英伟达不断加固自己的 GPU 护城河,而 Google、亚马逊、微软,乃至后来的中国厂商,都纷纷走上自研 AI 芯片的道路。算力竞赛从此不只是「买多少 GPU」的问题,而变成了「谁的芯片更适合 AI」的问题。

回看 2016 年的这颗 TPU,它像一颗投入水中的石子。涟漪先是扩散到搜索与语音业务,再扩散到整个数据中心,最终扩散到全球 AI 芯片产业。今天大模型训练动辄上万张卡,算力成为最稀缺的资源之一,而这颗小小的定制芯片,正是这场军备竞赛最早的注脚之一。

In May 2016, at Google I/O, a new piece of hardware appeared: the Tensor Processing Unit. Not a consumer chip, but a processor Google had built specifically for neural-network inference. To the audience it was one segment of a keynote; to the compute industry it was a signal—the largest internet company had started building its own AI chip.

Why would Google build its own chip? The answer sat in data-center bills. By around 2015 deep learning had moved from papers to production: speech recognition, search ranking, and Street View were all consuming GPU compute. General-purpose GPUs drew power and cost money, yet neural-network inference has a very fixed compute pattern—Google could design a chip that does only this one thing. The TPU was the product of that "custom silicon for a specific workload" mindset.

Nor was the TPU born in a vacuum. It had been running inside Google for a long time: much of the simulation behind AlphaGo's famous match against Lee Sedol ran on TPUs; search ranking and Street View text recognition already used it. By the time it was announced at I/O, it was not a lab prototype but a chip proven in production. Google claimed order-of-magnitude inference gains over contemporary GPUs with lower power.

The TPU's significance reaches far beyond one company. It proved to the whole industry that building specialized chips for AI is viable and worth it. Nvidia has since hardened its GPU moat, while Google, Amazon, Microsoft, and later Chinese vendors all walked down the custom-AI-chip path. The compute race stopped being "how many GPUs can you buy" and became "whose chip fits AI better."

Looking back at that 2016 chip, it is like a stone dropped into water. The ripples spread first to search and speech, then to the whole data center, and finally to the global AI-chip industry. Today large-model training runs on tens of thousands of cards and compute is among the scarcest resources; that small custom chip is one of the earliest footnotes to this arms race.

展开完整事件档案人物、主题、模型与产品
人物
模型
产品
来源

原始资料

  1. 01Google TPU blogGoogle Cloud · official

试试搜索