OpenAI 自研推理芯片 Jalapeño 公布测试成绩

Hot Chips 披露每瓦性能超英伟达 GB200/GB300,年底前开始部署

OpenAI 在 Hot Chips 大会公布自研推理芯片 Jalapeño 的测试成绩:峰值吞吐下每瓦性能为参照系统(英伟达 GB200/GB300)的 1.5~1.9 倍,交互式工作负载优势达 2.1~4.1 倍;芯片由 OpenAI 与博通联合开发,2026 年底前开始在自有基础设施部署,与英伟达并行采购而非替代。

时间2026 年 8 月 25 日 级别B · 领域级 组织OpenAINVIDIA 状态已核验 · 3 个来源
编辑插图:一块发光的芯片悬浮在服务器机架上方,芯片表面有辣椒轮廓的散热纹路
AI Chronicle 原创插图:悬浮的芯片与辣椒纹路,对应 Jalapeño 的命名与推理定位。 AI Chronicle

2026 年 8 月 25 日,OpenAI 在 Hot Chips 大会公布了自研推理芯片 Jalapeño 的测试成绩:峰值吞吐下每瓦性能为参照系统(英伟达 GB200/GB300)的 1.5~1.9 倍,高度交互式工作负载优势达 2.1~4.1 倍。测试覆盖 GPT-OSS 120B、DeepSeek R1 670B、Kimi K2.5 1T 等多厂商模型——这个细节值得单独读:Jalapeño 不是只对自家模型优化,适配性被当作卖点公开。

先看背景。Jalapeño 是 OpenAI 与博通 6 月 24 日联合发布的定制推理 ASIC,从设计到流片仅 9 个月,用于 ChatGPT、Codex 等模型上线后的推理阶段而非训练。当时只有工程样片与「每瓦性能优于现有最先进产品」的模糊说法;Hot Chips 的披露把「自研芯片」从战略叙事变成可核验的测试数据。额定功耗 700W、实测持续功耗 550W 以下,配备 216GiB HBM4、内存带宽 15.4TB/s;架构放弃 prefill/decode 分离,采用统一芯片池与乱序核心,减少数据移动。

需要带着厂商自述的边界读这些数字——Hot Chips 是行业会议,但测试由 OpenAI 自己设计与呈现。不过方向是清楚的:推理专用 ASIC 可以在每瓦性能上超过通用加速器,而每瓦性能正是推理成本的核心变量。SemiAnalysis 的独立基准也显示其在几乎所有场景下每瓦吞吐量领先英伟达、AMD、Google 的芯片——独立分析师的结论与厂商自报方向一致,这比任何单一来源都更有分量。

更值得读的是定位。OpenAI 明确表示与 NVIDIA、AMD 并行采购,而非替代——「并行采购」四个字说明算力多元化的本质是成本策略与议价筹码,不是对英伟达的宣战。2026 年底前开始在自有基础设施部署,第二代已深度开发、第三代初具雏形。对开发者,推理成本可能随自研芯片部署进一步下降,API 定价空间打开;对行业,每瓦性能成为推理芯片竞争的新标尺,自研 ASIC 路线获得数据背书。

Jalapeño 的意义不在单个数字。它把「自研芯片」从战略叙事变成可核验的测试数据,证明推理专用 ASIC 可以在每瓦性能上超过通用加速器;同时「并行采购」的定位说明算力多元化是成本策略,而非替代宣言。对开发者,推理成本可能进一步下降;对行业,每瓦性能成为新标尺,英伟达的定价权面临长期压力。芯片的名字是墨西哥辣椒——辣不辣,看数据说话。

On August 25, 2026, OpenAI posted benchmark results for its Jalapeño inference chip at Hot Chips: 1.5–1.9x the per-watt performance of the reference Nvidia GB200/GB300 systems at peak throughput, and 2.1–4.1x on highly interactive workloads. Tests covered GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T—a detail worth reading on its own: Jalapeño is not tuned only for OpenAI's own models, and broad compatibility is being sold as a feature.

On background: Jalapeño is a custom inference ASIC co-developed with Broadcom, unveiled June 24, nine months from design to tape-out, serving post-deployment inference for ChatGPT and Codex rather than training. At the time there were only engineering samples and a vague "better per-watt than the best available" claim; the Hot Chips disclosure turned "self-designed chip" from strategic narrative into verifiable test data. Rated at 700W with sustained test power at or below 550W, it packs 216GiB of HBM4 at 15.4TB/s; the architecture abandons prefill/decode separation for a unified chip pool with out-of-order cores, reducing data movement.

Read these numbers with the vendor-self-reported caveat—Hot Chips is an industry conference, but the tests were designed and presented by OpenAI. The direction is nonetheless clear: inference ASICs can beat general accelerators on per-watt performance, and per-watt is exactly the core variable of inference cost. SemiAnalysis's independent benchmarks also show it leading Nvidia, AMD, and Google chips on per-watt throughput in nearly every scenario—independent analysts and vendor claims pointing the same way carries more weight than any single source.

The positioning deserves even more attention. OpenAI explicitly says it buys alongside Nvidia and AMD, not instead of them—"buy alongside" reveals that compute diversification is fundamentally a cost strategy and bargaining chip, not a declaration of war on Nvidia. Deployment in OpenAI's own infrastructure starts before year-end, with Gen 2 deep in development and Gen 3 taking shape. For developers, inference costs may fall further as self-designed chips deploy, opening API pricing room; for the industry, per-watt performance became the new yardstick for inference chips, with the ASIC route gaining data-backed credibility.

Jalapeño's meaning is not in any single number. It turned "self-designed chip" from strategic narrative into verifiable benchmark data, showing inference ASICs can beat general accelerators on per-watt performance; the "buy alongside" positioning also frames compute diversification as a cost strategy, not a replacement declaration. For developers, inference costs may keep falling; for the industry, per-watt became the new yardstick, putting long-term pressure on Nvidia's pricing power. The chip is named after a chili pepper—how hot it is, the data will tell.

展开完整事件档案人物、主题、模型与产品
人物
模型
产品
来源

原始资料

  1. 01OpenAI's Jalapeño – update (Hot Chips)Jon Peddie Research · report
  2. 0236氪:OpenAI 自研推理芯片性能超越英伟达36氪 · report
  3. 03澎湃新闻:OpenAI 推出首款芯片澎湃新闻 · report

试试搜索