GPT-4o 发布
多模态从“外挂能力”变成统一模型
OpenAI 发布 GPT-4o,以同一神经网络处理文本、视觉与音频,让实时语音和视觉交互第一次成为通用模型的原生能力。
GPT-4 时代的多模态,常常是多个模型串起来:语音先转写,再进语言模型,再合成嗓音。延迟与信息损失都写在缝上。用户说完一句话,要等管道里每一节都走完。2024 年 5 月 13 日,OpenAI 发布 GPT-4o——名称里的 o 表示 omni——以同一神经网络处理文本、视觉与音频。
发布账本必须分项。文本与图像能力进入 ChatGPT 与 API;低延迟音频体验以分阶段预览推进,不是所有用户、所有入口同一天完整打开。把 5 月 13 日写成“万物皆已原生且全面可用”,会夸大首发清单。演示里的实时语音往返,展示的是交互设计空间被打开:打断、语气、视觉指认可以落在同一条对话时间线上——仍受预览范围与安全策略约束。
成本与速度也是产品主张的一部分:多模态推理被描述为更便宜、更快,具体数字随价目表与地区策略变化,应回看来源页。对应用开发者,统一模型降低拼接管道的工程量;对安全与隐私,麦克风与摄像头进入默认想象,也扩大了攻击面与误用面。能听能看,不等于更该被默认打开。
GPT-4o 没有结束“多模态是否等于更理解世界”的争论。它改变的是默认产品形态:通用助手开始被期待能听、能看、能在接近人类对话的节奏里回应——而工程与政策仍决定哪些通道真正打开。omni 是目标词;分阶段交付才是首日事实。
把 omni 当成已经完成的能力清单,会忽略预览标签的重量。麦克风与摄像头一旦进入默认想象,产品要同时设计拒绝与确认:何时听、何时看、记录保留多久。实时语音的惊艳来自节奏;节奏背后是工程上能否稳定地低延迟,以及策略上能否在滥用与误用之间划线。文本与图像先全面铺开,音频后到,这个顺序本身也是史实:通用模型的多模态,从来不是所有通道同一天打开。
“接近人类对话节奏”是体验主张,受网络、设备与负载影响。实验室演示与日常高峰时段的体验可以相距很远。把演示节奏写成普遍 SLA,会超出发布账本。多模态产品史要同时记下峰值体验与开放范围。
同一网络处理多模态,降低了级联错误,也让失败模式更难拆开归因:究竟是听觉、视觉还是语言环节出错?可观测性成为多模态产品的新成本。发布史若只写“原生”,不写调试变难,账本就不完整。
In the GPT-4 era, multimodality was often several models in series: speech transcribed, then a language model, then a synthetic voice. Latency and information loss lived in the seams. After a user finished a sentence, every stage of the pipe still had to run. On 13 May 2024, OpenAI released GPT-4o—the o for omni—a single neural network trained across text, vision, and audio.
The launch ledger must be itemized. Text and image capabilities entered ChatGPT and the API; the low-latency audio experience advanced as a staged preview, not every user and every surface fully open the same day. Writing 13 May as “everything native and generally available” inflates the checklist. Real-time voice demos showed interaction design space opening: interruption, tone, and visual reference could sit on one conversational timeline—still bounded by preview scope and safety policy.
Cost and speed were part of the product claim: multimodal inference was described as cheaper and faster; concrete numbers move with price lists and regional policy and should be checked at the source. For developers, a unified model reduces pipeline glue; for safety and privacy, microphones and cameras entering the default imagination also widen attack and misuse surfaces. Being able to hear and see is not the same as ought-to-be-on by default.
GPT-4o does not end the argument over whether multimodality equals deeper world understanding. It changes the default product form: general assistants begin to be expected to hear, see, and answer at near-human conversational tempo—while engineering and policy still decide which channels truly open. Omni is the target word; staged delivery is the day-one fact.
Treating omni as a finished capability checklist ignores the weight of the preview label. Once microphones and cameras enter the default imagination, products must design refusal and confirmation together: when to listen, when to look, how long records keep. Real-time voice dazzles through tempo; behind tempo sit engineering questions of stable low latency and policy lines against abuse and misuse. Text and image rolled out first, audio later—that order is itself historical fact: multimodality for a general model has never opened every channel on the same day.
“Near-human conversational tempo” is an experience claim, subject to network, device, and load. Lab demos and everyday peak-hour experience can sit far apart. Writing demo tempo as a universal SLA exceeds the launch ledger. Multimodal product history must record peak experience and access scope together.
One network across modalities reduces cascade error and also makes failure modes harder to attribute: hearing, vision, or language? Observability becomes a new cost of multimodal products. Launch history that writes only “native” and not “harder to debug” leaves the ledger incomplete.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- gpt-4o
- 产品
- chatgpt