字节发布 SeedRealtime 音视频全双工模型
豆包 App 全量上线「打电话」视频通话,端到端架构融合音视频与文本
字节跳动发布 SeedRealtime:原生音视频全双工大模型,统一端到端架构融合音频、视频与文本,豆包 App 全量上线视频通话;官方称相比级联模型,节奏问题减少约 50%。
2026 年 8 月 5 日,字节跳动把 SeedRealtime 推到台前时,演示里最动人的画面不是参数曲线,而是一个普通人拿起手机,对着镜头说「帮我看看这台咖啡机为什么不出水」。电话那头,模型看着画面,纠正了操作步骤。语音助手第一次不只是听见你,还看见你。
过去十年,实时语音助手走的是同一条装配线:语音识别把声音变成文字,大模型处理文字,语音合成再把回答变回声音。这条级联路线能工作,却有两个天生的毛病——延迟让对话像对讲机,节奏生硬;更重要的是,它看不见。你说「帮我看看这个」,它只能问「什么?」。SeedRealtime 用统一端到端架构把音频、视频、文本揉进同一个模型,听、看、说不再分家。
官方称相比级联模型,节奏问题减少约 50%。这是厂商口径,需要带着自述边界读;但演示场景本身说明了很多:多人聚餐时认出每个人、川菜馆里看着菜单点餐、河北博物院主动提醒展品、大兴机场的嘈杂环境里保持对话。这些不是实验室里的合成场景,而是真实生活里「需要看见才能回答」的时刻。模型第一次要为画面负责。
豆包 App 全量上线视频通话,说明这套能力不是研究演示,而是亿级用户产品里的默认功能。对用户,交互从「对讲机」变成「面对面」:可以打断、可以指着东西问、可以同时和几个人说话。对行业,实时多模态对话的架构路线从级联转向端到端,语音助手开始比拼视觉理解与节奏感,而不是谁的 TTS 更自然。
SeedRealtime 没有宣称延迟已经消失,也没有说模型永远不会看错。它做的是把「看见」写进对话的默认路径。当电话那头第一次能看见你面前的世界,语音助手的竞争就从「听懂」升级为「看懂」——而看懂,是 Agent 走向真实世界的第一步。
On August 5, 2026, when ByteDance put SeedRealtime on stage, the most moving image in the demo was not a parameter curve but an ordinary person holding up a phone and saying, "Help me see why this coffee machine is not dispensing water." On the other end, the model looked at the frame and corrected the steps. A voice assistant was hearing you and seeing you for the first time.
For a decade, real-time voice assistants ran on the same assembly line: speech recognition turned sound into text, a large model processed the text, and speech synthesis turned the answer back into sound. The cascade worked, but it had two congenital flaws—latency made conversation feel like a walkie-talkie, and, more importantly, it could not see. When you said "help me look at this," it could only ask "what?" SeedRealtime folds audio, video, and text into one end-to-end model; listening, seeing, and speaking are no longer separate jobs.
The company reported roughly 50% fewer pacing problems versus cascaded models. That is a vendor claim and should be read with that boundary; but the demo scenes themselves said a lot: recognizing each person at a group dinner, ordering from a menu at a Sichuan restaurant, proactively reminding visitors at a museum, holding a conversation in the noise of Daxing Airport. These are not synthetic lab scenes but real moments where "seeing" is required to answer. The model became accountable to the picture.
Doubao shipping video calls app-wide means this is not a research demo but a default feature inside a mass-market product. For users, interaction shifted from walkie-talkie to face-to-face: interruptible, point-at-things, multi-party. For the industry, real-time multimodal dialogue moved from cascaded to end-to-end architecture; assistants now compete on visual understanding and pacing rather than whose TTS sounds more natural.
SeedRealtime does not claim latency has vanished, nor that the model never misreads a scene. What it did was write "seeing" into the default path of conversation. When the voice on the other end can finally see the world in front of you, the assistant race upgrades from understanding speech to understanding scenes—and understanding scenes is the first step of agents entering the real world.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- seed-realtime
- 产品
- doubao