OpenAI 发布 Sora 2
视频与音频生成进入同一旗舰,并长出社交形态
OpenAI 发布 Sora 2 旗舰视频与音频生成模型,同步 iOS 应用与水印策略,把 2024 年底的订阅生成能力升级为更强的多模态产品,并探索信息流式分发。
2025 年 9 月 30 日,OpenAI 发布 Sora 2 时,展示的不只是下一代视频模型。配套的 iOS 应用带有作品信息流,用户可以生成、浏览和重混短片,并把获准的人物形象带进新的场景。Sora 从 ChatGPT 订阅里的创作能力,长成了一个有独立图标和消费习惯的媒体产品。生成按钮旁边,第一次出现了“刷视频”的入口。
模型层的重点是视频与音频共同生成。对白、环境声和动作可以在同一次输出中配合,OpenAI 还强调更好的物理表现、时序一致性和指令遵循。竞争对手已经把运动与镜头控制推成日常卖点,单纯提高画质不足以定义第二代。Sora 2 试图让一段生成视频在声音打开后仍然成立——无声的演示片时代结束了,观众会听到脚步、门响和台词。
独立应用改变了风险的形状。生成工具面对的是一个主动写提示的人,信息流面对的是大量准备滑动和转发的观众。人物进入视频、作品被重混、相似内容被推荐,都会让身份同意、冒充和内容来源变得更难管理。首发采用地区与邀请限制,并在视频上加入可见动态水印;这些措施提供提示和追踪线索,却不能保证画面永远不会被裁切、重录或脱离原始语境。水印因此不是一层贴纸,而是 Sora 2 产品设计的自我说明。
OpenAI 预期这些片段会离开生成页面,进入公共传播。模型团队关心声音和动作是否同步,应用团队还要处理谁能生成谁、谁能看见什么、推荐系统怎样放大某类内容。视频模型一旦拥有社交形态,治理就与推理同时消耗产品资源。这也是后来停服时被反复提起的伏笔:一个带信息流的生成应用,运营成本远不止算力账单。
Android 与更高能力档随后扩展了这条产品线,但 9 月 30 日的关键选择已经完成:Sora 不再只是创作者偶尔打开的工具。它尝试成为人们持续观看、模仿和参与的地方。第二代真正扩大的不是片段长度,而是生成内容抵达观众的路径——从工具到媒体,这一步跨出去,就再也退不回来。
When OpenAI released Sora 2 on September 30, 2025, it presented more than a next-generation video model. The accompanying iOS application included a feed in which users could generate, watch, and remix short clips, as well as place approved personal likenesses into new scenes. Sora had grown from a creation capability attached to a ChatGPT entitlement into a media product with its own icon and its own habit loop.
At the model layer, the central change was joint video and audio generation. Dialogue, ambient sound, and movement could be produced as parts of the same output. OpenAI also stressed better physical behavior, temporal consistency, and instruction following. Competitors had already made motion and camera control everyday selling points; improving image quality alone could not define a second generation. Sora 2 attempted to make a generated clip remain convincing after the viewer turned the sound on.
The standalone app changed the shape of risk. A creation tool served one person who deliberately wrote a prompt. A feed served a larger audience prepared to swipe and share. When likenesses entered clips, works were remixed, and similar material was recommended, consent, impersonation, and provenance became harder to manage. The initial rollout used regional and invitation limits and placed visible moving watermarks on videos. Those measures supplied disclosure and a tracing clue; they could not guarantee that an image would never be cropped, rerecorded, or separated from its original context.
The watermark was therefore more than a sticker. It was an admission embedded in the design of Sora 2: these clips were expected to leave the creation screen and circulate in public. Model teams had to make sound and motion align. Application teams also had to decide who could generate whom, who could see which content, and how recommendations amplified particular outputs. Once a video model acquired a social form, governance consumed product resources alongside inference.
Android access and higher-capability tiers later extended the line, but the important choice had already been made on September 30. Sora was no longer only a tool a creator opened occasionally. It attempted to become a place people continually watched, imitated, and entered. What the second generation expanded most was not clip duration, but the number of paths by which generated media could reach an audience.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- sora-2
- 产品
- sorachatgpt