Google 发布 Veo 视频生成模型
搜索巨头把文生视频写进 I/O 主舞台
Google 在 I/O 2024 公布 DeepMind 的 Veo 文生视频模型,宣称可生成更高分辨率、更长片段的视频,与 Sora 演示后的全球赛道正式对位。
2024 年 5 月 14 日,Google I/O 的大屏幕上出现了 Veo 生成的长镜头。Google DeepMind 宣称模型能够生成 1080p、超过一分钟的视频,理解“延时摄影”“航拍”等电影语言,并在较长提示中维持人物、物体与场景的一致性。三个月前,Sora 的演示已经抬高公众胃口;Veo 让 Google 第一次用一个清楚的品牌名正面进入同一赛道。
舞台上的可见度与产品里的可用性并不相同。Veo 当天只通过 VideoFX 向部分创作者提供私密预览,其他人需要加入等待名单。Google 还邀请 Donald Glover 与其 Gilga 工作室等创作者试用,希望从实际制作流程中获得反馈。对大多数观众而言,I/O 展示的是路线图和经过挑选的样片,不是一项回家就能打开的普遍服务。
超过一分钟这个目标很具体,因为视频错误会随时间累积。单帧好看不难,人物走动时衣服、肢体和背景关系保持一致更难;镜头继续推进,先前出现的物体还要遵守同一个空间。Google 把“理解真实世界物理”和“准确遵循较长提示”写进发布材料,但这些仍是厂商对模型能力的描述。演示能够证明模型生成过这样的片段,不能证明每次提示都能稳定复现。
Veo 也把创作者控制放在单纯画质之外。电影术语可以进入提示,已有画面可以被进一步编辑或延展,模型被设想为镜头草稿工具,而不只是随机短视频生成器。与创作者合作的意义正在这里:电影制作需要角色连续、镜头衔接和可修改素材,漂亮的八秒样片只是工作流的一个部件。
Google 同时宣布,VideoFX 生成的 Veo 视频将嵌入 SynthID 水印,并配合安全测试、过滤和对记忆内容的检查。视频越接近真实摄影,来源标识越不能等到滥用发生后再补。水印不能独自解决版权、冒充与检测问题,却说明发布方已经把生成能力与内容溯源放在同一个产品入口里。
Veo 首发那天,Google 真正完成的是从研究项目到可命名平台旗舰的转换。1080p 和一分钟以上定义了野心,VideoFX 等待名单定义了现实边界,SynthID 定义了它试图承担的治理方式。大屏幕上的镜头向所有人播放;谁能按下“生成”,仍由一扇尚未完全打开的门决定。
On 14 May 2024, long shots generated by Veo filled the screen at Google I/O. Google DeepMind said the model could produce 1080p videos longer than a minute, understand cinematic language such as timelapse and aerial shot, and preserve people, objects, and visual style across longer prompts. Sora’s demonstration three months earlier had raised public expectations. Veo gave Google a clear flagship name with which to enter the same contest.
Visibility on the keynote stage was not the same as availability in a product. Veo entered private preview through VideoFX for selected creators; everyone else was directed to a waitlist. Google invited filmmakers including Donald Glover and his Gilga studio to experiment with the model and provide feedback from actual production work. It also said some capabilities would reach YouTube Shorts in the future, another promise rather than launch-day access. For most viewers, I/O offered a roadmap and selected samples, not a service they could open when the presentation ended.
The promise of more than a minute mattered because video errors accumulate with time. A convincing frame is one problem. Keeping clothing, anatomy, objects, and background relationships stable while a character moves is another. As the camera continues, things introduced earlier must remain in the same world. Google described Veo as understanding real-world physics and following longer prompts accurately. Those were publisher claims. A demonstration established that the model had generated such clips, not that every prompt would reproduce them reliably.
Veo also emphasized creator control beyond image quality. Cinematic terms could be written into prompts, and generated material was presented as something that could be extended or further edited. The intended role was closer to a shot-development tool than a machine for isolated random clips. Collaboration with filmmakers mattered for precisely that reason: production requires character continuity, transitions, and revisable material. A beautiful eight-second sample is only one component of a workflow.
Google announced that Veo videos made in VideoFX would carry SynthID watermarks, alongside safety tests, filters, guardrails, and checks for memorized content. As generated footage approached ordinary photography, provenance could not be treated as an afterthought. Watermarking could not resolve copyright, impersonation, and detection by itself, but it placed generative capability and source marking inside the same product doorway.
The first Veo release converted a research direction into a named platform flagship. The 1080p, minute-plus claim defined the ambition. The VideoFX waitlist defined the practical boundary. SynthID defined part of the governance Google wanted attached to it. The keynote video played for everyone; the ability to press Generate still depended on a door that had opened only a little.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- veo
- 产品
- —