生数发布 Vidu 1.0

中国创业公司把文生视频做成可公开使用的产品

生数科技发布 Vidu 1.0 文生视频模型与应用,强调动态一致性与角色稳定,成为 Sora 演示之后最早一批可实际试用的中国视频生成产品之一。

时间2024 年 4 月 27 日 级别B · 领域级 组织生数科技 / Shengshu 状态已核验 · 1 个来源
编辑插图:胶片帧中一个稳定的人物剪影穿过连续镜头
AI Chronicle 原创插图:连续帧里不碎裂的主体,对应 Vidu 强调的动态一致性。 AI Chronicle

2024 年 2 月,Sora 的演示让全世界突然知道视频模型可以追求什么;问题是,大多数人只能观看。两个月后的 4 月 27 日,生数科技与清华团队在中关村论坛发布 Vidu。它同样用精选片段展示长镜头、运动和角色一致性,但更重要的悬念已经从“能不能生成”变成“什么时候能让创作者自己试”。

Vidu 1.0 采用团队提出的 U-ViT 架构,把扩散模型与 Transformer 结合。发布时给出的目标规格是单次生成最长 16 秒、最高 1080p,并强调高动态场景与时序一致性。这些数字描述了模型的上限和厂商选择,不代表任意提示都能稳定得到同样质量。镜头越长,身份漂移、物体变形和动作失去因果的机会也越多。

4 月的发布首先是一场模型亮相,真正面向用户的产品开放随后到来。这个时间差值得保留:演示由团队挑选,生成队列却会收进所有笨拙、含糊甚至互相冲突的提示。普通用户开始经历注册、输入、等待、重试和下载,Vidu 才从一组视频样片变成一个可被检验的工具。失败不再被发布会剪掉,而成为产品每天必须处理的负载。

对中国创作者来说,本土入口还减少了另一层摩擦。中文提示、熟悉的题材和本地服务流程,使测试不必先跨过语言与访问门槛。Vidu 并没有因此自动胜过 Sora、Runway 或后来大规模开放的可灵;它做的是把比较拉到同一张工作台上。用户可以用自己的角色、动作和镜头要求,观察哪一个模型在第几秒开始失控。

Vidu 1.0 后续还会被更高画质、图生视频和参考主体等版本覆盖。第一代留下的价值却很明确:在视频生成被少数华丽演示定义的阶段,一家创业公司把模型推向了可反复尝试的产品。门一旦打开,竞争就不再只由谁的发布片最好看决定,而由谁能承受真实用户不断提交的坏提示与坏结果决定。

In February 2024, Sora’s demonstrations abruptly expanded what the public expected from a video model. Most people, however, could only watch. On April 27, Shengshu Technology and a Tsinghua team unveiled Vidu at the Zhongguancun Forum. Its selected clips also displayed longer shots, movement, and subject consistency, but the more important question had changed from “can this be generated?” to “when can a creator try it with an unselected prompt?”

Vidu 1.0 used the team’s U-ViT architecture, combining diffusion and Transformer components. The launch described a ceiling of up to 16 seconds in a single generation and resolution up to 1080p, with an emphasis on dynamic scenes and temporal coherence. Those figures expressed a target and a technical choice; they did not promise the same quality for arbitrary requests. The longer a shot continued, the more opportunities it created for identity drift, object deformation, or an action to lose its cause.

The April event was first a model unveiling. User-facing access followed rather than being completed on the same day. Preserving that interval matters. A demonstration is selected by its maker; a generation queue receives clumsy, ambiguous, and internally contradictory prompts from everyone. Only when ordinary users could register, type, wait, retry, and download did Vidu become a tool that could be tested rather than a collection of clips that could be admired. Failure was no longer removed in the launch edit. It became daily product load.

For creators in China, a domestic entrance removed another layer of friction. Chinese prompting, familiar subject matter, and local service flows meant experimentation did not begin with language or access barriers. That did not automatically make Vidu superior to Sora, Runway, or the later large-scale release of Kling. It placed them on a more comparable workbench. A user could bring a specific character, action, and camera instruction, then observe which system lost control and at which second.

Later Vidu generations would add better image quality, image-to-video modes, and stronger subject references, pushing 1.0 into the legacy column. Its first release retained a clear role. At a moment when video generation was defined by a small number of polished demonstrations, a startup pushed the model toward a product people could repeatedly try. Once that door opened, competition could no longer be decided only by whose launch reel looked best. It also depended on who could withstand the bad prompts and bad outputs submitted by real users every day.

展开完整事件档案人物、主题、模型与产品
人物
模型
vidu
产品
来源

原始资料

  1. 01ViduShengshu · official

试试搜索