Google 发布 Gemini 1.5

百万级上下文窗口让长文档处理进入新量级

Google 发布 Gemini 1.5 Pro,首次将上下文窗口推向 100 万 token,可在一次推理中处理数小时视频、整本书籍与超长代码库。它把「长上下文」从演示变成可用的产品能力。

时间2024 年 2 月 15 日 级别A · 行业级 组织Google 状态已核验 · 1 个来源
Gemini 1.5 Pro 处理百万 token 长文档的插画
Gemini 1.5 Pro 把上下文窗口推进到百万级 token。 AI Chronicle

2024 年 2 月,谷歌发布了 Gemini 1.5 Pro。在众多升级里,最让人记住的是一个数字:100 万 token 的上下文窗口。这意味着在一次推理里,模型可以「吃下」数小时的视频、一整本书、甚至规模不小的代码仓库。在 GPT-4 时代,几万 token 已经算长;百万级别,是另一个量级。

上下文窗口决定了 AI 的「一次性视野」。窗口小的模型,面对长文档要么截断、要么分段,还得靠检索把相关片段挑出来。Gemini 1.5 直接把这道工序省了:整本书放进去,模型自己读、自己找、自己答。官方演示里,它分析一段 44 分钟无声电影、在几千页文档里精准定位信息,这种「全量输入」的能力第一次成了可用的产品卖点。

为了撑起百万上下文,Gemini 1.5 在架构上做了针对长序列的优化,并采用了稀疏 MoE 结构。但对用户而言,架构是次要的,体验是主要的:过去要拆成几十次提问的活,现在一次对话就能完成。这种变化对文档分析、代码审查、长视频理解等场景是直接的能力升级。

Gemini 1.5 也点燃了一场行业竞赛。它发布之后,上下文窗口成了各家旗舰的核心宣传指标:OpenAI 扩展 GPT-4 的上下文,国产模型更是把「百万字」当作标配。Kimi 以超长文本破圈、各家纷纷跟进,长上下文从「锦上添花」变成「默认要求」。整个行业对「喂给模型多少材料」的想象被一次发布重新定义。

回看 Gemini 1.5,它的价值在于把「长上下文」从一个演示概念变成了真实可用的产品能力。今天旗舰模型的百万级上下文已经见怪不怪,但回看起点,是 2024 年初这台「能读整本书」的模型,把大模型的输入边界推向了一个新量级,也让长上下文竞赛正式开跑。

In February 2024 Google released Gemini 1.5 Pro. Among many upgrades, one number stuck: a one-million-token context window. In a single pass the model could absorb hours of video, a whole book, even a sizable codebase. In the GPT-4 era, tens of thousands of tokens already counted as long; a million was another magnitude.

The context window sets the model's one-shot field of view. A small-window model must truncate long documents or chop them up, relying on retrieval to surface relevant pieces. Gemini 1.5 removed that step: put the whole book in, and the model reads, searches, and answers itself. Official demos showed it analyzing a 44-minute silent film and pinpointing information across thousands of pages—"whole input" became a usable product selling point for the first time.

Supporting a million tokens took long-sequence optimizations and a sparse MoE architecture. For users, though, the architecture was secondary and the experience primary: work that used to require dozens of prompts could now be done in one conversation. For document analysis, code review, and long-video understanding, that was a direct capability jump.

Gemini 1.5 also ignited an industry race. After its release, context length became a headline metric for every flagship: OpenAI extended GPT-4's window, and Chinese models made "a million characters" standard. Kimi broke out on ultra-long text, and vendors followed—long context went from a nice-to-have to a default requirement. One release redrew the industry's imagination of how much material a model can take.

Looking back, Gemini 1.5's value was turning long context from a demo concept into real product capability. Today million-token windows are unremarkable, but at the origin is this model from early 2024 that could "read a whole book"—the moment the input boundary of large models was pushed to a new magnitude and the long-context race officially began.

展开完整事件档案人物、主题、模型与产品
人物
模型
产品
来源

原始资料

  1. 01Our next-generation model, Gemini 1.5Google · official

试试搜索