Meta 发布 SAM「分割一切」
像 ChatGPT 之于文本,SAM 想成为图像分割的基础模型
Meta AI 发布 Segment Anything Model(SAM),一个可在零样本下分割任意图像中任意物体的模型,配套 1100 万张图像、10 亿掩码的数据集。它被视为视觉版「基础模型」。
在 ChatGPT 之前,「基础模型」还只是文本领域的概念;2023 年初,人们开始问:视觉领域能不能也出现一个「什么都能干」的模型?图像分割——把图片里的物体一个个圈出来——恰恰是视觉里最基础、也最繁琐的任务之一。Meta 把赌注押在了这里。
2023 年 4 月,Meta AI 发布了 Segment Anything Model,简称 SAM。它的用法和 ChatGPT 有几分神似:给模型一个提示——点一下、画个框,甚至输入一段文字——它就能输出精确的分割结果,而且是零样本:不需要针对某个具体任务重新训练。背后的数据同样惊人:1100 万张图像,超过 10 亿个掩码。
SAM 让人惊叹的地方在于「通用」。过去要做图像分割,几乎每个任务、每类物体都要训练专门的模型;SAM 却用一个模型应对所有提示。它把「分割」从需要专家、需要定制的工作,变成了像填空一样简单的操作。模型和数据全部开源,很快被集成进无数工具。
当然,SAM 也有边界。它擅长「把东西圈出来」,但不理解圈出来的东西是什么;它的能力依赖提示,没有提示就无法工作。它并不是视觉智能的全部,只是视觉智能里「分割」这一块拼图。但正是这块拼图,让整个视觉生态的效率上了一个台阶。
回看 SAM,它最珍贵的贡献是验证了一个范式:基础模型加提示,可以复制到文本之外的领域。它让研究者相信,视觉也有自己的「ChatGPT 时刻」。当后来的多模态模型把视觉理解推向更高处时,SAM 已经在图像处理的最底层,铺好了那张所有人都能用的网。
Before ChatGPT, "foundation model" was still a text-world concept; in early 2023 people began asking whether vision could have a model that "does everything." Image segmentation—circling every object in an image—is among the most basic and most tedious tasks in vision. Meta put its bet there.
In April 2023, Meta AI released the Segment Anything Model, SAM. Its usage is oddly reminiscent of ChatGPT: give the model a prompt—a click, a box, even a line of text—and it outputs a precise segmentation, zero-shot, without retraining for a specific task. The data behind it is equally staggering: 11 million images and over 1 billion masks.
What made SAM remarkable was its generality. In the past, segmentation nearly always required a dedicated model per task and per object class; SAM handles all prompts with one model. It turned segmentation from expert, custom work into something as simple as filling in a blank. The model and data were fully open-sourced and quickly integrated into countless tools.
Of course, SAM has boundaries. It is excellent at "circling things" but does not understand what it circles; it depends on prompts and cannot work without one. It is not the whole of visual intelligence, just the "segmentation" piece of the puzzle. But that piece lifted the efficiency of the entire vision ecosystem.
Looking back, SAM's most precious contribution was validating a paradigm: foundation model plus prompting can be replicated beyond text. It convinced researchers that vision had its own "ChatGPT moment." When later multimodal models pushed visual understanding higher, SAM had already laid, at the very bottom of image processing, a net that everyone could use.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- —
- 产品
- —