Amazon Mechanical Turk 上线
众包平台把「人工」藏进「人工智能」流水线
亚马逊上线 Mechanical Turk(MTurk),让企业把人类难以自动化的任务拆成微任务分发给在线工人。它成了 AI 数据标注的基础设施,也把「众包人工」嵌入机器学习流水线,改变了数据生产的方式。
2005 年,亚马逊上线了一个叫 Mechanical Turk 的服务。名字来自 18 世纪那台号称会下棋、实则在柜子里藏人的「土耳其机器人」——这个命名本身就是个精准的隐喻。MTurk 把人类难以自动化的任务拆成一个个几美分的微任务,分发给在线工人完成,企业通过 API 按件计价。人工,被悄悄嵌进了机器学习的流水线。
MTurk 最早是亚马逊的「内部工具」。它处理过图像识别、去重、内容审核这类机器搞不定的活,随后对外开放。对 2000 年代中期正为数据发愁的研究者来说,它几乎是天降甘霖:标注一张图几分钱,几天就能攒出过去几个月的人力工作量。ImageNet 这样改变 AI 进程的大型数据集,背后就有众包标注的影子。
它的影响远超「一个外包平台」。MTurk 让「数据标注」从研究团队的杂活变成一种可规模化的产业形态,催生了一大批平台型公司与专业化标注团队。今天 AI 行业动辄谈论训练数据,而这套「把任务切碎、按件分发、人机协作」的模式,正是从 MTurk 开始的。
但硬币的另一面同样真实。众包工人时薪低、无保障、缺少权益通道,标注工作的枯燥与高强度长期被行业忽视。当「AI 背后站着海量匿名工人」的现实被揭示,围绕数据标注的劳动伦理成了 AI 行业绕不开的议题。
回看 MTurk,它提醒我们一个容易被遗忘的事实:人工智能里始终有「人工」。数据是 AI 的燃料,而众包平台是这套燃料体系的早期基础设施。它让数据生产规模化,也把公平与尊严的问题留给了整个行业——这些问题,直到今天仍在被追问。
In 2005 Amazon launched a service called Mechanical Turk. The name came from the 18th-century chess-playing "Turk"—a machine with a person hidden inside—and the naming was itself a precise metaphor. MTurk split tasks humans couldn't automate into microtasks worth a few cents, distributed them to online workers, and let companies pay per unit through an API. Human labor was quietly embedded into the machine-learning pipeline.
MTurk began as an internal Amazon tool. It handled image recognition, de-duplication, and content moderation—the work machines couldn't manage—then opened to the public. For researchers desperate for data in the mid-2000s it was a gift: label an image for a few cents, and in days you had what used to take months of in-house effort. Large datasets that changed the course of AI, like ImageNet, carried crowdsourced labeling in their veins.
Its impact went far beyond "an outsourcing platform". MTurk made data labeling a scalable industry and spawned platform companies and specialized annotation teams. Today the AI world talks endlessly about training data, but this pattern—chop tasks into pieces, distribute per unit, human and machine cooperate—began with MTurk.
The other side of the coin is equally real. Crowd workers earn little, lack protections, and have few channels for their rights; the tedium and intensity of labeling work has long been ignored by the industry. When the reality that "AI stands on countless anonymous workers" was revealed, the labor ethics of data annotation became an unavoidable issue.
Looking back, MTurk reminds us of an easy-to-forget fact: there has always been "human" in "artificial intelligence". Data is AI's fuel, and the crowdsourcing platform was the early infrastructure of that fuel system. It scaled data production—and left questions of fairness and dignity for the whole industry, questions still being asked today.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- —
- 产品
- —