一盏灯
首页文章视频生词本复习
首页文章视频生词本我的
科技互联网高级15:12ai frontier

With Spatial Intelligence, AI Will Understand the Real World

空间智能:AI 将理解真实世界

AI 教母、ImageNet 缔造者李飞飞从五亿年前寒武纪「视觉的诞生」讲起,勾勒计算机视觉如何一步步从「看见」走向「理解」再走向「行动」。她提出下一个前沿——空间智能:让机器不只是看和说,而是能在三维世界里动手、学习并越做越好,并展望医疗、机器人等场景。适合学习 AI 前沿、科技演讲的高阶英文表达。

播放器加载中…

台词(75 句)

00:04
Let me show you something. To be precise, I'm going to show you nothing. This was the world 540 million years ago.

我给大家看样东西。准确地说,我要给你们看的是「什么都没有」。这就是五亿四千万年前的世界。

00:15
Pure, endless darkness. It wasn't dark due to a lack of light.

纯粹而无边的黑暗。这黑暗并非因为缺少光。

00:22
It was dark because of a lack of sight. Although sunshine did filter 1,000 meters beneath the surface of ocean, a light permeated from hydrothermal vents to seafloor, brimming with life, there was not a single eye to be found in these ancient waters.

它之所以黑暗,是因为缺少「视觉」。尽管阳光确实能穿透到千米深的海面之下,海底热液喷口也透出微光,那里生机盎然,可在这片远古之海里,却找不到哪怕一只眼睛。

00:47
No retinas, no corneas, no lenses. So all this light, all this life went unseen.

没有视网膜,没有角膜,没有晶状体。于是这所有的光、这所有的生命,都无人看见。

00:57
There was a time that the very idea of seeing didn't exist. It [had] simply never been done before.

曾经有那么一段时期,连「看见」这个概念本身都不存在。因为在此之前,从来没有谁看见过。

01:06
Until it was. So for reasons we're only beginning to understand, trilobites, the first organisms that could sense light, emerged.

直到它出现了。出于一些我们才刚刚开始理解的原因,三叶虫——第一种能感知光的生物——登场了。

01:18
They're the first inhabitants of this reality that we take for granted. First to discover that there is something other than oneself.

它们是这个我们习以为常的现实世界最早的居民,也是第一批发现「除了自己之外还存在别的东西」的生物。

01:28
A world of many selves. The ability to see is thought to have ushered in Cambrian explosion, a period in which a huge variety of animal species entered fossil records.

一个由无数个体构成的世界。人们认为,正是「看见」的能力开启了寒武纪大爆发——在那段时期,大量各式各样的动物物种进入了化石记录。

01:43
What began as a passive experience, the simple act of letting light in, soon became far more active.

最初,这只是一种被动的体验——不过是让光进来这么简单的动作,但很快它就变得主动得多。

01:53
The nervous system began to evolve. Sight turning to insight.

神经系统开始进化。视觉逐渐转化为洞察。

02:00
Seeing became understanding. Understanding led to actions. And all these gave rise to intelligence.

看见变成了理解,理解催生了行动,而这一切,又孕育出了智能。

02:10
Today, we're no longer satisfied with just nature's gift of visual intelligence.

今天,我们已不再满足于大自然赐予的这份视觉智能。

02:17
Curiosity urges us to create machines to see just as intelligently as we can, if not better.

好奇心驱使我们去创造能像人一样聪明地「看」——甚至看得比人更好——的机器。

02:25
Nine years ago, on this stage, I delivered an early progress report on computer vision, a subfield of artificial intelligence.

九年前,就在这个讲台上,我做过一次关于计算机视觉的早期进展报告——它是人工智能的一个分支领域。

02:35
Three powerful forces converged for the first time. Aa family of algorithms called neural networks.

当时,三股强大的力量首次汇聚到了一起。第一,一类叫做神经网络的算法。

02:43
Fast, specialized hardware called graphic processing units, or GPUs.

第二,一种快速的专用硬件,叫图形处理器,也就是 GPU。

02:49
And big data. Like the 15 million images that my lab spent years curating called ImageNet.

第三,大数据。比如我的实验室花了数年时间整理的一千五百万张图像,我们把它叫做 ImageNet。

02:57
Together, they ushered in the age of modern AI. We've come a long way.

它们合力开启了现代 AI 的时代。一路走来,我们已经走了很远。

03:04
Back then, just putting labels on images was a big breakthrough. But the speed and accuracy of these algorithms just improved rapidly.

在当年,光是能给图像打上标签就是重大突破。而如今,这些算法的速度和准确率都在飞速提升。

03:14
The annual ImageNet challenge, led by my lab, gauged the performance of this progress.

由我的实验室主办的年度 ImageNet 挑战赛,衡量着这一进展的表现。

03:21
And on this plot, you're seeing the annual improvement and milestone models. We went a step further and created algorithms that can segment objects or predict the dynamic relationships among them in these works done by my students and collaborators.

在这张图上,你能看到逐年的进步和一个个里程碑式的模型。我们又向前迈了一步,做出了能够分割物体、或预测物体之间动态关系的算法——这些成果都出自我的学生和合作者之手。

03:41
And there's more. Recall last time I showed you the first computer-vision algorithm that can describe a photo in human natural language.

还不止如此。还记得上次我给你们展示过第一个能用人类自然语言描述照片的计算机视觉算法吗?

03:52
That was work done with my brilliant former student, Andrej Karpathy. At that time, I pushed my luck and said, "Andrej, can we make computers to do the reverse?"

那是我和我才华横溢的学生安德烈·卡帕西一起做的。当时我贪心地问了句:「安德烈,我们能不能让计算机反过来做——由文字生成图像?」

04:02
And Andrej said, "Ha ha, that's impossible." Well, as you can see from this post, recently the impossible has become possible.

安德烈说:「哈哈,那不可能。」可你从这个帖子里也看得出来,最近,这个「不可能」已经变成了可能。

04:12
That's thanks to a family of diffusion models that powers today's generative AI algorithm, which can take human-prompted sentences and turn them into photos and videos of something that's entirely new.

这要归功于一类扩散模型,它驱动着今天的生成式 AI 算法——能把人类输入的文字提示,变成全新的照片和视频。

04:28
Many of you have seen the recent impressive results of Sora by OpenAI. But even without the enormous number of GPUs, my student and our collaborators have developed a generative video model called Walt months before Sora.

你们很多人都见过 OpenAI 的 Sora 最近那些令人惊艳的效果。但即使没有海量的 GPU,我的学生和合作者早在 Sora 之前几个月,就开发出了一个叫 Walt 的生成式视频模型。

04:47
And you're seeing some of these results. There is room for improvement. I mean, look at that cat's eye and the way it goes under the wave without ever getting wet.

你们看到的就是它的一些效果。当然还有改进空间——你瞧那只猫的眼睛,还有它钻进浪花底下却一点也没湿的样子。

04:59
What a cat-astrophe. (Laughter) And if past is prologue, we will learn from these mistakes and create a future we imagine.

真是一场「猫」难。(笑声)如果说往事是序章,那我们会从这些错误中学习,创造出我们所设想的未来。

05:11
And in this future, we want AI to do everything it can for us, or to help us.

在这样的未来里,我们希望 AI 能替我们、或帮我们,做尽一切它能做的事。

05:19
For years I have been saying that taking a picture is not the same as seeing and understanding.

多年来我一直在说:拍下一张照片,和真正看见并理解,是两回事。

05:26
Today, I would like to add to that. Simply seeing is not enough.

今天,我想再补充一句:仅仅看见,还不够。

05:33
Seeing is for doing and learning. When we act upon this world in 3D space and time, we learn, and we learn to see and do better.

看见是为了行动和学习。当我们在三维的空间与时间中对这个世界施加行动时,我们就在学习,并学会把「看」和「做」都做得更好。

05:46
Nature has created this virtuous cycle of seeing and doing powered by “spatial intelligence.” To illustrate to you what your spatial intelligence is doing constantly, look at this picture.

大自然创造了这样一个「看」与「做」的良性循环,而驱动它的正是「空间智能」。为了让你们体会自己的空间智能时刻都在做些什么,请看这张图。

05:59
Raise your hand if you feel like you want to do something. (Laughter) In the last split of a second, your brain looked at the geometry of this glass, its place in 3D space, its relationship with the table, the cat and everything else.

如果你有种「想动手做点什么」的冲动,请举手。(笑声)就在刚才那一瞬间,你的大脑已经看清了这只杯子的几何形状、它在三维空间中的位置,以及它和桌子、猫,还有其他一切的关系。

06:16
And you can predict what's going to happen next. The urge to act is innate to all beings with spatial intelligence, which links perception with action.

而你能预判接下来会发生什么。这种想要行动的冲动,是一切拥有空间智能的生命与生俱来的,它把感知和行动连在了一起。

06:30
And if we want to advance AI beyond its current capabilities, we want more than AI that can see and talk.

如果我们想让 AI 超越现有的能力,我们要的就不只是一个能看、能说的 AI。

06:39
We want AI that can do. Indeed, we're making exciting progress.

我们要的是一个能「做」的 AI。而事实上,我们正取得令人振奋的进展。

06:46
The recent milestones in spatial intelligence is teaching computers to see, learn, do and learn to see and do better.

空间智能近来的一系列里程碑,就是在教会计算机去看、去学、去做,并学会把看和做都做得更好。

06:57
This is not easy. It took nature millions of years to evolve spatial intelligence, which depends on the eye taking light, project 2D images on the retina and the brain to translate these data into 3D information.

这并不容易。大自然花了数百万年才进化出空间智能:它依靠眼睛接收光线,把二维图像投射到视网膜上,再由大脑把这些数据转译成三维信息。

07:14
Only recently, a group of researchers from Google are able to develop an algorithm to take a bunch of photos and translate that into 3D space, like the examples we're showing here.

直到最近,谷歌的一组研究者才开发出一种算法,能把一堆照片转化成三维空间——就像我们在这里展示的这些例子。

07:29
My student and our collaborators have taken a step further and created an algorithm that takes one input image and turn that into 3D shape.

我的学生和合作者又进了一步,做出了一种算法:只需输入一张图像,就能把它变成三维形状。

07:40
Here are more examples. Recall, we talked about computer programs that can take a human sentence and turn it into videos.

这里还有更多例子。还记得吗,我们刚才提到过能把一句人类语言变成视频的计算机程序。

07:51
A group of researchers in University of Michigan have figured out a way to translate that line of sentence into 3D room layout, like shown here.

密歇根大学的一组研究者想出了一种办法,能把那一句话转译成三维的房间布局,就像这里展示的这样。

08:03
And my colleagues at Stanford and their students have developed an algorithm that takes one image and generates infinitely plausible spaces for viewers to explore.

而我在斯坦福的同事和他们的学生,开发出了一种算法:只需一张图像,就能生成无穷无尽、看起来都合情合理的空间,供观者去探索。

08:17
These are prototypes of the first budding signs of a future possibility.

这些都是原型,是一种未来可能性初露端倪的最初征兆。

08:23
One in which the human race can take our entire world and translate it into digital forms and model the richness and nuances.

在那样的未来里,人类可以把我们整个世界转化为数字形态,并把它的丰富与微妙一一建模呈现。

08:35
What nature did to us implicitly in our individual minds, spatial intelligence technology can hope to do for our collective consciousness.

大自然曾在我们每个人的头脑里悄然完成的事,空间智能技术有望为我们的集体意识去完成。

08:47
As the progress of spatial intelligence accelerates, a new era in this virtuous cycle is taking place in front of our eyes.

随着空间智能不断加速发展,这个良性循环的新纪元,正在我们眼前徐徐展开。

08:56
This back and forth is catalyzing robotic learning, a key component for any embodied intelligence system that needs to understand and interact with the 3D world.

这种看与做的往复,正在催化机器人学习——而对任何需要理解并与三维世界互动的具身智能系统来说,机器人学习都是关键一环。

09:12
A decade ago, ImageNet from my lab enabled a database of millions of high-quality photos to help train computers to see.

十年前,我实验室的 ImageNet 提供了一个包含数百万张高质量照片的数据库,帮助训练计算机去「看」。

09:23
Today, we're doing the same with behaviors and actions to train computers and robots how to act in the 3D world.

如今,我们正用行为和动作做同样的事,来训练计算机和机器人如何在三维世界里行动。

09:34
But instead of collecting static images, we develop simulation environments powered by 3D spatial models so that the computers can have infinite varieties of possibilities to learn to act.

但这次我们不是去收集静态图像,而是构建由三维空间模型驱动的仿真环境,让计算机拥有无穷无尽的可能性,去学习如何行动。

09:50
And you're just seeing a small number of examples to teach our robots in a project led by my lab called Behavior.

你们看到的,只是我实验室主导的一个叫 Behavior 的项目里,用来教机器人的一小部分例子。

10:00
We’re also making exciting progress in robotic language intelligence. Using large language model-based input, my students and our collaborators are among the first teams that can show a robotic arm performing a variety of tasks based on verbal instructions, like opening this drawer or unplugging a charged phone.

我们在机器人语言智能方面也取得了令人振奋的进展。借助以大语言模型为基础的输入,我的学生和合作者是最早一批能让机械臂根据口头指令完成各种任务的团队之一——比如打开这个抽屉,或是拔掉一部充好电的手机的充电线。

10:26
Or making sandwiches, using bread, lettuce, tomatoes and even putting a napkin for the user.

又或者做三明治——用上面包、生菜、番茄,甚至还给使用者备上一张餐巾纸。

10:34
Typically I would like a little more for my sandwich, but this is a good start. (Laughter) In that primordial ocean, in our ancient times, the ability to see and perceive one's environment kicked off the Cambrian explosion of interactions with other life forms.

一般来说,我的三明治我还想多加点料,不过这已经是个不错的开端了。(笑声)在那片原始的海洋里,在我们的远古时代,「看见并感知周遭环境」的能力,引爆了与其他生命形式互动的寒武纪大爆发。

10:55
Today, that light is reaching the digital minds. Spatial intelligence is allowing machines to interact not only with one another, but with humans, and with 3D worlds, real or virtual.

今天,那束光正照进数字的心智。空间智能让机器不仅能彼此互动,还能与人类互动,与三维世界——无论真实还是虚拟——互动。

11:12
And as that future is taking shape, it will have a profound impact to many lives.

随着那样的未来逐渐成形,它将深刻地影响许许多多人的生活。

11:18
Let's take health care as an example. For the past decade, my lab has been taking some of the first steps in applying AI to tackle challenges that impact patient outcome and medical staff burnout.

以医疗为例。过去十年里,我的实验室率先迈出了一些步子,把 AI 用于应对那些影响患者治疗效果、以及医护人员职业倦怠的难题。

11:34
Together with our collaborators from Stanford School of Medicine and partnering hospitals, we're piloting smart sensors that can detect clinicians going into patient rooms without properly washing their hands.

我们和斯坦福医学院的合作者以及合作医院一起,正在试点一种智能传感器:它能检测出医护人员在没有正确洗手的情况下走进病房。

11:49
Or keep track of surgical instruments. Or alert care teams when a patient is at physical risk, such as falling.

它也能追踪手术器械的去向,或在患者面临身体风险(比如跌倒)时,向护理团队发出警报。

11:59
We consider these techniques a form of ambient intelligence, like extra pairs of eyes that do make a difference.

我们把这类技术视为一种「环境智能」,就像一双双额外的眼睛,而它们确实能带来改变。

12:08
But I would like more interactive help for our patients, clinicians and caretakers, who desperately also need an extra pair of hands.

但我更希望为我们的患者、医生和护理人员提供更具互动性的帮助——他们同样迫切地需要多一双手。

12:19
Imagine an autonomous robot transporting medical supplies while caretakers focus on our patients or augmented reality, guiding surgeons to do safer, faster and less invasive operations.

想象一下:一台自主机器人在运送医疗物资,好让护理人员专注于照顾患者;又或者借助增强现实,引导外科医生做出更安全、更快、创伤更小的手术。

12:35
Or imagine patients with severe paralysis controlling robots with their thoughts.

再想象一下:严重瘫痪的患者,仅凭意念就能操控机器人。

12:42
That's right, brainwaves, to perform everyday tasks that you and I take for granted.

没错,就是用脑电波,去完成那些你我习以为常的日常事务。

12:49
You're seeing a glimpse of that future in this pilot study from my lab recently. In this video, the robotic arm is cooking a Japanese sukiyaki meal controlled only by the brain electrical signal, non-invasively collected through an EEG cap.

在我实验室最近的这项试点研究里,你正瞥见那个未来的一角。在这段视频中,机械臂正在做一顿日式寿喜烧,而操控它的,仅仅是通过脑电帽无创采集到的大脑电信号。

13:10
(Applause) Thank you. The emergence of vision half a billion years ago turned a world of darkness upside down.

(掌声)谢谢。五亿年前视觉的出现,把一个黑暗的世界彻底颠覆。

13:23
It set off the most profound evolutionary process: the development of intelligence in the animal world.

它引发了最深远的一场进化进程:动物世界中智能的诞生与发展。

13:31
AI's breathtaking progress in the last decade is just as astounding. But I believe the full potential of this digital Cambrian explosion won't be fully realized until we power our computers and robots with spatial intelligence, just like what nature did to all of us.

过去十年 AI 令人屏息的进展同样惊人。但我相信,这场数字版寒武纪大爆发的全部潜力,要等到我们用空间智能去驱动计算机和机器人——就像大自然当年赋予我们所有人的那样——才能真正完全释放。

13:55
It’s an exciting time to teach our digital companion to learn to reason and to interact with this beautiful 3D space we call home, and also create many more new worlds that we can all explore.

这是一个激动人心的时刻:我们教会数字伙伴学会推理,学会与这个我们称之为家园的美丽三维空间互动,还能创造出更多我们都可以去探索的新世界。

14:11
To realize this future won't be easy. It requires all of us to take thoughtful steps and develop technologies that always put humans in the center.

要实现这样的未来并不容易。它需要我们所有人审慎行事,去开发那些始终把人放在中心的技术。

14:23
But if we do this right, the computers and robots powered by spatial intelligence will not only be useful tools but also trusted partners to enhance and augment our productivity and humanity while respecting our individual dignity and lifting our collective prosperity.

但只要我们做对了,由空间智能驱动的计算机和机器人,就不仅会是有用的工具,更会是可信赖的伙伴——它们在提升与增强我们生产力和人性的同时,尊重每个人的尊严,并抬升我们共同的繁荣。

14:45
What excites me the most in the future is a future in which that AI grows more perceptive, insightful and spatially aware, and they join us on our quest to always pursue a better way to make a better world.

未来最让我兴奋的,是这样一个未来:AI 变得更善于感知、更富洞察、更懂空间,并加入我们的求索,与我们一同,永远去追寻一种更好的方式,去缔造一个更好的世界。

15:05
Thank you. (Applause)

谢谢大家。(掌声)