heading · body

Transcript

Jitendra Malik Vision And Robotics For Embodied Ai

read summary →

---TRANSCRIPT--- Thank you Mark. It’s a great pleasure to be here. I always enjoy visiting ETH Zurich. The energy is amazing and all the demos we saw in robotics were amazing. So, I’m here to tell you about what I think about robotics and a particular perspective on robotics. Can everyone hear me? Great. Excellent. Okay, so let’s start. So, I I will start by talking about a particular perspective on robotics. So, what are the central problems of robotics? So, there’s a definition connecting perception to action and I would argue that the central problems that most people would agree are navigation, locomotion, and manipulation. So, let me show you some results on navigation. These are now practically 2 years old. Uh it’s we had a system called go to go to anything. And the challenge was the following. Uh we have all this technology which has been developed over the years. Uh mapping, finding objects, going exploring, etc., etc. So, so I I talked with the team and I said the challenge is the Airbnb test. So, we’re going to rent an Airbnb. You don’t know anything about it. You open the door, you put down the robot, and then the robot is given a task. And the task is go find the television. Go find a potted plant. The task could be specified linguistically or with an image. Right? Or it could even be a a location specified geometrically. And and that’s what we did. And let’s see what happens here. So, this Spot robot is is is the one which is being used, but you could use any robot. That’s not the main point. The main point is the software running on top of from the input from the camera. And this is what the robot sees. So, it’s right now seeing a sofa and it’s been given a goal. What’s the goal? It’s like this kitchen sink, but in this case specified as an image. So, it’s going to go around, explore until it finds it. But, in this process it’s going to build up a representation, which is going to be a partial map. It’s not going to be a complete map, but it’ll be a semantic map in which it records objects. And uh let’s see what happens. So, it’s moving. See what’s Okay, so the robot is going around. It’s identifying objects, couch, chair, chair, couch, etc. It’s not yet found the goal. It’s seeing various objects. It’s built a certain semantic map. And now it it found the goal. Then we give it a new goal, the red cup on the brown leather chair. It’s going to keep moving around. It has success. Keeps moving around. Etc. etc. And uh because it’s built a map and it has a a memory, if you ask for the same goal again, then it will be efficient. It won’t have to do the search. It doesn’t need to build a full slam. It builds what is needed. This is how we humans operate. And it’s been trained in simulation. So, therefore it knows what are more productive paths. Where are kitchens more likely to be because it was trained in an simulation with uh lots of environments of homes and so on. So, this is what we can do in uh navigation. Uh I think it these are very robust systems. The failure rates are often due due to failures of computer vision, which basically, if you don’t like the results, you just wait for the next generation of your foundation model and it’ll be better. And we evaluated it on 200 object instances and so on. Uh there are some scientific findings. End-to-end policies did not do very well. A modular approach where for each module you have kind of the best vision system, for example, is actually wise. What you do need learning for is efficient exploration policies. Okay, now let me talk about locomotion. We have in my group, we have worked on quadruped locomotion, we have worked on humanoids, bipeds. I’m going to show you results. We started with quadrupeds, moved to humanoids. I’ll show you results. These are basically also about 2 years old with these series of papers from 2024. And the test is that of uh robustness. I mean, the robot was trained in simulation, zero shot in the real world. And these are long hikes around Berkeley. And and then uh then we decided, “Okay, so let’s test how well we can do.” So, these are our hikes in San Francisco. San Francisco is known to be a city built on hills. So, this is a pretty serious slope, 29%. And uh it’s all zero shot stuff, right? And then I So, this is work led by Ilya Rasazovich, who was my grad student at the time. And I said, “Uh okay, what’s the steepest hill in San Francisco?” And if you do a Google search, it turns out that there is this hill. Okay, which has a slope of 41%. And uh in fact, it it was not trained with that big a range in simulation. But uh there you have it. This is of course the Digit robot from Jonathan Hurst. Hurst, it it I mean, this is one of the uh actually the humanoid that’s manufactured in the US and is available for you to buy. So, uh there’s some scientific ideas here. We There’s a stage of self-supervised learning and there’s a stage of reinforcement learning. The self-supervised learning is some kind of next token prediction on sensory motor sequences and that’s the kind of stuff you can read in the paper. So, let me make a sweeping remark. Probably offend people in this process, but I would argue locomotion, we have good progress. Navigation is nearly done. But I will argue that in manipulation, we have a long way to go and most of my talk will focus on manipulation. And there is like the generic answer which everybody knows which is that the big challenge is data. It’s much harder to obtain than for language and vision because a very fundamental reason is that the relevant data here is sensory motor trajectories. And we need sensory motor trajectories in the body of the robot. So, we don’t even have sensory motor trajectories for humans. Much less of robots. So, there’s a clear gap. And I’ll argue that this imposes significant constraints on how we should think about this problem and we cannot afford to be as wasteful in this domain as we could be in language. So, in language it’s sort of well known that an LLM has maybe on the order of 10,000 times more data than a human child. So, there’s a number like a human child hears a million words a month. So, at the age of 10 you have about 100 million words. And obviously we are operating at close to a trillion. So, there’s a 10,000 factor of redundancy in the data that is used for training an LLM. Okay, so let’s think about robotics. So, uh I think it’s very important to analyze the problem the right way. So, I’m going to start with a philosophy and how to frame the problem. I think motor control and action, the space of action is fundamentally hierarchical. So, if I talk about how I can go from here to New York, I cannot plan it at the level of muscle commands. There has to be a hierarchy. Here are three levels of the hierarchy which I think are quite conceptually convenient. The top one is that of goals and plans and it connects to cognitive vision. Middle level is motion trajectories in 4D and then there’s a level of forces and torque. So, let’s start with the top one. The top one is like, okay, I want to make an omelet. I can type that into any LLM and in fact I did. This is the answer according to Google Gemini. But there’s a sentence in there, crack eggs into the bowl. And now the question is how do you operationalize that in the body of a robot? But you do need this level and this is a level where LLMs, VLMs, I mean these are appropriate. They have that kind of knowledge for this. But how do you crack eggs into the into a bowl and it has to be done in the by the robot in the body of the robot. And I think this is a fake equation which is I think important. Action equals movement plus goal. And uh uh goal-directed action is a goal-directed behavior. We are trying to achieve some goal and there are some steps and plans and so on. Uh but the actions are achieved by movements in the physical world. So, both aspects so there’s a duality here which should be kept in mind. And language is good for the for the goal part but not for the movement part. Language is a very compression device. So, therefore you do not specify trajectories in detail. So, this is what I mean by trajectories. So, this is in the case of humans. So, So of all these actions. And I’m trying I’ve tried to pick atomic actions which are can be described at the scale of 1 or 2 seconds. And we have you can match them to verbs in a language and their verbs and nouns. And notice of course that there’s a lot of variety and generalization. So pick up and put down. I mean these are this is just like one phrase in English, but how you pick up an object you could be picking up a sugar cube which is 1 cm, you could be picking up a big tray which is much heavier. We use the same term. So the semantics is what is captured. The actual motion trajectory can vary a lot depending on the nature of the object. But these are kinematic trajectories. And this is not enough. I will argue that there is a third level. And the third level is that of forces. F equals MA. Uh what Newton told us. Uh to change things in the world, you need to apply forces. Here’s a very simple example from a neuroscience paper of how what it what happens when you pick up an object. So you apply a a normal force that causes with the because of friction it causes some force which enables the object not to slip. So there’s a grip force and there’s a vertical movement which requires that you compensate for gravity and so on. As you do various events, these the various contact events have happened, various forces change, positions change, and then there are sensors in the in the skin in the the tactile signal and and your fingertips which is being used to measure this these are these have names like Merkel receptors, Meissner, etc. etc. So all of this plays a role. So uh so to so these so that’s sort of an organizing principle for thinking about robotics. Before we get into the various machinery and the flavor of the day, you know, this VLA, this VLA, this world model. Let’s remember the these central aspects which any theory will have to address. Then I want to start with a statement in which I since I don’t want to argue about it, I’m just going to state that as a religious belief. So it’s like credo, right? It’s my credo. So if it was a Middle Ages, you didn’t like it, you could burn me on the stake, but but I’m going to stick to this. Which is that to achieve human-like physical intelligence, we need multi-fingered hands. Okay. And here is a kind of proof for some people. And so these are experiments done by my student Toru Lin, which are examples of simple tasks that we do every day where if you try to do them with parallel jaw grippers, you might be able to do some of them, but with very fairly high failure rate. So if that bottle of wine was worth $250, you probably don’t you want really want a power grasp rather than two-point grasp. And so forth. So these are different arguments, but I I’m I’m not going to argue about this. This is my credo. Okay. Uh how should So everything I do is in this in my lab, we don’t do parallel jaw grippers. Okay. We we need some to draw some line somewhere. Okay. So how should we train robots? Okay, so this is of course the question that everybody’s talking about. And again, I’m giving you the clichéd answer. It’s well behind natural language processing and computer vision. And this is because of the shortage of data. So how will we solve it? Again, there are three standard answers now. This probably there’s it’s not controversial. Teleoperation, videos of humans, simulation, and there are variations on the theme. Okay, so let’s do teleoperation. Okay. So again, there are flavors of teleoperations, so I’m going to critique this particular one. So this is the most common one. I think it’s common where you are essentially have what’s called a leader-follower system. A human operates a robot. A a a human has the is So this is I think Zipeng. He’s trying to operate something and then exactly the same actions are mirrored in the follower robot. And then if this video plays, you’ll see you know, okay, so now you’re recording trajectories corresponding to how to remove a paper towel. And these trajectories are now in the body of the robot. So people love this. This is the dominant approach. This is where there are maybe thousands or millions of people all around the world doing this this I think really painful job. Okay, to convince you just try to do it. Why is it painful? It is painful because it is making a human behave like a robot. It’s really it’s like a torture session almost, right? You have to you are because what’s happening is the main cue sensory modality being used is vision. You’re not getting haptic feedback. You’re looking at this object which are far away and you’re trying to be this puppeteer essentially getting the robot to do this. And so essentially when trajectories are recorded like this, they are very slow. They are very slow because this poor human is suffering. And so therefore all those results are after that sped up by 5x so as to make them appealing to you. So I I think these are the two problems. First is human sensory motor limitations. The trajectories are in the right embodiment, but they are collected by a human who is not being natural and therefore they are very slow and clumsy and difficult. And then there’s a second argument which I think is a weaker argument which is the one about domain gap between training and test conditions and probably with enough money you can solve that. But you can’t use money to solve the first problem because those are essentially the limitations of human. The human visual system is being used as the only sensory system whereas when we actually do task we also make use of haptics. Okay, this is the next thing we could do videos as a computer vision person or an ex computer vision person I love this. I think this is the natural signal and these are this is this is again now being done collected by thousands of people now. It’s now I think probably this is the the derivative of this is high. Lots of people are doing this. This is egocentric video. You could also have exocentric video. And cool. Okay, what’s the problem with this? So the problem with this is that ultimately you are doing collecting this data in the human embodiment and which is not the same as the robot embodiment. So there is an embodiment gap because ultimately it has to be executed by the robot. So the teleop people will say, “Oh, you’ve got a human trajectory. It’s not a robot trajectory. You’ll have to do some translation.” Fair point. Fair point. Uh another point is I think is that you’re missing information about contact forces because I told you that we should study robotics at those three levels and the level of forces is very important and you don’t have that information if you just collect visual data. Okay, so then we go to approach three which is simulation. And in simulation you have everything, right? Because you are it’s a physics simulator so you have forces. And uh and and this is what? By the way, the one existence proof we have, I mean I gave you the I talked about navigation and I talked about locomotion and both of those were cracked with simulation. So in manipulation say people say, “Oh, we can’t use simulation.” But I see the one existence proof you have is in fact of simulation. So this should be in fact be your prior. And then there’s a critique of these. I mean, there’s a first critique which is a sim to real gap. This is like the standard answer. Anybody who creates a simulator but oh, is it realistic enough? Is it varied enough? I don’t buy this. Yes, those are problems, but those are fundamentally surmountable problems. It’s just more engineering and we’ll deal with it. The second problem is a real problem. Which is defining RL rewards for specifying tasks. And it’s easy for locomotion and hard for manipulation. So, let me give you a flavor. So, this is from a project from student of mine. And he’s trying to train this robot to pick up and do some pouring task. Okay. Like I initially would have thought, okay, reasonable. What do you need to do? You need to do a lot and lot of this kind of engineering. You need to specify every stage painfully. And then each of those stages corresponds to a certain term in the in the reward function. And and this is like 10, but there might be some task where this list goes to 150. And then there are coefficients for this and I think this is actually the reason why people have given up. Whereas when you do behavior cloning on top of teleop trajectories, you are good to go on day two. You just download some software on training diffusion policies or whatever and you’re you’re good to go. Okay. So, now I’ve put myself into a corner. I’ve took the three standard approaches. I’ve said, this is bad, this is bad, this is bad. So, what the hell am I going to do, right? Okay. So, so so now after you know, conveying pessimism, now let’s have hope. Okay. And I I I I always I I this this quote I’ve been using for like 10 years and so long as I’m in research, I’ll keep using this. This is This is This is from Turing and there’s a sentence there which is instead of trying to produce a program to simulate the adult mind, why not rather try to produce one which simulates the child’s and then if this were then subjected to a course of education, we could build a robot. I think this is fundamentally right. The devil is in the details and there are lots of details. And Turing was writing this paper in 1950 and now it’s 26, so that’s 76 years later. And in the meanwhile, all our colleagues in psychology have been doing a lot of hard work and they have come up written hundreds and thousands of papers. I will not summarize them in the next 15 minutes, but read one of these review papers, talk to your favorite child psychologist, one of these. But there are some insights from all that work. I’ll just flash them here. Multimodality is important. Vision, touch and proprioception. Hierarchies, I already mentioned that. Physical search is okay. Social imitation learning, language, there’s a role for all this. But I I I think this is essentially opening up a whole can of worms. So I would like to pick one nugget and and I’m I’m going to focus on that for the rest of my talk. And this is how imitation learning actually is. You are being fed a story when you are being told imitation learning is behavior cloning. Behavior cloning would mean that the adult controls the neural circuits of the child so that the child can pick up the object. This ain’t happening. This is not how it happens. This is how it happens. So there’s this father and the little daughter and he’s washing some dishes and the girl is observing him. Right? So she’s This is how cultural knowledge is transmitted. You know, your parents or your peers, or your teachers show you various tricks. If you are a bicycle mechanic, you go and learn in as an apprentice in a shop. Okay, but she cannot directly do this because I told you visual imitation has some problems, and the problems are that you don’t know about contacts and forces, not in detail. I mean, you get some approximate idea, but not a lot. And the body is different. There’s an embodiment gap. So, we talked about the embodiment gap for robots and humans, but there is an embodiment gap for this girl and her father. I mean, she has itty-bitty hands. Her hands are much smaller. Right? And she has missing data, which is like forces. So, what she has to do is she has to learn to do the task in her own body. With the guidance, a little hint provided by the staff from observing her father. And she’s doing this. She’s playing. She’s doing a lot of trial and error and play. That’s it. So, to me, this is the formula. Just replace the girl by a robot. Done. I intend to devote the next few years of my life to making this slide as the way to solve robotic manipulation. Okay. Now, how far have we got? So, I’m going to start with an example which is not for manipulation. It’s for locomotion. It’s so-called perceptive locomotion. This was a coral paper last year. They actually won a prize. And uh, what we did here was we are teaching So, this is perceptive locomotion. So, there’s a distinction which is drawn between blind locomotion and perceptive locomotion, where you’re using vision. And this is what you need when you are going to go up and down stairs and difficult terrain. So, this is what we have here. So, this is the robot roaming around in various places in Berkeley. And many people can do this now. I mean this is some time ago, but I want to emphasize I won’t let Okay, I think that it looks a little bit drunk, but the answer in all RL answers is if I just train it a bit more it will not no longer look at But how we trained it? How did we train it? I think this is the important point. How we trained it is by collected RGB video of humans. And then we do 3D reconstruction. And this is computer vision technology. Okay? And then on we then we put that into a physics simulator and then we do RL on top of that to make it robust because just open loop is not going to be enough. So So this is Arthur and he’s going to do various actions. Now you have the 3D reconstruction of Arthur. And then there is the humanoid the retargeting of Arthur to a humanoid and then in some terrain. And now you just do collect more of this, but not a lot. I mean in this case I think there were like a total of 100 trajectories. So it’s a much smaller number than the thousands of hours that people collect. It’s much smaller because the variations are created inside the simulator. This is providing you the the knowledge that you need to learn from others and then the rest you have to do in your own body. Because the body of this humanoid is not the same as the body of a graduate student. And and of course because we have this physical reconstruction, you can construct what the view looks like from various points of view and so on and so forth. So and you learn how to sit etc. etc. And this formula now is by the way widely being deployed. So I think a lot of humanoid stuff with the perceptive side of it is always done by these tracking approaches where you use a human to define a set of trajectories and then you copy them in the body of the robot. Okay. Okay. So now how much time do I have? Somebody keeping track?

25 minutes left. Perfect. Thank you. Okay. So visual imitation accompanied by trial and error. Back to that philosophy slide. So we have a success story which is in locomotion but I already said locomotion is easier than manipulation so I need to figure out convey the story for the more complex case. Okay. So how do we do this? So I was So let’s just think through what is needed. So what’s needed is learning from video the cultural aspects, the affordances and so on. And then practicing in your own body which is to be done in a physics simulator of some sort. So now the data problem is gone. The data problem is I mean YouTube has 5 million plus videos. I looked a few years ago it was 156 million hours. So you do not need to pay all those people who are suffering trying to do this teleoperation task somewhere, you know, making poor humans operate like robots. No. There are humans who have consciously done this as humans. Okay. So uh And and then this is the story. I mean this is really a hot off the press because uh my students who are have submitted a paper to Coral. This is a photo from there. So the top row you have some I mean I’m showing you frames. Uh these are actually videos. But uh what you have done is we’ve taken that human action being performed by human and what do you see in the bottom row? The same action being performed by a robot. Okay, which is in this case equipped with sharper hands. Okay, so there’s re-targeting. So, there’s this kitchen whisking operation or there is this brushing operation and there’s a pouring operation. Of course, these are cherry-picked videos like they always are. But, you get this is meant to illustrate the spirit. You have the data of human videos. You will do 3D-fy it in some way just like we did for walking. 3D-fy it or rather 4D-fy create a 4D thing. And then you do the re-targeting. So, just like the girl has to solve the problem of re-targeting from the adult’s body to her body, here we have to re-target from the human body to the robot’s body. Okay. Now, this is the This is where we are as of yesterday, okay? But, I have been in this enterprise for years now, for five or six years. Ever since I got into robotics, I felt I have a little bit of experience in computer vision. And if I can do something from it, I should be able to solve this problem. And it has proved to be much harder than I thought it would be. Because in 2020, I thought, “Oh, we can do it in 2 years or 1 year.” And it’s 2026 and even now it’s not fully solved. It’s close to solved, but not fully solved. But, let me take you through the journey. This is a recent paper, actually. I mean, we and I’ll show you some results from this. Okay, so these are This is the first stage. I think the hard part is the first stage. How to 3D-fy robustly from general RGB videos. By the way, there’s one trick you can use. Collect RGBD videos. So, if you have depth already, it makes life much easier. So, that lots of people are doing and and we are also doing. So, I think that’s a practical solution now. But, I want to solve this with RGB video because then it becomes then the data problem just disappears. Everything is there on YouTube. So, these are examples of those reconstructions. And here is a way to show I mean, look, this is the beauty of 3D, right? Once you get if you 3D-fy once, you’ve got you can now you get it from all viewpoints and you get at least approximate contacts and so on and so forth, right? And now, of course, this is not the end point. After this, you have to retarget to the robot. And then, it’s not the end point because then you get an open-loop trajectory. On top of that, you have to do RL training to to get a robust policy. So, there are steps, but anyway, you get the flavor. Okay, in the course of this so, this led us to focus a lot on what are the relevant problems in computer vision which are solved and not solved. So, we for several years, starting about 2018 with Anju when Anju Kanazawa came to my lab as a post-doc and now she’s a professor, we worked on the problem of 3D-fying humans. 3D-fying humans, then we did hands. This was with George Pavlakos. And after fighting that battle for 6 years, I feel okay, we’ve solved it. We are done. Okay, I’m I’m not I mean, I’m joking a bit. I mean, there’s more to do, but but we got a solution which was good enough for the robotics version of the problem. So, that was the human part. Then there was the object part and it turned out that even though that has been worked on very for a very long time, it actually turns out to be the challenge. So, you might think, oh, 3D vision, it’s solved. It’s like there are so many versions. No, the typical way in which people approach 3D is as multi-view geometry, but I I want to approach it as multi-view geometry. I’m going to just have some video and often the viewpoint may not change much. There’s may not be much of a parallax signal. So, I want to approach 3D-fication really by recognition. It’s reconstruction by recognition. The way I can recognize any object is often by uh by So, here is the scene. There are chairs, plates, etc. At some time in the past, I’ve seen it from different viewpoints and I’ve built a 3D model. model. So, when I see it as just a single image, it connects to my memory and then I retrieve a certain 3D shape. And and that’s uh what this uh paper is about, Sham 3D. So, this was something we did with a team at Meta and this paper will appear at CVPR next week. But, look at the reconstructions here. So, the reconstructions are not just a 2.5D sketch. It’s not a point cloud. It’s for distinct objects and you have sort of the full 3D reconstruction. Okay? And and you might ask like why did this have to wait till 2026? I mean, we’ve been doing multi-view geometry for like 150 years, right? The photogrammetrists were at it. And no, what The reason is that it’s a shortage of data. So, ultimately we are going to solve this as a supervised learning problem, but a supervised learning problem needs data. So, you need paired data and that’s hard to get for the full variety of a complex scene. So, I won’t get into the details here. I mean, there’s a You can read the paper. It’s at CVPR. It’s It’s as, you know, sort of the usual ingredients, lots of transformers here and there. Uh predicting uh you know, you know, first a volumetric representation, then meshes, Gaussians, splats, etc., etc. I think the innovative part is actually the training. And what we had to do was how do we get this data? So, yes, for isolated objects you have data like these CAD models, but the CAD models don’t give you how the objects look like in real scenes. But you can use that for pre-training and then there are various stages of mid-training and fine-tuning. And what we used was the fact that humans are not necessarily good at reconstructing objects metrically, but if you give them two reconstructions, they can choose which one is better. So, use humans in the loop for judging that. And then we, at the ultimate stage, we use these computer graphics artists for some really tough nuts which for which you can’t get a reconstruction. Anyway, long story, but but let me do this show off video. So, here’s a scene, here’s an image. And then okay, sorry. Let’s see if there’s some video which will play. Uh I was trying to play a video, but okay, I I won’t get to play the video here, but each of these objects in that scene, you just have one image. And for each of these there’s a 3D reconstruction here and you can spin them around. And look at the detail in each of these reconstructions. And these objects are not there in the training set. The way it works is that once you have reconstructed a million objects, every new object is like a combination of old objects. And and so it works. Okay, how much time? Somebody Huh? 34. Okay. 10 minutes. Okay, great. So, so this works. And by the way, this is open. You can download it and use it. And this is what was needed at the end of the day. We really needed a very powerful model for doing 3D reconstruction from a single image. And then And course, there are some tricks when you’re reconstructing from a video stream. And these are examples. So at finally after like years of this struggle, these are examples of the kinds of reconstruction we can do. So these are all in 3D. So there’s a 3D hand and a 3D object. Okay. And I think this is getting there. I think there’s probably still a bit more work to do, but but this is like 80% of the way, I think. So now the data problem just disappears. You have just go to YouTube, download videos, 3D-fy them, you got the trajectory, retarget to a robot, you’re you’re off to the races. Okay, I’ll show you some a few slides. So now I I need to show some results. And as I said, I think on dexterous manipulation we are a long way to go. These are my results are not going to be the most impressive demos that you can see in dexterous manipulation, but I want to show these because these are consistent with my philosophy and paradigm. And these are examples which show dexterous manipulation tasks. So this is these are this is a with a my student Toru Lin, and she was doing did an internship at Nvidia, and this is very much in the simulation paradigm, but it’s all as you see dexterous manipulation. And and here there is has to be there is some amount of that reward hacking going on, which we want to sort of avoid. And some lessons from this I think sim-to-real, I want to repeat, sim-to-real is just an engineering problem. You put in the effort, you’ll solve it. It’s a bit like when people started to use neural networks in 2012, 2013, they would always complain, “SVMs are so easy. With neural networks I have to tune so many parameters.” Well, you learned how to tune those parameters. Okay. Okay, I’ll skip this. Uh these are just more dexterous manipulation policies. I think I’ll Okay, just to give you a flavor that this is real. But I mean this is just one small set of examples. I mean the the set of manipulation tasks is infinite. So, there’s a long way to go. Okay. Okay, now I need to address the issue of forces because we didn’t so far. And how do we do this? So, I think there are there are several ingredients to a solution. One ingredient is just better sensing, which means understanding forces. And you can do it at the finger level and you can do it at the joints. I mean humans do it with sensors in various places. So, there’s tactile sensing, but also sensors at your joints and in your muscles. So, this is an example of the tactile sensing solution, the the gold-plated Cadillac version, I should say, Digit 360. And this is like an amazing sensor and it’s it’s probably overkill. I mean, I think I think that right now the downstream uses of tactile sensing have been rather limited and we really need to do a lot of research showing that tactile sensing is indispensable, which is what I believe. I don’t think that we can solve robotics without solving the problem of tactile sensing. But okay, there’s an And there are other designs. I mean, this design itself builds on a long tradition going back to GelSight, but there are other technologies. I I’ll give you this task not because this task itself is it was challenging 3 years ago. I mean, now everybody can do this, but the one of the insights from that was is this figure, which was we compared the performance uh with just vision, with just proprioception, just vision, vision plus touch, only touch. And it turns out that the the rightmost bar bar is the ground truth kind of thing, but the bar next to it is what you get from vision plus proprio plus touch, which almost approaches ground truth knowledge. And I think this is going to be the answer always. You want to use all your sensory capabilities, and if you don’t use touch, and lots of people don’t use touch, you will get a certain performance, you will get a certain cherry-pick demo you if you like, but your performance will and the numbers you will lose. You will not get to 99, you will get to

  1. And that may be all the difference between being real or not. Here’s another uh result, which is uh this is a more recent paper. Uh this is with the Sharper Hand, it has a certain sensing technology. And uh and then the the nice finding is that when you have more powerful sensing, you need far fewer demos. So, the number of examples being used here is like on the order of a few hundred, not on the order of a few thousand. So, the policies, like the typical VLA type policies, they need a few thousand trials to get to like 80% or something. And uh here you get quite high performance. And then we have the same finding, which is yeah, the same finding, which is that uh you the different sensory modalities all uh need to be there. And if you don’t use touch, you you lose. And there’s then there’s a question of architectures which exploit touch, and basically you need to do next touch prediction modeling. And uh uh okay, these are this is just the advertisement copy. Uh uh but here’s another example which shows that there is value in real-world reinforcement learning. So, we were trying to take to study a task which requires a combination of position and force control. And this is like peeling vegetables with a knife and without using a special end effector. I mean, you just Okay. And if you have to have force sensing here, and it turns out that the kind of machinery that has been developed in LLMs is RLAIF idea that you apply like how do you know what is a good peeling operation? Well, you have to have a human who judges this is a good peeling job, this is a bad peeling job. And and then you can use that to train the policy. And then here are some results. I’m sure that this is not ready for a a fancy restaurant, but uh Michelin three-star restaurants do not have to worry just yet, but it’s a step in the right direction. I’m going to make my last concluding slide. So, so more general remarks. So, [snorts] I I feel that I feel that focusing on VLMs and VLAs is a distraction for the robotics community. They’re focusing on the wrong parts of the problem. Okay, we can take making an omelet and translate it into a recipe. LLMs can do that, but this is not where the problem is. The problem is how to crack an egg and how to crack it successfully 99.99% of the time and to do it with generalization. So, generalization is important. Generalization across location, generalization across different objects in a category, generalization across clutter. And that means fighting the So, there are these three levels that I believe in. Right? The first level I call the level of Aristotle, the level of high-level planning. The middle level is the level of Euclid. The bottom level is the level of Newton. I think most of the action is the level of Euclid and Newton. And at the end we’ll need to have language because we need to communicate to the robot. But that’s the easy part. The hard part is the good old fight of contact forces, trajectories. Thank you very much.

[applause] [applause] Time for one question or whatever. Okay. So we have time for questions. I don’t know who’s managing the I think Mark had to go for something. But I I’ll I’ll take any questions if there’s a way to get whoever asked the question to have a mic. Or or you can come up if you have a question. Okay, any questions? Yes. Okay, can you take it to her? Yeah. A philosophical question. To the very In the very beginning you said robotics is about linking perception to action. Shouldn’t one turn this thing around? Robotics is about solving some tasks. So first you have some goals, some task. And then depending on the task, you know, you open your eyes or you listen to different channels. I think that turns the the whole problem around. You don’t have to perceive everything. You perceive what you need to perceive. Depending on your task. Uh yeah, but nothing that I said here is counter to that. I’m I’m I I mean I this is Okay, so the person who articulated this was Gibson. Okay, and then we see in order to move and we move in order to see. And actually before Gibson there was this gentleman called Uexküll. Okay, we talked about the Umwelt. I I believe in this. Okay, so here but and if you have very small animals like flies, I mean, it’s going to be very the the it’s it’s a very clear direct thing. At the level of humans, what has happened is that there’s a such a wide range of tasks that we can perform that our vision system effectively lands are being somewhat general purpose. Now, what you focus on in a particular for a particular task will be just one subset of the input, but we have a pretty powerful uh general purpose vision system. And I think that for a task there isn’t that much fine-tuning of vision. So, I’m I I think generally the idea of totally doing pixels to tasks is I think a bad idea. Because the vision system that you’ll develop in the course of that is going to be a very flimsy vision system. It’ll have had access to rather little data. Whereas why not throw a beefy foundation model of vision which has been trained on like millions and billions of images? And but its use needs to be modulated according to the task. And how precisely to do that is is our research problem. So, I’m not saying I have a pat answer for that. Maybe to just say with human analogy, I don’t think that we let all our cortex light up when we solve a particular particular task. So, we kind of don’t throw the large lights up. I mean, the early stages light up. Later stages are very much modulated by the task. So, I mean, your your what your retina sees is the same independent of what I mean, the retina doesn’t get modulated from top-down input. I At the later stages, it does. I’m not trying to make this point too extreme. Uh but uh you you have a fairly general So, we we we have we’ve been playing this game, right? And like the game is essentially robotics people think vision people are no good. Everything they did we can throw away. I will just train my vision encoder in the context of a task. And I think this does not work. I will I will argue that it does not work. I have very hard evidence for it in the navigation task. Because what you land up doing is over fitting to those trajectories that you explored, and that’s usually a very small subset. So, that encoder doesn’t work. So, we need to have build vision encoders on huge amounts of data, millions of hours of video, whatever. And then just fine-tune them for a particular task. Thanks. Okay, one more question there then probably we It’s Vladlen’s turn. Hello. So, you said that the sim-to-real gap is just an engineering problem. Um so, my question is, do you think that we’re able to model the physics and its complexity, and especially for like dexterous manipulation, um with like classical physics engine, also like for deformable objects, which are like harder to model, right? And its friction and so on. Or do you think that like um like they have the hypothesis that like through video models we could like learn in latent space latent space um like extract physics by learning through video itself instead of handcrafting these uh features. Um do you think that like like neuro simulators based on world model approaches would be like the next step? Or do you think that it’s possible just with handcrafted physics features? I I I think everything should be tried. So, certainly the kind of simulators that you can learn from world models trained on lots of videos as a plausible way to go. I I uh we can also use physics simulators. I want to make one point which is that we do not need detailed physics. So in particular just take locomotion as an example. We have these locomotion systems which are trained with simulators. So like the work from Berkeley the the RMA work from like 2021. The simulator was just some used a fractal terrains. The fractal terrains do not capture the full variety of the real world like walking on a beach or something like that. But it turns out that these are proxies. See the world manifests itself to the to the robot through a low dimensional interface. The low dimensional interface is the the the parts of the actuators and the joints and the hands and the feet and so on. So so essentially you’re going to have a metamerism problem. But metamerism is it is good. So okay. In in human perception color vision is mediated through three cone types. The distribution of light wavelengths is infinite. But it’s all projected down to the three-dimensional subspace, right? That’s why you can mix colors and create new and with with just three a basis set of three. So the same thing is happening in manipulation. You have ultimately the you have your your end effectors and you’re going to feel forces through them and you have to stay balanced and so on. So you essentially have a low dimensional interface and a low dimensional latent. The full complexity of the world doesn’t matter. It is the projection onto this latent space. But could one not argue that human actually get feedback from the real world whereas in the simulator you only get feedback from the simulator. So if like the physics priors are wrong for let’s say deformable object I I the classic answer all models are wrong but some models are useful. So, the question is are they good enough for certain purposes? And I I think that depends. So, I’m not cannot give a crisp a universal answer, but that’s where there is some modeling to be done. And uh yeah. So, I I like in locomotion, we have simulators which are very crude and yet we could train policies which worked. In uh dexterous manipulation, the demands are much greater. So, currently, we could build better simulators, but they would be too slow. But, I think all this will get solved over time. Okay, cool. Thank you. Thank you. Thank you very much, well. Okay, thank you, Chaitanya. Uh let’s thank speaker again. [applause]