Why Robots Still Struggle With Simple Tasks (And What Might Finally Change That) | Karol Hausman, Co-Founder & CEO of Physical Intelligence
Karol Hausman is the co-founder and CEO of Physical Intelligence, a robotics company building a general-purpose “AI brain for the physical world.” The company has raised more than $1 billion in funding to develop foundation models that allow robots to operate across many machines, environments, and tasks rather than being programmed for a single purpose. The core thesis: the same scaling dynamics that transformed language models may also unlock robotic intelligence. But only if you resist every commercial pressure pushing you toward specialization. The central challenge isn’t mechanical design. It’s intelligence: how robots learn, generalize, and interact with a physical world that is far harder to simulate than it is to describe.
Appears in
- Uploaded
- Uploaded Jun 5, 2026
- File type
- POD
- Queried
- 1
Full transcript
Showing the full transcript for this episode.
[00:00] The more robots you deploy, the better the models should get, because I think the models will keep on getting better. They'll be able to absorb more and more data. As I started looking into robotics and started studying it, it was very disappointing. It seemed like there are no intelligent robots out there and no one was working with them. Every robot you would see would be this pre-programmed machine that just goes from A to B and does it repeatedly and has no intelligence whatsoever. That was maybe like a wake-up call.
So many people excited about this future that we saw in the movies, that we read about in the books, and it seems like no one was working on it. [00:30] particular experiment with a coke can in front of a robot and three pictures of different celebrities in front of it. And the prompt we gave to the model we built was put a coke can on a picture of Taylor Swift. The robot picked it up and then slowly moved it towards Taylor Swift, all from internet data. That was the moment where it clicked for us, where you can bring in a lot of prior knowledge from LLMs, from the internet, and connect it to robot motions.
And it felt like it opened another door. Maybe, just maybe, if we do everything right, if we combine it with internet knowledge, [01:00] if we do all of the pieces that need to be done, it might work. And at that point, it became clear that the way to accomplish this is to create an organization whose sole purpose is to solve physical intelligence. [01:18] We taught machines to beat the world's best chess players, long before we taught them to reliably fold a towel or carry a coffee cup. [01:27] That paradox is at the heart [01:29] of the robotics industry, responsible for many of its false dawns over the years.
[01:33] Carol Hausmann, [01:35] believes that he's found the way forward. [01:38] The CEO of Physical Intelligence has raised more than a billion dollars in venture funding. [01:42] to build an AI brain [01:44] for the real world. [01:45] one that helps robots [01:47] generalize accurately across different tasks [01:50] and environments. [01:51] he's doing so by taking a very different approach, [01:54] from many of his best funded competitors. [01:56] focusing on gathering data from actual interactions [02:00] rather than through stimulation. [02:02] In today's episode, [02:04] Carl and I discuss lessons learned from the inner game of tennis, [02:07] Why reinforcement learning is making a comeback [02:10] and the challenges of creating true [02:13] physical intelligence.
[02:15] I'm Mario. [02:16] And this is The Generalist. [02:18] This episode is brought to you by Brex. If you're a founder, the hardest part isn't the idea. It's scaling fast without getting buried in back office work. That's where Brex comes in. Brex is the intelligent finance platform for founders. [02:32] With Brex, you get high-limit corporate cards, easy banking, and high-yield treasury. [02:37] plus a team of AI agents [02:39] that handle manual finance tasks for you. [02:41] They take care of things like expenses, [02:44] all according to your rules, so you can move faster while staying in full control.
[02:49] One in three startups in the S. already runs on Brex. You can, too, at com slash Mario. [02:57] I'm really excited about today's sponsor, Granola. Simply put, Granola is the AI notepad for people in back-to-back meetings. [03:05] I've been using Granola for over a year now, and honestly, it's a tool that has transformed the way I work. [03:11] Granola takes meeting notes for you without any intrusive bots joining your calls. You can jot down rough notes like you always do, [03:18] And in the background, Granola transcribes and turns those notes into clear, useful notes when the meeting ends.
[03:25] You can also chat with your notes, which is one of my favorite features. [03:29] If someone says something on the call that you didn't quite catch or want to learn more about, [03:33] Granola can help you out. It's an amazing way to be better informed during a conversation without having to interrupt everyone else's flow. [03:41] You can also have Granola review all your recent conversations [03:45] to pull out to-dos, write a weekly recap, or surface interesting ideas you might have forgotten. Another thing I love. To get started with Granola, head to
ai slash mario. [03:56] And for new users, you can get three months free with the code Mario. So go to ai slash Mario and use code Mario for three months free. [04:07] I'm so excited to have this conversation today. In studying [04:11] what you're building with physical intelligence, [04:14] Naturally, much of the conversation is around the technology. But the more I looked into your story and the company, it feels like what you're building is. [04:23] really personal and connects to some of your [04:26] your background in philosophy and all these other things.
So I'd love to start maybe with some of [04:32] Some of that first. When were your earliest memories of being excited by robots? [04:37] Probably with watching movies as a kid. [04:39] I think I was a big Star Wars fan. [04:43] I was reading all the books, watching all the movies multiple times. [04:48] And I really love the way that robots were portrayed in that world, the diversity of them, different functionalities, the way they interacted with people. [04:57] the way that there are good robots and bad robots. But I don't know if I can pinpoint a specific point.
[05:03] um [05:04] I remember just growing up being fascinated by that concept, that you can build something that is so human-like. And I always thought that if you were able to do that, [05:16] it would allow us to understand ourselves a little bit better. [05:19] And. [05:21] I was fascinated by... [05:22] Kind of the things that I think people are just fascinated with as they're growing up, you know, asking deep questions about... [05:30] the nature of reality, how the brain works. [05:33] Um... [05:34] what does it mean to think? And I thought that robots [05:39] Probably you're going to have some answers to these questions.
[05:43] because reading other kinds of books or studying philosophy didn't have good answers to these questions. [05:49] Then, more importantly, as I started looking into robotics and started studying it, [05:54] the first impression I had was that it was very disappointing. In that every robot you would see, [06:00] with [06:01] Basically, it seemed like there are no intelligent robots out there and no one was working with them. [06:06] every robot you would see would be this pre-programmed machine that just goes from A to B. [06:11] and does it repeatedly. [06:13] and has no intelligence whatsoever.
[06:16] So I think that was maybe the... [06:18] the most defining moment for me when I saw that. [06:21] Because that was like a wake-up call of [06:24] You know, there's so many people excited about this future that we saw in the movies, that we read about in the books. [06:29] And it seems like no one is working on it. You said that, you know, in some ways you had, you know, that you're asking the classic questions of every child about the nature of the universe and the human mind. I don't know if that's as universal as as you might imagine it to be.
I think probably there are some kids who really become fascinated by that, but maybe not all. Were your parents scientists? Were there things that they did that sort of encouraged that natural curiosity that you must have had? [06:54] Not really, no. My dad was a mechanic and my mom was an entrepreneur. So not really, there wasn't a lot of science at home. [07:08] I think it was mostly coming from watching movies and being fascinated by it. [07:13] I also went to, so I come from a small town in Poland. I went to high school that had very high talent density for that.
[07:21] for that area. [07:22] and [07:23] I remember that being an environment where [07:26] a lot of these questions were being asked and it was very encouraged. So I found [07:31] a group of friends that were really fascinated by similar things. And I think that would really increase my interest in that area. It was just a thing that you would talk about in high school. You had a tweet that I saw where you were talking about Fei-Fei Li's biography and how it sort of spoke to you in some sense that, you know, the immigration story and also the fact of, you know, having teachers that believed in you along the way were sort of parallels to your own life.
In high school, was that when you sort of had that first batch of those teachers who [08:01] saw some of that, you know, really remarkable early talent? I think it was throughout. It wasn't just high school. It was, it was throughout. There were multiple people that now looking back, I'm really grateful to. [08:13] that either saw something or just encouraged me to dig deeper. [08:18] or motivated me or gave me some opportunities that I wouldn't have otherwise. [08:23] And I think at every single stage, there is someone... [08:27] I could point to.
And reading Feifei's book just reminded me of that. And I remember after I read it, [08:33] I pinged some of them and thanked them. [08:36] for everything they did for me. So yeah, it was a long journey from [08:41] From Poland through Germany. [08:43] to the US, going through multiple places in the US and [08:47] It felt like at every single stage of the journey, there was someone kind of rooting for me and [08:52] and helping me. [08:53] you know, before... [08:55] you realized that you wanted to work in robotics in particular.
As an undergrad, you studied sort of computer science and philosophy. But I'm curious if there were interests before that, like as a kid, were you thinking I'm going to be a astrophysicist or I don't know, a neuroscientist? [09:12] Yeah, I was really into physics actually. Right after high school I applied to a lot of physics and math programs. [09:19] And I thought I was going to be an astrophysicist or a theoretical physicist. [09:24] I was really fascinated by it. I was reading a lot of books. I was really into physics in high school.
But then a lot of my friends decided to go for engineering degrees and we're all going to go to Warsaw. And I thought that. [09:36] I could maybe try to study both or I could start with engineering and then go back to physics if it turns out to be boring. It wasn't like a very intentional choice. [09:45] At that time, [09:46] It was more like... [09:47] These are really smart people. I really want to stay with my friends. [09:51] And robots are really, really cool. [09:53] So let me try that and then see how it goes.
[09:56] So I went to Warsaw to study [10:00] It was called mechatronics at the time, which was basically like a mixture of electronics, mechanical engineering, computer science. [10:07] Then I also studied philosophy. [10:09] But I think... [10:10] That's when, once I moved to Warsaw, that's when I started realizing that robotics is a thing, that it's a thing that is very disappointing. [10:18] I remember studying robotics in Warsaw. The kind of the dream job you could get was [10:24] being a roboticist at an Avon factory that was right next to Warsaw.
[10:30] That was the job that everybody was competing for, if you could work for Avon on the production line. [10:36] That's how you made it. [10:38] in the robotics industry. [10:39] And I remember I went there once and saw that production line, and it was the most uninspiring thing. [10:45] I've heard a scene with robots. As I said before, robots just moving from A to B. [10:50] not being intelligent whatsoever. And I think that's when [10:54] when it started clicking for me that I would really like to do something beyond that.
[10:58] and it seemed like no one was working on it, which sparked my interest even more. [11:02] So that's where I started thinking who is doing intelligent robots and is it even a thing. [11:07] And that's then what defined, I think, the rest of the journey. [11:10] of trying to find those people [11:13] and move from place to place as I try to find them. [11:17] And then eventually start working on it myself. The sort of study of philosophy, as you've looked back on that time, you know, maybe it's a line on the CV at this point and you're like, you know, I don't reflect on it necessarily that much.
But I'm curious if there are writers that you've ended up coming back to more and more as you start to do this sort of work around philosophy. [11:39] I don't know, the nature of consciousness or the, you know, the certain senses that make us more or less human. [11:44] It was a really interesting period. [11:47] I think there was a few realizations I had while studying it. [11:51] I think one... [11:52] was that... [11:54] A lot of philosophy is the history of philosophy. It's not [11:59] Philosophy itself. [12:00] It's talking about what others thought about.
[12:03] Years ago. [12:04] From the outside looking in, it always sounds like Plato or... [12:09] or people like that, they're the... [12:12] They were these geniuses that kind of like figured everything out way before their time. [12:17] But then when you actually read their works, and that was part of the deal of studying philosophy, [12:23] you realize that there's a very small percentage of things that kind of sound right. And there are a lot of things that they said that are absolutely wrong. Yes. [12:31] And that was quite eye-opening. The history of philosophy shows that there are people who made many, many mistakes along the way.
[12:39] and thought that the world worked a certain way when it's now very clear that it doesn't. They had some really interesting observations and I think we often over-index on those. [12:48] But we were wrong a lot. [12:50] There are still some philosophers that I remember studying that I think were very relevant and I was a big fan of. One of them is Spinoza. [13:00] That was, I think, my favorite philosopher. When I look back at what I studied, [13:06] But then there were some subjects, very, very few, that were more about philosophizing today rather than looking at the history.
[13:15] And, um, [13:16] I think those were the ones that were the most fun for me. [13:20] There was this one subject called ontology. [13:24] which was... [13:25] a study of [13:27] of things, of how they are. I think it's even difficult to comprehend what this means. But it was one of those subjects where you just like sit and contemplate reality. It just was very fun, even though it was very unstructured and it was [13:41] It wasn't clear what is right and what isn't. [13:44] It just seemed like it touched something real.
[13:48] And probably does the subject I remember the most. That's where I had the most fun. Much more so the history of philosophy. [13:55] so interesting. And these things I love to explore because there is, to me at least, such deep parallels between some of the things that you're trying to build to fruition. Why Spinoza? What was it about him that [14:06] that seemed deep to you are true. [14:09] The one thing that really resonated with me was this idea that [14:13] you know at that time [14:16] A lot of philosophers were contemplating the nature of God and the nature of reality.
[14:20] and there was [14:22] often this idea that God is this separate entity that either was the creator of everything or had some kind of way of controlling everything that is that is happening in the universe. [14:35] And I think with him, what was... [14:37] what is so interesting [14:41] was the... [14:42] He said that [14:43] everything around us, like all of reality, that is God, that is what it is. [14:48] and there is some underlying structure to it, and that structure in it itself [14:52] That's where the beauty lies and where the intelligence lies.
[14:56] And I thought that was... [14:57] very thought-provoking and [15:00] as you study physics and things like robotics or AI. [15:04] I think more and more you realize that there's some underlying structure to things. [15:08] And in some ways that structure is quite counterintuitive, like why it's there. Why is it so that you can describe all the complexity of the world in a simple equation rather than in a million different equations. And the more beautiful and short that equation is, the more real it seems. [15:26] or the more accurate it is.
[15:28] And I thought this was kind of the closest to... [15:31] what it feels like today given the current state of science. You had a tweet where you were talking about a book that I tried to read as much of it as I could. [15:41] in preparation for this called the inner game of tennis and you sort of mentioned how i mean you were really talking about tennis but it seemed like some deep parallels to robotics you sort of make the point that they talk about the fact in this book that [15:53] If you really focus with your conscious mind on improving your stroke, you sort of get to this local maximum where maybe you are doing a little bit better at hitting a forehand.
But really what you need to do to get to an optimal stroke is sort of allow the unconscious mind to take over in some respect. [16:16] How do you think about emulating this aspect of almost the unconscious and the role it plays in fluid motion and... [16:25] I don't know, proprioception and all these sort of things when it comes to robotics. Is that were you thinking about robotics when you were reading that book? [16:33] A lot. Yeah, definitely I was. I don't think it just applies to Notion. I think it's, there's a big difference between how we learn things and how we think we learn things or how we teach things to others.
You know, even if you look at something like learning a language. [16:48] the way we thought you're supposed to teach it or the way we would want to teach it to others in school, [16:54] is by defining all the rules and, you know, you first learn about grammar and all the different concepts and how tenses work and what word follows what other word. So you learn all of that structure so that the thinking is that if you know all of the structure, you can follow those rules. [17:09] And then. [17:10] speak like a native speaker would.
[17:13] But that's not how we learn language. [17:16] Right? Like you're just [17:17] immersed in it, you just hear everybody [17:20] using the structure that they kind of intuitively know, and then all of a sudden you emerge with a full understanding of how language works, even though you can't really pinpoint any of these rules. [17:30] There are many native speakers that speak [17:33] you know, perfectly follow all the rules exactly, but I have no idea what these rules are. [17:38] what the underlying structure is, what the declination is, what tenses are, they have no idea of any of that.
[17:43] but they still follow them perfectly. With that book, I think that was... [17:47] that was shown, but in the motion aspect of things. [17:53] where [17:54] You could read all you want about how to play tennis, or you can try to describe where your elbow should be exactly, where you hit the forehand, or... [18:04] how to adjust your swing, but the only way to really learn it is by just doing it and kind of immersing yourself in it. And I think there is a lot of parallels to the robotics world and to the world of AI, how we thought that the machines should think.
[18:19] or how they should learn things versus how they actually learn them. You, I think in 2023, say that you had a moment where for the first time you really felt like, [18:29] you could see a bit of the future of the robotics industry and things were really clicking into place. [18:35] But, [18:36] You had really been in that field for quite a while already, it strikes me. So I wondered what the period before that moment was like. Like, was that... [18:46] quite a lonely time to be working in a space where you don't know how it's going to play out.
So maybe I can go back and tell you a little bit more about the journey. I would love that. [18:56] So you understand that kind of the feeling. After I moved to Warsaw, as I mentioned, it was very disappointing to see the state of robotics. [19:04] And for a long time it felt like no one was working on the intelligent robots that I really wanted to work on. [19:10] And I remember at the time, [19:12] I was searching for really anything, anybody who is doing something that is a little bit closer to what I imagined robots to be.
[19:21] And for a long time it felt like no one was. [19:23] Then I remember I finally found someone. I found a professor in Germany that was... [19:29] that was writing papers that seemed a little bit closer to what I would imagine robots to be. And there was this master's program in Munich, [19:38] called Robotics, Cognition and Intelligence. [19:41] And I thought if I apply for this program and get in and they don't do intelligent robots, then I'm basically lost. [19:49] The program with that title doesn't do anything that I imagined, and something is very, very wrong.
[19:55] So I applied and moved to Munich. [19:58] And I remember the first day of classes, I went to one of the very first classes and I got stuck in the metro somewhere and was super light. So I missed the entire class. [20:11] and I showed up, no one was there. [20:14] The class was over at that point. But I saw a janitor and I was looking for a bathroom, so I asked him where the bathroom is. He pointed me to the third floor of the building. I remember it very vividly.
And as I went upstairs... [20:26] I remember seeing these big, big doors and there was some mechanical sound coming from behind them. [20:33] I don't know why, but I just decided to open them and see what's going on there. The sound was very intriguing. [20:39] And as I open these doors, I see these two basically humanoid robots. [20:44] doing, one of them was making popcorn and the other one I believe was spreading butter on a toast. And that was the moment where, you know, this is something I've been searching for at this point for years.
[20:56] So that was the moment that [20:58] That was probably one of the most powerful moments I had during that journey where I finally found it. [21:03] There's robots that look intelligent, that do something that is very human-like. [21:07] or people who work on this. So I immediately [21:11] left the room, because there's basically no one there, and I knocked at the room next door and asked if I could work there. [21:19] And basically I was completely unqualified. [21:22] I had no idea how to program these robots, I was not familiar with any of the tools, but the person who was there was not one of those people that gave me a chance.
[21:31] and asked me to just come in on Monday. [21:33] and see what that could do. [21:35] Wow. I got involved, got a job at that lab. [21:39] That's how I got into finally working on intelligent robots. So, yeah, I think this is yet another person that I'm really grateful to. And then afterwards, I moved to the US. [21:50] got into a PhD program studying intelligent robots. And at that point, I was so stoked that I finally get to do this, and I finally found a small group of people that work on those things.
[22:01] It didn't matter that it doesn't work that well. I was just so excited that I finally get to meet the people that are working on this. I can learn from them. I can... [22:10] I can push it forward and [22:12] and I can finally work on the thing that I always wanted to work on. So I did this throughout my PhD, and I had another moment like this during my PhD, this kind of defining moment where... [22:23] I was working on a certain sort of problems where it was referred to as active perception.
And I can go a little bit deeper into this if you're interested. [22:33] But, um, [22:35] At that point, it kind of felt like there are ways to write papers, to progress in your PhD, but there isn't anything that felt like it could actually solve this problem. [22:46] It all felt like the world is too complex. [22:49] You can push it. [22:50] to some extent this technology, but you won't really get to the finish line. [22:56] They're not kind of adding up. You don't really see a path out of it.
[23:00] And there was another moment where a postdoc candidate stopped by our lab. [23:06] And at that time, when a postdoc candidate stops by, you are supposed to, as a PhD student, show them your work and kind of [23:13] with them as part of the interview. So I showed to that post that kind of that stuff that I was working on. [23:20] and he listened to all of it, and then at the end of it he said that I should drop all of this and switch to deep learning. [23:27] That's the way to solve the problem that I was actually working on.
At that point, I just thought to myself that he had no idea what he's talking about, and he probably didn't listen to me. [23:34] Because at that point, I think I spent already two years or something like this working on one of the topics I was working on. But then later that day, I went to his lecture where [23:44] he showed what he was working on. [23:47] And that was another one of those moments where... [23:50] It was extremely eye-opening because for the first time I saw something that could actually work.
[23:56] where [23:57] it was less of like a, you know, here's just a little paper over there and a little paper over there, but they don't add up. [24:03] but it's something that kind of brought all the pieces together. [24:06] And yeah, that was another one of those moments, like the similar one to the one I've experienced in Munich, where at that point I decided to drop my PhD topic. [24:17] Change it completely. [24:18] start collaborating with that person. [24:21] and do everything I can to push deep learning in robotics.
[24:25] And that person was Sergey Levin, who is now my co-founder at Physical Intelligence, and one of the pioneers of deep learning in robotics. And yeah, I never looked back after that. [24:36] But. [24:37] This is a long way of answering your question, but [24:40] It didn't feel lonely in that [24:42] You know, I was just really excited to be working on that problem. [24:46] That's such good stories there. The first one is almost like, you know, um, [24:51] when a wizard is accepted to Hogwarts or something. It's like he finally found your collection of magical people and the magic that they were doing.
The Sergei story is also so interesting to me because you seem to have gone from skeptical to convinced [25:07] very, very fast. Like, what was the cycle on that? Was that, like, literally the same day, the same week? I think it was within that lecture. Wow. I think, like... [25:16] in the middle of that talk, I was like, yep, this makes sense. [25:20] this is 100% it, whatever I was doing is wrong, this is the right way. [25:25] I want to do everything I can to work on this topic. [25:29] And, [25:29] I had feelings like this.
I think every researcher has a feeling like this every now and then. [25:36] where they just see something and it clicks. [25:38] And at that point, it's a very powerful feeling. You kind of want to drop everything you've been doing and you just feel like you've [25:44] you found something that [25:46] That is much closer to truth than what you thought before. [25:49] I think we all had a similar moment afterwards when we were working on combining large language models with robot learning. [25:59] And that was, I think, one more of these, another of these powerful moments where things start to click.
[26:05] That was the Taylor Swift demo? That's right. Yeah, maybe you could tell that story because that is a fascinating one. Yeah, so what happened afterwards is I started working with Sergey. There was a few of us who were, [26:16] really getting into that field. It kind of felt like we were... [26:22] as a small group of renegades within the robotics community, because there were a lot of problems with deep learning at the time. And some of these problems remained. Like, it was not interpretable. There was no way of really making it modular.
[26:35] It seemed like nobody fully understood how it worked. It wasn't very sample efficient. And there were all of these problems. So like at that time, it was extremely unpopular to write any deep learning in robotics paper. And it was kind of like between these two worlds of machine learning world. [26:49] were. [26:49] Any robotics paper seemed kind of like the... [26:52] wrong paper for that venue. The robotics world that never fully embraced or didn't embrace deep learning at that time. But it was still really cool to work with a small group of people and push these methods forward.
[27:06] Um, so then afterwards, [27:08] the only place that was really embracing that view was Google Brain. So I decided to join Google Brain right out of my PhD, continue working with Sergey and others. [27:18] on those set of methods with the basic premise being that the way to really get it to work is to scale it up. [27:26] So we're scaling it up, we're trying to figure out how to really [27:29] get it to work at scale. But again it started feeling like [27:33] There is so much complexity in the world.
[27:36] that if robots need to... the only way for robots to learn all of this complexity is to experience all of it firsthand, [27:43] It will be very difficult to scale. It will be very difficult to have them learn about logic and how to break down a task if they had to do it all by themselves. Yes. So then I think around 2022... [27:56] I would need to look back to see what year it was exactly. [28:00] we started combining it together with Brian Liechter, my other co-founder, we started combining these robotic methods, [28:07] with large language models.
[28:09] This was before the chat GPT moment. This is where people just started experimenting with large language models and kind of understanding them better and better. There was this one particular demo. So we were really excited about this because there was one. We thought that this would be a path [28:25] to bringing in a lot of prior knowledge that robots didn't experience first hand that we learned from the internet, [28:31] into the robotics world. And that would kind of solve this problem of you having to collect so much data to understand how the world works.
[28:39] because that understanding is already embedded in large language models. So if you figure out a way to combine these two, [28:45] That should really, really help. [28:46] And we had this one particular experiment. [28:50] were, um, [28:51] we [28:52] We were testing how much transfer do you get from this internet-scale knowledge to robotic behavior. [28:58] And usually when you work on robots, you work on this very specific task, and then you test that specific task. So you're very rarely surprised. It's mostly disappointing because you really want this task to work.
You've been working on this task for a very long time. [29:13] And then it usually doesn't work, but if you work hard enough, eventually you get it to work. But very rarely you are in a situation where, [29:19] you didn't expect the task to work or like it's not the task that you're working on and it works so very early you're positively surprised [29:26] So we set up a few of these experiments trying to test how much knowledge transfers. [29:31] And one of them was this experiment with a Coke can in front of a robot and three pictures of different celebrities in front of it.
And the prompt we gave to the model we built that combined LLM and robot models was: "Put a Coke can on picture of Taylor Swift." [29:46] And one of the pictures was Taylor Swift. And you can see this video on the internet. It's a pretty pathetic demonstration of what robots could do at the time. [29:54] But the robot picked it up. [29:56] and then slowly move it towards Taylor Swift. [30:00] And that was another one of these moments of huge, huge excitement, even though if you watch the video, it's [30:05] totally unimpressive.
[30:07] Because this was the robot models had never had the chance, had never had any of Taylor Swift in their data. It had to understand the concept of Taylor Swift connected to the image of Taylor Swift and then connected to the right motion that would move Coke and to the picture of Taylor Swift. [30:24] all from internet data. So that was the moment where it clicked for us that it actually worked. [30:29] where you can bring in a lot of prior knowledge from [30:32] LLMs from the internet and connected to robot motions.
And even though the demonstration itself was very unimpressive. [30:39] just the fact that you can combine these two knowledge sources. [30:42] That was really, really... [30:44] really impressive for us. And it felt like it opened another door. Yeah, it feels like, I don't know if you're familiar. I'm sure you are. I can't believe I'm asking if you're familiar. You know much more about this than I do. But there was some experiment in like the 70s, I think, called like Shardloo or something like that, where the entire approach was to try and teach AI all the rules of the world programmatically.
And to say, you know, [31:14] And then here's the category of a bat. It has some overlapping character, you know, all of these sorts of things. But essentially, you know, people are familiar with this concept now because of the way we use LLMs. But it was almost as if the robot suddenly inherited all of these rules and, you know, pieces of knowledge from... [31:31] tying them up with the LLMs that allows the Taylor Swift demo to happen. [31:36] Yeah, that's exactly right. And I think there's maybe two points there. One is that we made the same mistake in robotics.
We wanted to write all of these rules [31:45] We thought that if only we had enough of these rules, the robots would be able to follow them and do the right thing. [31:51] But kind of like we said about the inner game of tennis, you know, you can't just [31:54] write all of the rules. You kind of have to [31:57] Do it. [31:58] And there is some underlying structure, but you can't just fully put your finger on it of what it is. [32:03] You need to learn it. [32:05] from data.
[32:06] And I think people thought about it similarly in language as well. They thought that, you know, if only we could write all of the rules, that would be enough. But it turns out that [32:15] There is trillions or billions of these rules, and sometimes we can't fully even express them in language. [32:21] You just need to learn them. And if you do, [32:25] then you would be able to follow them, even though you still don't understand what they are. [32:30] So I think we're learning this lesson over and over again.
And so, you know, you have this... [32:35] almost third eureka moment for yourself where the technology has taken yet another jump, [32:41] By that point, it sounds like you have... [32:43] most of the physical intelligence co-founders around the table, but how did you sort of pull the last few members aboard and decide to make that leap? [32:54] Yeah, I think at that point it started becoming clear that [32:59] That could be possible. [33:01] and [33:02] You know, if you worked on something for so long, at that point, that was basically my entire adult life.
[33:07] 15 years or so. And for a long time you thought that there was no solution to this problem. There were some times where it felt a little bit more [33:15] tangible but it never felt like it could actually be solved. [33:18] And then you have this moment where for the first time you see the light at the end of the tunnel. Like maybe, just maybe, if we do everything right, [33:25] if we combine it with internet knowledge, if we scale it up, if we do all of the pieces that need to be done, [33:31] it might work.
If you see that light at the end of the tunnel, you can't unsee it. [33:36] You want to do something about it. [33:38] And at that point it became clear that the way to accomplish this is to create an organization whose sole purpose is to solve physical intelligence, is to solve this problem. It can be solved as like priority number 20. [33:52] in another organization. [33:54] The only reason to exist for this organization would be to solve physical intelligence. [33:59] So then the next question is how do we actually do it and what are them?
[34:04] what are the conditions to make it happen. [34:07] And what became clear immediately is [34:10] You need to have truly the best people, truly the best researchers to make it happen. [34:15] Because the second best team in those areas doesn't really work with a company like that. You need to have the top top people in the field. [34:23] Um, [34:24] So that one actually wasn't that hard because we were already working with each other for a very long time. So I happened to be friends with the best people in the field already.
[34:33] So it was just a matter of making sure that we all want to do it. [34:36] and we're ready for it. [34:38] The other condition was that you need to get access to a lot of funding. It would require very long-term bets. [34:47] and investors that are fully aligned with this taking some time and this starting as a research company. [34:53] and not being oriented around revenue or around the short-term revenue. So that was the second requirement. [34:59] And the third requirement was just like build a very, very incredible, build an incredible company with the right focus, with the right people.
[35:07] with the expertise in all the other areas that we didn't have expertise in, like hardware operations and things like this. I spend most of my time figuring out these two points. How can we find the right investors that can... [35:19] that can really help us in this adventure and [35:24] they're fully aligned with how we want to do this. And then how can we find the right people to [35:30] to fill all the missing pieces. [35:32] How did Lockie Groom join the crew? Because... [35:36] He's obviously had a very impressive career as an operator, as an investor, but isn't someone who [35:41] One would automatically think, you know, this is someone who's going to devote the next chapter of their life to robotics.
[35:47] So I didn't know Lockheed before we started thinking about starting a company. [35:52] But then, and he should tell this story rather than me, but I believe what was happening is that he's been investing in that point for a few years. [36:02] And as long as he's been [36:05] he always thought that [36:07] He doesn't want to be investing forever. He wants to find something else that he can fully devote himself to. [36:12] And he wants to build things. [36:14] And the one area he was always fascinated by was robotics.
And he was seeing over the past few years that there was [36:21] There's a lot of innovations happening there. There was some kind of moment. Robotics was having its moment. [36:27] And particularly, he was impressed. He was seeing over and over papers coming from Chelsea, Finn's and Sergey Levin's lab. [36:35] as well as our papers from Google. So he talked to a lot of his friends that if ever Chelsea, Sergey or Carl or anybody from those teams are thinking of starting a company, please connect me to them. [36:49] I would at least want to invest or at least talk to them.
[36:51] So that's when we got introduced by a friend, and we pitched to Lockheed as a way to have him invest. And at the end of that pitch, it became clear that, [37:03] He would like to do much more than just invest. This is the opportunity he's been looking for for all the years he's been investing. [37:09] and she really wanted to dive in. [37:12] Um, so then the next step was. [37:14] us trying to get to know each other as quickly as possible, spend as much time together as we could.
[37:20] And then as we were doing that, it also became clear that this is the person. [37:25] that I've been looking for that could help us on the commercial side, on the fundraising side, on the operations side. Kind of one of the missing links that we really needed. [37:35] Were there certain... [37:37] philosophical alignments that you needed to find in someone who you hadn't worked with before? What were the sort of [37:44] you know, even beyond maybe the tactical pieces, the, the traits that you, you saw in him that made you feel like, yes, this is someone I can, can bring into this very trusted group where, you know, we've already built all this history together.
[37:57] The ref you, um, [37:59] I think maybe the biggest one was that he immediately got it. [38:04] At that point, there were a lot of investors or other people that I talked to, other business people that I talked to. [38:10] And it was very hard to get that idea across, that idea that, [38:14] You have to do research right. You need to get... [38:18] You need to build the technology first and you can't be distracted by short-term revenue. [38:22] And if we do do this right, this is going to completely change the world and it's going to be [38:27] most valuable business of all time, but you need to, you need to have the patience to let it [38:32] let us do it the right way, rather than short-circuit that path and [38:37] I'm gonna cap the ceiling.
[38:39] And he got it, I think, within like the first minute. [38:42] So that was very reassuring. And then on top of that, [38:46] What? [38:47] What I really liked about our conversations was that [38:50] In many of those previous conversations with other people, I always felt that I'm the one pushing the ambition of this project, of how far it could go. [38:59] if you set it up right [39:01] And I think this was the first time where I had somebody else [39:05] set up even higher ambitions for us. That was really cool to see because it felt like we're all pushing in the same direction.
He will make us even better and more ambitious. [39:14] So, [39:15] I was really glad to find someone who will [39:18] you know, make sure that we are not gonna [39:21] We're not gonna short circuit this, he's fully aligned in doing this in the biggest way possible. [39:26] He makes us more ambitious. [39:28] And he's the person in the areas that we really need more expertise. [39:32] and he'll make sure that those areas are fully aligned with how we want to [39:36] develop this company. So, [39:38] It was basically a no-brainer at that point.
You had this... [39:42] insight that this was the right moment to to go and solve this problem. [39:47] But, [39:48] at least from the outside, one could imagine many different form factors that might take. Like another version of this company might be saying, hey, we're going to try and build the best humanoid robot ourselves, or we're going to take this approach to getting the data we need to run these models very effectively versus another approach. How did you sort of land on the version of physical intelligence as it is today?
[40:13] I think from the get-go we had an... [40:16] we had an idea of, we had the thesis for the company. And that thesis was around all the research that we've done up until this point. [40:24] that similarly to what happened with language, [40:27] It's not going to be specialist models that are really going to solve this problem. [40:32] is going to be generalist models that work across. [40:35] all kinds of different tasks, all kinds of different environments, and all kinds of different robots. [40:39] And it was similar to what we've seen in language where, you know, you would think that the best translator would be just specialized for translation or the best translator.
[40:49] coder would be just specialized or just trained on code. [40:52] But it turned out that the way to be the best at all of those fields is to train one generalist and, you know, generalist that takes in poetry data and coding data and translation data. And then it turns out to be much better than all of those specialists at those specialist tasks. [41:08] And we started seeing something similar. [41:10] in Robot Learning, in Robotics. [41:12] that if only we could collect enough data, if that data was very, very diverse and if we do it the right way, [41:19] we should be able to build the best generalists.
And if we solve that, that is really going to allow us to have this world of many neighbors, form factors and robots that can be very intelligent. So I think the kind of the two thoughts that we had from the beginning is that one is, [41:36] Intelligence has always been the bottleneck for robotics. But rather than trying to start a robotics company that focuses on a specific robot, [41:44] How can we tackle this problem head on and just focus on the intelligence? And then the second thought being that, [41:50] The way to solve intelligence is to take the lessons that we learn from vision and language and other fields.
[41:57] and really take the foundation model approach, which includes things like cross embodiment learning, large diversity of data, a lot of real world data, and do the research necessary to figure out how to build these models. [42:10] So you sort of land on this idea of, you know, the AI brain for these different robot types, to put it sort of simply. How did you start to think about the right way to gather the data? Because that's clearly such a big piece of it. And obviously there are different players that take
[42:27] different approaches there. So I imagine you must have had to reason through that in a million different ways to land on the one you have. [42:35] I think we're still reasoning through it. I don't think we have all the answers yet. [42:40] I think the important piece about how [42:43] how to think about physical intelligence is [42:45] We are not very dogmatic. [42:47] It's not that we sit down and think very, very hard and then come up with the solution and this is our bet. [42:54] I think the better way to think about it is that we're really...
[42:57] truth seeking and we run experiments. [43:00] and we don't know the solution and [43:03] And we know that we don't know the solution, so we want to follow the scientific method and really try to find [43:08] the truth and really try to find what works and what doesn't. Because I think there is a true answer out there. We just need to be very humble. [43:17] in finding it. So that's I think how we arrive at the current sort of answers. We run a lot of experiments, we try many different ideas.
[43:24] And we see which ones stick. And then we double down on them. Now in terms of [43:30] based on what we've seen, based on the evidence we've seen, how we think it's going to work or how I think it's going to work. [43:36] these models will need to be able to absorb very diverse data sets, where it's less about picking the right data or picking the right way of collecting data, [43:47] More about building the engine that allows you to absorb all kinds of data, literally. [43:52] videos of people, whether it's teleoperation data from robots, data from handheld devices,
[43:58] um video data really or simulation data really anything and the more data they can absorb [44:04] the better the model will be. And I think we're kind of at this stage of robot learning where we [44:10] kind of try to throw anything we can at these models, build them in a way that they can absorb as much of it as possible. [44:17] and get them to the threshold of being able to be being deployable. [44:21] where you can actually deploy them in the world and [44:24] have robots out there collecting data for real doing economically valuable tasks.
[44:29] And I think once you're at that stage, once you're at that threshold of [44:32] they can actually work and do valuable tasks. [44:36] That's when you enter this next stage of now deploying robots at scale, and deploying it in multiple verticals and diverse environments and actually delivering value. [44:46] And I think that second stage is actually going to be the stage where we get most data from. [44:51] where the more robots you deploy, the better the models should get, the more you can deploy them and there's a natural flywheel.
[44:57] to it. And what's exciting about the moment today, this moment right now, is that I believe we're very close to this threshold. [45:05] already above that threshold. [45:07] And that's something that is really, really exciting. Because I think the months will keep on getting better, they will be able to absorb more and more data, but they will also have this sustainable source of data that is... [45:17] very, very valuable because that's the data that is the most real, that is the closest to how you actually want to deploy the robots.
For folks that maybe haven't spent as much time [45:25] digging into this or sort of coming to it fresh. [45:28] I would say that one of the experiments, as you think about this data collection, that physical intelligence has done maybe more than others, is this real-world data approach. Why is that so important and so valuable to get and so valuable to enter that phase two that you talked about that can sort of start the flywheel? We have a few theories why this is really important, but I think maybe the meta point here is that [45:54] If there was a different path that worked much better, like maybe for simulation or from...
[45:59] through videos or something else, we would happily pick that path. [46:03] So again, it's not that we sat down and thought that this is the best path and this is the only thing we're going to do. [46:09] We run a lot of experiments, we tried out a lot of ideas, and this is the one that seems to be working very, very well. [46:15] Why do we think that this is important? [46:18] Um... [46:19] I think it's kind of difficult to describe it in the absolute. It's, I think, a little bit easier to compare it to other alternatives.
[46:27] So one popular alternative is, um, [46:29] assimilation. [46:31] And this is in particular an area that had a lot of success, especially in locomotion use cases, where you have robots walking around or doing stunts like backflips and things like this. Most of these methods are training simulation first. [46:44] And I think for those kind of tasks, the main complexity is about how you move your own body. [46:50] It's less about interacting with the world and more about [46:52] How do I move my legs correctly so that I can walk or I can run?
[46:57] Um. [46:58] And in that sense, as long as you model your own body accurately, [47:03] You should be good. [47:04] If you model it very very well, your own particular robot, [47:07] and that transfers from simulation to the real world, you should be able to learn that behavior in simulation, and then that's good enough to them to work in the real world. [47:17] Now, when it comes to the problem of manipulation, where you're manipulating the world around you, the difficulty is less about modeling your own body, [47:25] how you move your arm from A to B.
But it's more about... [47:31] modeling how the world will react to it. [47:33] So the problem of manipulation is more about this interaction with the world that you have. [47:37] or the interaction of the world that you're interacting with. And I think in this case, it's just much harder to simulate the world around you than it is to simulate your own body. It's no longer about just a single robot that you need to simulate. You need to simulate everything. [47:51] and [47:52] We don't really know how to simulate everything at scale, how to do this accurately enough.
[47:57] and scalably enough, [47:59] so that it would work [48:01] for anything, for every single one of tasks or objects. It takes really long to get it exactly right, to get all the friction parameters right, to get the simulation behaviors right. And it's just not as scalable. That's been our finding so far. [48:14] And so it's sort of the case to try and distill that [48:18] That, you know, with simulation, you can do some of these... [48:21] I don't know, maybe this isn't the perfect word, but sort of coarser, larger actions that are more self-contained.
But once you start trying to implement picking up the coffee cup or the towel or whatever it might be, you're starting to rely on a simulation that would have to be so good you'd effectively have to create a simulation that is perfect. [48:43] perfectly similar to reality and that's where the the real world data starts to become so important [48:50] Yeah, you wouldn't have to simulate all of the external world. Yeah. Any one of those tasks. [48:55] And that's just too costly, too difficult to do. [48:59] But as long as you can get away with just simulating your own body, [49:02] I think that works perfectly fine.
And that's what we've seen in locomotion or backflips or dances or things like that. [49:08] There was a... [49:10] A moment, I think about a year and a little bit ago where you talked about... [49:14] reinforcement learning making a comeback. And that has since become a really interesting part of the way that physical intelligence [49:22] seems to work with recap. What were you seeing at the time that made you think, you know, uh, this, this technique might have sort of a second or, you know, additional wind here. And why has that been so useful for you?
[49:37] Yeah, there's been a long history of applying reinforcement learning to robots. [49:41] And the problem of reinforcement learning is such that [49:44] you need to be learning from your own experience. [49:48] And, [49:49] while gathering that experience you need to encounter some successes and then [49:54] You want to increase the probability of good actions that lead to those successes and decrease probability of the actions that lead to failures. [50:01] What this means is that you need to have [50:04] As you explore the world, as you collect your own experiences, you need to have some of these successes.
Because if you don't see any of them, you basically don't really know where to go. You're kind of lost. And if you start from scratch, if you just try to command random commands to a robot,
Want to learn more?
Ask about this episode