The Token Budget Problem Nobody Is Talking About (Matan Grinberg, Co-Founder & CEO of Factory)
Matan Grinberg is the co-founder and CEO of Factory , an AI company valued at $1.5 billion that helps enterprises like Nvidia, Morgan Stanley, and Adobe automate software development through “Droids,” intelligent agents designed to streamline software engineering. Before Factory, Matan spent more than a decade in theoretical physics, studying string theory at Princeton and UC Berkeley. His work now centers on a different kind of complex system: how software gets built in an era of increasingly capable AI agents, open models, and shifting compute economics.
Appears in
- Uploaded
- Uploaded Jun 30, 2026
- File type
- POD
- Queried
- 0
Full transcript
Showing the full transcript for this episode.
[00:00] The biggest threat to OpenAI, Anthropic, or any of the model apps is not each other, it's open models. Because if open models are really good, then it really puts a lot of pricing pressure on them. So their interest is it's just, you know, the big four and no one else, no one can compete. And so the best way to do that is say, you know, it's either us or the Chinese model. [00:17] The future of software engineering is you're not building the software, you're building the factory that builds the software.
It's crazy that we've lived in a world where engineers who are some of the smartest people spend years becoming experts in their craft, spend hours and hours and hours fixing CI. Yes. It is not a good use of their intelligence. And so we're figuring out ways to automate the low leverage use of their time so that those key moments where it's like their deep insight is essential, they can do that more. What do you think looks different about AI in a year's time? And what do you think looks different about factory?
[00:47] I think AI a year from now, there will be less focus on... [00:58] In 2023, Matan Grimberg, then a Berkeley physics PhD, [01:03] took a 30-minute meeting with the venture capitalist Sean Maguire. [01:07] It became a three-hour walk. [01:10] At the end of which, McGuire issued a dare. [01:13] drop out, [01:14] and start a company. [01:15] Three years later, [01:16] and Matan's company, Factory, is valued at $1.5 billion. [01:22] Factory helps enterprises like NVIDIA, Morgan Stanley, and Adobe automate software development through the use of droids. [01:29] AI agents capable of chewing through the endless drudgery that consumes so much of developers' time.
[01:36] factories growing rapidly. [01:38] consistently doubling revenue month over month, [01:40] and growing from 25 people to over 100. [01:43] since just the start of the year. Today, Matan and I discuss his decade in string theory, [01:48] The existential crisis that drove him to entrepreneurship [01:52] AI's brought power dynamics, [01:55] and our mutual love for the film's [01:57] of Paolo Sorrentino. [02:00] Matan also shares the unusual ways he runs his company, [02:04] and why he believes both the AI space and his business will look very differently. [02:09] in 12 months time.
[02:11] I'm Mario. [02:12] and this is The Generalist. [02:14] - This episode is brought to you tech domains. I spend a lot of time speaking with founders and builders who are building the next generation of technology companies. For all of them, [02:24] Finding a compelling and distinctive company identity [02:27] is essential to breaking through the noise. [02:30] That starts with a great name and a great domain. That's exactly the thinking behind dot tech domains. [02:36] for companies building in tech, [02:37] tech domain gives your project a clean, confident identity from day one, instantly communicating what you're building.
tech, tech, tech, tech. The list of companies tech is growing quickly. [02:53] It's not surprising that many venture-backed startups secure tech domain early. If you're building a technology company, it's worth thinking about how you want to show up from the start. [03:03] Secure tech domain today from any registrar of your choice. [03:08] The best founders aren't spending their time on expense reports. They're busy building. Brex is the agentic finance platform that makes that possible. High limit corporate cards, banking, and AI that handles the back office automatically so your team never has to.
[03:23] Expenses get captured, books get closed, and spend stays in policy without anyone chasing it down. [03:30] your team gets their time back to focus on what actually moves the company forward. [03:35] Vercel, Opening Eye, Anthropic, Granola, and Deepgram. [03:39] already run on Brax. [03:41] One in three startups in the US does too. [03:43] It's time to get bricks. [03:46] Go to com slash solutions slash startups. [03:51] Matan has been... [03:53] A long time that I've wanted to have this conversation and a long time that we've actually known each other.
So I'm excited to have you here. Thank you for having me. It's a pleasure to be here. You know, I've been thinking about [04:02] how we should start this conversation because it's such a busy week in AI when we're talking. [04:09] But you have such an unusual background that I think it would almost be a waste to start too much on the AI side. There are not... [04:16] many string theorists to CEOs that I've come across. [04:21] So let's start somewhere different. [04:23] eminerta [04:24] Her... [04:26] First law, [04:27] is a [04:30] physical law that you apparently are [04:33] a big fan of.
[04:35] You can see that already I'm off my footing here. Tell me what that means to you. Yeah, so Emmy Neuter, that was a great German pronunciation. Thank you. I had the look of the phonetics. Oh, it's good. I mean, she's a legend, and I think every good physics undergrad learns how I call it Neuter's theorem. Basically says, with every continuous symmetry that exists in the universe, there is a corresponding conserved quantity. [05:03] Now, there's something so beautiful about it because... [05:06] Yep. [05:06] Let's go abstract, dumb it down. Basically, if things are symmetrical in some way, then there's something in the universe that is conserved accordingly, like something that we know.
And let's put some exact examples. Like, suppose we're in empty space. [05:20] If you are... [05:22] you know, [05:23] at one point in and one point in space or another point in space, or if you're moving versus not moving, these are things that you can't really tell the difference. [05:31] If you're in an empty space or if time is translated, like if an hour passes and you're in an empty space. [05:36] There's no difference of empty space, it all looks the same. [05:39] And so naively, it's just, you know, okay, these are things that are symmetrical.
The universe doesn't change. [05:44] based on [05:46] change in position of space or in time, [05:49] But it turns out that [05:51] secretly behind this we get [05:53] conservation of energy. [05:55] and conservation of momentum, [05:57] because of this. And it's like, you know, maybe the mathematics of it tells a more beautiful story than I can, but there's just something so fundamental about, and so beautiful about [06:08] these very simple things that like a three-year-old would understand or a five-year-old would understand like if you just imagine you know you're in a black abyss [06:16] and an hour goes by, nothing is going to change about your world.
[06:18] But that is like [06:20] conservation of energy. That is the source of where that comes from. The mathematics makes it even more fascinating because... [06:26] in kind of different toy universes that you can build, you can create different conserved quantities, and it leads to these different... [06:33] theories of physics, but yeah. Do you remember when you sort of like encountered this as an idea, as a theory? I think the first time was probably like junior year in high school at some point. Okay, wow. Yeah. What is it about it that appeals to you so much?
Are there sort of connections you make to it? [06:50] today to it? You know, people in Silicon Valley always talk about first principles. [06:54] and first principles thinking and all that. That is an example of like in physics, one of the most like, [06:59] you know, basic first principles. It's like you always look at what are things that are conserved? [07:04] What are things that are symmetries where if this thing changes, it doesn't matter? [07:08] right and obviously like you know the first principles thinking leads to [07:13] a lot of good kind of mantras or decision making ideas in you know business or startups or um tech and i think what really appealed to me was it's kind of this like base that you can always no matter what complicated problem you're dealing with in physics [07:27] you can always think, okay, what are the things that are conserved?
Or what are the things that are symmetries here that I can then find something that's conserved? And that kind of gives you a baseline upon which you can stand and then look at [07:36] some of the more complex things that are not symmetrical or not conserved, but it gives you like a basis to stand on. [07:42] And there's something so deeply satisfying because you could break down a lot of problems into [07:47] what symmetries are conserved and which ones are violated. And then when you separate out, okay, these things are conserved, but then, so for example, suppose we were in an empty space [07:55] and there's a sun.
[07:56] nearby. [07:57] Now it is no longer a symmetry where if I move forward or backward, [08:02] uh nothing changes because now if i you know [08:05] go light years away i'm affected less by the gravity of this or i even have a reference point of how it's going to interact with me and so now suddenly there's a new dynamic and there's a [08:15] a symmetry that's violated. However, one that remains conservative is there's a son or whatever object there and I wait an hour nothing changes. [08:22] So, you know, we still have that conserved.
I'm dwelling on it in a way because you are operating in like the most changeable part of the economy at the moment. [08:32] ai and then within that code generation which is the absolute like heat center of all of this and so [08:39] The idea that you've really trained your mind on what is conserved in these moments and thinking through these things, like I almost... [08:45] Maybe I'm over deriving it, but like I see some of that in the strategy that you guys have been. I'm so glad you mentioned that.
I actually think it is so important because there are a lot of people that are going out there and saying things like all jobs are going away or we're going to automate everything or, you know, this language about like the permanent underclass or things like this. [09:04] And it really upsets me because not only is it inaccurate, but it's really harmful to like the psychology of a lot of people that I think have a very important role to play in the future of humanity. [09:14] And the reason you can actually break it down to some of these things that are kind of conserved or symmetries or however you might define it.
And so. [09:21] One thing, [09:22] I think we can say is, first of all, there are a lot of problems in the world. [09:26] I think we could all agree. There is a lot of problems, many of which we are not solving. Okay, so let's take this group of problems that we are currently not solving. I would argue that a very large percent of them could be solved with software, but they are not currently being solved with software. [09:40] Okay, so we know that there is a huge amount of problems, many of which could be solved with software, but are not being solved with software.
And we'd also probably agree that the number of problems that exist generally tends to grow over time, not shrink. As every new technology comes about, or even as, you know, the quality of life for every human grows over time, so too does the problems that they have and the things that they expect and the things they want to solve. Do we think that's true? Do we think that the problems multiply? Or doesn't the fact that [10:06] The quality of life improves so much suggests that we're somehow [10:09] getting rid of a lot of problems along the way?
I think we kind of climb the ladder of what type of problem. So I think right now, mental health is a very important problem. [10:18] I don't think that existed in like hunter gatherer times. [10:21] Yes, probably not. You didn't have to worry about it. It was just like, where is the boar so I can, you know, hunt it and bring food for my family or something. So maybe there wasn't kind of. [10:29] that it wasn't something that emerged at the time. We don't have robust data on this. We don't have data.
I mean, I think that's a fair point. But I think even putting that aside of if the number of problems grows or it's static or it shrinks, [10:42] What's clear is there are way more problems that are currently being solved. [10:45] And the tools that we are creating, and in particular for factory, for software engineers, these give every engineer more leverage. [10:52] And so, [10:53] That means that they kind of have more power to solve any given problem. Now, [10:57] it might mean that certain problems need fewer people to solve it.
[11:01] And so if you pick one particular problem, you could say, OK, people are being replaced because they're not, you know, we used to need a thousand engineers to solve this problem. Now we only need a hundred. [11:11] And so on a local... [11:12] basis that could be true on a per problem accounting but [11:17] there are so many other problems that we are not solving that we can solve and [11:22] I think the kind of beauty of what is happening here is by giving everyone more leverage, we will solve more of those problems.
And that is a net good for the world because there are so many problems that need to be solved. [11:32] There's so many problems that also need to be solved with great software, not shitty software. Like, [11:37] Even something like the DMV. [11:39] Why do we live in a world where the DMV has the most atrocious software? That doesn't need to be true. [11:44] Like it has been true because of, you know, funding reasons or maybe it doesn't attract the best talent, but it doesn't. We don't need to live like this.
Like we can live in a world where the DMV is nice. And I'm very excited for that world. I think I think we will live in that world on the topic of the job since you bring it up. It sounds like your impression is. [12:03] or your view would be that there is a huge amount of displacement [12:08] from these positions where because of, you know, [12:11] Frontier models and, you know, OpenWeights models being so capable at this point that [12:16] Plenty of companies get more leverage. [12:18] They cut down their teams, but that there's sort of such an abundance of problems that humans are still very, very relevant to addressing them.
Is that a fair encapsulation? Yes, I think that's right. I think also some people are experimenting because there's going to be more dynamics that emerge, which is pick any random vertical. You might say, great, now we can get the same done with fewer people. [12:37] but [12:38] Generally, the free market has competition for a reason. And so maybe you have a competitor who doesn't do that and doesn't lower their staff, and they can actually give a much better experience to the consumers. [12:48] And then they might end up winning. And whoever ended up [12:50] taking that like preemptive layoff or whatever it might have been, maybe they end up losing a huge amount of market share.
[12:56] And so there are probably going to be verticals where [13:00] you know, actually writing software wasn't a core competency. It was just a kind of a thing that had to be done, in which case maybe they will have fewer people doing software engineering. But I think there will be other verticals where actually every software engineer you have gives you more leverage and every tool that you give them gives them more leverage. And so that helps the business. [13:18] even more and so you might not want to stop hiring at all.
I think it just depends kind of vertical by vertical. [13:24] It sounds like that version of the future, which I could very well see being the case, is predicated on like some version of a jagged intelligence where humans are more capable at some things for a period of time. [13:37] And that maybe we remain like very compute constrained or models are so expensive that humans are still like on a relative basis. The the ROI on a human is good. Like, is that sort of how you think about it? Yes. And I think.
[13:50] kind of every company will have a resource allocation problem that I think is very interesting, where basically, [13:57] CIOs or CTOs or kind of CEOs, [14:01] I'll have a problem where they will have for every incremental dollar, they will have to decide, do we put that to more headcount? [14:08] Do we put that to more like data on the market or understanding our users? Or do we put that... [14:14] to more compute. And by compute, I mean like token budget per person. Yes. And so within that even, and I think this is kind of related to what we're doing is, you know, we're in the middle of, there was a crazy, you know, spike in usage of all these tools.
They didn't know what the ROI was. So they're going in and putting in user limits. [14:28] We're in a very interesting interim period where... [14:33] People are kind of preventatively putting in user limits of $1,000 of tokens. [14:38] per month for everyone in the company. There is absolutely no way, and we can timestamp this and check, there is no way in 12 months [14:45] orgs are going to do such a kind of blind allocation of tokens where everyone gets the same? Absolutely not. There is going to be a resource allocation problem of if I have an incremental token to move the needle for my business, where do I put it?
Does it go to the sales team, the marketing team, within the engineering team? Is it platform? Is it customer support? Is it billing? Yes. You're investing in some capacity. You need to get a return on it. And so we just have the most right now we have these like resource allocation things, [15:15] the dumbest way possible, which is like, hey, should we hire more people? Should we give people token party? It's like, I don't know. Yeah, let's do this. Let's do that. But it's very not data, like not data backed at all.
Yes. And I think, [15:27] um we're about to enter a world where we'll be able to have very clear metrics on actually [15:32] We want to spend more tokens in this department. Actually, tokens don't matter for this other department. We need more people. [15:37] Maybe for sales. Because sales, I don't know if tokens do that much. Sales, it's like you want the face-to-face time. You want those people who are great at connecting with other humans. And so it's going to be very interesting to see this resource allocation play out. That's super interesting.
I hadn't thought through how we will need to attribute that and how that sort of [15:54] becomes possible. Maybe to give folks a sense of where you're coming into this conversation from, how would you describe factory? [16:01] Yeah, so it's funny, you know, it's kind of it's changed over time in terms of what the quick description is, but the mission has remained the same. And our mission since we started at the very beginning, since you first mentioned us in that newsletter three years ago, the mission has been to bring autonomy to software engineering.
And, you know, in the first couple of years, the focus has been the software development agents that we build, which are called droids. [16:25] focused on not just coding, but the full end to end software development lifecycle. They're kind of the fundamental unit of this new world of agent native software development. [16:35] But recently, as this is becoming more common and people are more familiar and using droids and [16:40] You know, they're orgs that have tens of thousands of users who are using this. Now the thing that becomes more relevant to discuss with them is not just the incremental, you know, droid, but rather the software factory that they build.
The visual that we have sometimes is like, [16:53] Have you ever seen videos of Tesla's factories? Where they have like these robotic arms that are going in and doing all these things? [17:00] The future of software engineering is you're not building the software, you're building the factory that builds the software. And so engineers are going in and saying, hey, actually, you know, we need to tweak this arm this way and this arm that way. So it's going to go in and attach this widget and that widget. And [17:15] It becomes much more of like a systems design.
[17:17] problem you're not locked into any one model but now you have a resource allocation problem again of in our assembly line of this software that we build we could naively use opus for everything but we probably don't need to it's probably pretty inefficient it's not good for our margins to do so so where do we you know for this robotic arm use an open model or for this one actually gemini might be really good at this language maybe we want to fine tune a model [17:41] for COBOL because it's some legacy code that, you know, the other models aren't good at.
And so this language of like software factories. [17:48] is what is kind of really top of mind for the enterprises that we work with. Basically figuring out [17:54] How do they give [17:56] every engineer more leverage. [17:58] Because it's crazy that we've lived in a world where engineers who are some of the smartest people spend years becoming experts in their craft. [18:04] spend hours and hours and hours like fixing CI or writing documentation or doing like PRDs and meetings with 20 people. [18:13] Like it is not a good use of their intelligence.
And so we're figuring out ways to automate the low leverage use of their time so that those key moments where it's like their deep insight is essential. They can do that more. Yes. And the great thing is that's what engineers like doing. So I've been droiding. I like I like using my droids. I think it's a really cool product. You've also shipped a lot. [18:34] of sort of new features it feels like the cadence at which the product is expanding is really exciting you know i'm thinking of droid computer i'm thinking of factory router there was even another one that i'm i'm oh automated security like how do you think about [18:47] expanding the scope of this.
[18:49] Yeah, um... [18:51] I think it's actually... [18:52] It's funny, the factory analogy is very helpful because it really does feel like that where, you know, you think about the assembly line. [19:00] The point is you want to figure out what are the kind of biggest bottlenecks. [19:04] that stop us from producing that whatever unit on the assembly line faster. And so, you know, one of the things that emerged as another bottleneck was some of our customers were saying, hey, you know, this is fantastic, but we're subject to these regulations.
Can you help us, you know, test for this in every code review and not just review for, you know, is this good code, but does this adhere to these regulations or these security standards? And so [19:26] kind of these are things that allow us to level up or droid computers yeah you know [19:30] If you're such a mature organization, such that everyone is using agents, you eventually come into a bottleneck where [19:37] You're delegating so many tasks to agents, you can't do it on your computer anymore. [19:42] Because also, what if you close your laptop, you know, then the agent's done.
And so having the ability to spin up a remote machine where the droid can go and use it in a sandbox environment, generate the code, but also test it and verify that that code is good and verify that it adheres to all these standards or, you know, spend a week [19:59] fixing CI if you have some really messy, you know, CI. By giving Droid that environment, then it is able to [20:05] tackle just so much more of that software development life cycle that engineers you know don't want to waste their time on [20:11] What's your view on...
[20:13] How many models? [20:14] people will [20:16] ultimately use at a sort of end state like are we trending towards a world in which [20:20] all of us are using literally dozens of models for very specific things and you know factory is helping us route across all of them do you expect there to be sort of i don't know a set of 10 that are really top of the line and you know optimizing these different ways like how do you think about that i think if we succeed we're [20:38] uh [20:39] our customers will know the model that they're using just as well as they know like where the electricity that powers their toaster comes yes sure which is like at the end of the day when you use a toaster you just want it to toast your toast [20:51] And I think similarly, [20:53] There are so many models, which is great.
Like more models, the better. But also it just becomes a lot of like intellectual overhead. [21:00] as an engineer to even know, are you going to use like, [21:04] Opus 4.8 for this, or GPT 5.5, ultra high, and it's 4.6, and this and that. Kimi does this. And the Kimi K 2.5, like... [21:12] It's just way too much when, you know, some businesses, it's a core competency to know their stuff, like us, for example. If you're an engineer at, [21:20] Pick a Fortune 500 company. [21:22] It is not worth your time to keep track of every single model that comes out.
And every time update your priors on, is it good at this? Is it good at that? This language, like... [21:31] You also might be doing Python sometimes, Java other times, Cobalt other times. So not only will you need to know the personalities of these models, but also how they perform in these different tasks. [21:40] Way too much. [21:41] not a good use of their time. And that's kind of what we want to take on for our customers is deal with all of that [21:46] route according to whatever you want to optimize for so whether that's cost performance um latency and then allow the engineers to focus on [21:55] their core competency, which is, you know, [21:57] moving the needle for their business, thinking about kind of the bigger picture of [22:01] their team or their org or whatever feature they own, instead of just scrolling Twitter constantly to see the latest updates of every model.
How have you found you are winning customers sort of opening the door? Is it working on... [22:13] difficult migrations I've seen you talk about? Is it, you know, a certain sort of wedge that appeals most to the enterprises that you're working with? Yeah, it's funny. I wish... [22:23] For many reasons, I wish it was like kind of one thing every time, but it actually, [22:27] ends up being a different thing depending on what the bottleneck is for that org. [22:31] So there are some orgs where, you know, they are very concerned about having vendor lock-in with just OpenAI or just Anthropic or just Google.
And they see that we're model independent and are like, okay, great. We need that because we don't want to be kind of subject to... [22:46] to the whims of one particular company. Yes. [22:49] Then there are others who are like, oh, my God, we've been working on this migration for three years. Please help us. I've heard that it works. Yes. And so that's, you know, in those cases. Then there are others where. [22:58] developers, [22:59] you know, found us on Twitter, started using it and start raising the pitchforks to their CTO say, hey, we want this.
[23:05] So it kind of varies case by case. And it's, yeah, it's just, it typically ends up being different parts of the software development lifecycle that ends up being a bottleneck for that org. [23:14] I saw, you know, after the news of sort of the provisional... [23:19] cursor acquisition from xai was announced [23:24] You were sort of having a debate on Twitter with someone about, you know, the position of the other person was basically, it makes sense that these... [23:31] model companies eventually buy cursors, you know, the harness companies, the environments, because you need to sort of co-optimize between the foundation model and, you know, this other layer.
And your take was like, there's no data for that. And it's not at all true based on our experience. Can you help me understand like what, why that's a position that you take and the data you have? Okay. So I guess there are two things here. So there's one thing, which is, [23:56] the relationship between the model providers and like coding generally. And then the second is the relationship between [24:04] agents, and the model itself. [24:07] So on the first one, I mean, I think something that I... [24:10] very much believe is that i think the cursor spacex acquisition is good for spacex [24:15] good for cursor and it's good for us and so [24:18] I guess first, the reason it's good for SpaceX or XAI is they have a ton of data centers.
They have all the infrastructure. They have good engineers. They don't have the distribution or data to make their models really, really good at coding. [24:30] cursor has that distribution and has that data but doesn't have the data centers and all the infrastructure and the like cache [24:38] to make that work. So that's great for them. They also, you know, fantastic valuation. Great for all of the teams. So huge congrats to them. [24:44] For us. [24:45] It's also helpful because there needs to be an independent kind of software development factory provider that is not tied to one model in particular.
[24:54] And, [24:55] that was the only other big one out there really that we were seeing in the market and so [25:00] it kind of cleared the way for us to some degree. And so the great thing is now, [25:06] OpenAI is going to have great models for code. [25:08] Anthropic is going to have great models for code. Yes. Google is going to have great models for code. And XAI. [25:13] As much as people don't believe it, they're going to have good models for code by the end of this year.
And that is good because the more choice we have, the better it is for the consumers or the enterprises or the people who are using these. [25:23] Now, [25:24] The thing I disagreed with was this claim that [25:27] If you make the models and the agent or the harness, you are going to make them better together, you know, naturally, right? Like, you know, an example here being Claude code and the Claude models or Codex and the Codex models. What we find empirically is that. [25:43] It's actually not the case. So by making Droid model independent, we actually end up making it perform better with Opus, better with Codex, better with any model than.
[25:54] those models do in their respective harnesses. So we can outperform Claude code with Opus by using Opus in Droid. Naively, it might not make sense. But I think a good analogy here is let's throw back five years ago. [26:06] If you wanted to make a... [26:08] AI that was good for you as a personal assistant, you can make the argument of, hey, you should train it on just your data because it's your personal assistant and you want it to know you better than anyone else. [26:18] What the last five years have showed us is actually, if you want it to be a really good personal assistant for you, you should train it on the whole internet.
Yes. And then it's going to be much better for you. And there's an analogy that emerges here where basically what data is to a model. [26:31] models are to the harness. Hmm. [26:34] Really? Yes. You find that? Yes. And so the more models you expose the harness to, the more kind of nuances and intricacies you find and you avoid overfitting. [26:43] the harness to the nuances of that model in particular. And so these are on axes like, [26:48] compaction, which is how the agent deals with [26:52] um kind of being over the context limits which is going to happen oftentimes when you're dealing with large code bases also on the axis of uh token caching so basically how often are you [27:03] Caching the tokens in a way that basically makes each query more cost effective.
Because a cached token is kind of like a repetitive token that costs a tenth as much. [27:12] So you want as much caching as possible. Yeah. We end up outperforming that. And then also like tool use and environmental feedback. By seeing the kind of intricacies of the different models, it ends up, [27:21] performing better with all of them, which is pretty crazy. And my friends who work at the model companies get very frustrated by this. Really? Yeah. I mean, I bet. You have to put this out in research, I think.
Then that would be amazing, right? Like, I feel like people will... That will really, like, change people's prior. Oh, yeah. We have a couple blog posts about, like, caching in particular or compaction in particular, but I think not the collective story of here is why you need multiple models to make it perform better. I think it would be really cool. You've also talked a lot about... [27:50] you know, the need for open weights models and the value of that. And talk to, I think, also a little bit about how you don't like it when folks sort of classify these as Chinese models because there needs to be a robust...
[28:02] US open source system. Yes. Where do you feel we are on that front at the moment? Like, are there enough [28:08] good players in that space yeah i think the reason why i don't like when people do that is because naturally [28:12] You know, if you're a... [28:14] Like the biggest threat to OpenAI, Anthropic or any of the model labs is not each other. [28:19] It's... [28:20] Open models. [28:21] Because if open models are really good, then it really puts a lot of pricing pressure on them. [28:26] So their interest is, it's just, you know, the big four and no one else, no one else in town, no one can compete.
And so the best way to do that is say... [28:34] You know, it's either us or the Chinese models. Yeah. Oh, like, you know, it's our enemy. You don't want to use their models. Right. When the reality is, no, it's like it's them versus the open models. And currently, yes, the best open models typically come from China. [28:47] Um, [28:48] However, you know, NVIDIA has Nemetron, which is actually catching up quite a lot. Still on the frontier, and I do think it's really important to see US Open models get better.
I think it's just a little like sinister phrasing to try and get you subconsciously. And it works. Like enterprise CIOs that I speak to, they will often say like, oh, we don't use Chinese models. [29:07] And I'm like, okay, we can use other open homes. Those aren't the only ones. Yeah, yeah. But it's crazy. Like the propaganda works to get them to think, oh no, you can only use these. And so just like kind of ringing the alarm bells on that is something that matters to me. It also matters to me that we have good US open source.
You had a spicy tweet earlier this week after the Fable 5 launch. That seemed like it actually had a, you know, maybe it contributed to a positive outcome in some respect. Safe to say you were not a fan of the way the Fable 5 rollout happened from Anthropic. [29:37] Yeah, I think there were two things that I was really disappointed to see. One being they changed their policy to have mandated data retention. [29:47] for the purposes of security, or safety rather, which is important. [29:51] But I think it sets a precedent for basically, and that's even for the enterprise.
Typically models, if you use it in the like, you know, self-serve plan, there might be data retention. That's pretty standard. [30:03] In this case, it is for everyone, even the enterprise, requiring data retention, which talk to a financial services company. They're going to be like, absolutely not. And all indications suggest that they want that to be the case for every model going forward. It sets a very bad precedent of basically we get to see everything that's happening. And I think what's worse is they also... [30:23] you know, at least initially we're posturing that, [30:26] they would...
[30:27] either deny your service or [30:30] hamper your performance, [30:31] not based on violating terms of service or violating the law, but doing things that they don't agree with. [30:37] And that I think is the line that's really important we don't cross. And even the fact that it's a possibility for any of these providers to do makes it so important that we have one competition in the closed models, but also open models. [30:49] Because, [30:50] The crazy thing that is still the case right now is that [30:54] If you are doing something that [30:56] they deem competitive or risky initially without telling you would degrade the performance of the model and have it give you kind of garbage answers specifically on ai research right i think on the bio research they tell you no
[31:09] Like AI research in theory. Okay. You think it's just like a little. Oh, whoops. We made a mistake. You know, like. [31:17] Who knows? It's hard to attribute it to like, you know, malice or just, you know, making mistakes. But I think the precedent that it sets initially was going to be without telling you it would dumb it down. Yes. And that's really, really dangerous. And so because of a lot of the pressure, I don't know if my tweets had an impact there, but I think a lot of people had a really bad reaction to it.
[31:37] They walked back just one part of it. They walked back to the part where it would do it without telling you. Now it would do it, [31:43] while telling you. Yes. Which I think is still really bad because it's still degrading performance. And I saw countless people and friends of mine as well where they're doing like, [31:53] cutting-edge bio research, looking into like prostate cancer markers. [31:57] And they were being denied their requests. And the thing is, even if they go and then fix it, the point is, they now have an ability to, at any point in time, deny requests.
[32:06] your service, [32:07] And then say, oh, sorry, it'll take a week to fix. Oh, it's there. [32:10] And so... [32:11] As people become more and more reliant on these tools, you want to make sure [32:15] no one can kind of quickly affect you. [32:17] over this period of time and this breaks that precedent and does so not just in the weights of the model where before in the weights of the model if you're like [32:25] whatever, hey, build me a nuclear weapon, it'll say no, like, I'm not going to do that.
Now it's kind of programmatically added on top, where in theory, there's someone over there who could like change a line of code. And now it's like, oh, anytime Mario asks about like world events, like, [32:39] convince him of this crazy political ideology and like just sent him in this other direction without you knowing. [32:44] And that's crazy. I do think the transparency piece for me, like, felt like the obvious misstep. But I'm... [32:52] almost, I'm quite sympathetic to the idea that you do have to have some limitations of what people can use these models for.
Like I, I, I do think flagging some cutting edge bio research or, you know, there are obviously bad uses there. Like, where do you think the right place to set [33:08] the line is when you're operating at the frontier here? First of all, the one thing that's an obvious kind of baseline is the law. [33:15] And like, [33:15] We live in a country that has laws and we should defer to the government to enforce those laws. And to the degree to which we can, we should work with the government on figuring out what the best way to help that is.
And obviously platforms like Facebook and Twitter work with the government a lot to make sure, on one hand, you can have free speech. But on the other hand, you adhere to laws about hate speech and things like that. [33:36] And so I think, you know, we should take from that and act [33:39] in a similar manner there. [33:42] And then I also think whichever way we go down, it has to be [33:46] very [33:48] clear who is getting what and it needs to be equal. [33:51] I think that's the thing that's really...
[33:53] not obvious to me is like, [33:56] Tech has a history of people doing [33:58] bad things and then getting punished for it later when it's already past the point of no return yep and that's the thing that i think worries me is like for example [34:05] all of these model companies, trained on the whole internet without asking anyone permission. Yes. Right? So we already know that there are companies, and I'm not saying any of these companies right now would do that or anything like that, but we know that tech has a history of, like, doing things that is irreversible and then suffering later, but, like,
[34:22] What penalty could you provide that would make up for that? There's nothing you could do. It's like once you pass the point of no return. [34:27] And so like take, for example, data retention. [34:29] If you retain all the data of all the financial services companies and all that, [34:33] You could train on all of it. Oh no, it's a whoopsie. And you try to make up for it 10 years later. It's going to be in courts forever. You're past the point of no return. You could have trained on every company's data.
It's illegal and you break the law, but that hasn't stopped tech companies before. Now, I'm not like super pro-regulation, but I think the important thing is like, [34:52] having the right checks and balances in place to make sure that [34:55] No one company can kind of go too aggressively monopolistic. [34:59] In any of these directions. Yeah, you know, I think, uh... [35:02] someone who we had on the podcast who, you know, in this, in this location, uh, was Vincent from Prime Intellect. And I liked the way that he frames. [35:10] The idea that, you know, a safer world is ultimately one with sort of [35:14] multiple super intelligences rather than just one and you know that was the first time i'd had someone put it to me in that way and the more i thought about it the more i sort of think that's [35:24] correct and so you know this is why you do need this robust ecosystem that that you are sort of helping companies [35:30] operate into and abstracting away that complexity.
[35:33] Maybe let's take a step back to some of your story. How does one decide to become a string theorist? Is that like you first want to be an astronaut and then make a jump? What was the childhood ambition there? I was quite a rebellious child. In elementary school and middle school, I was always skipping class, getting in trouble. I [35:56] Generally, I was just bored by things and would just kind of get into a lot of trouble. But I was always good at math. [36:02] But I didn't really try in school because I was like, oh, I want to do whatever other random things.
And then my eighth grade geometry teacher. [36:10] told me that I should retake geometry in high school. And that was like a trigger. I was like, what the hell? Like she thinks I need to retake geometry. Shout out Margie Greenwald, eighth grade geometry teacher. Look what she created. [36:23] And so I was like, I can't believe it. Like I'll show her. And so my first order on Amazon ever was textbooks for Algebra 2, Trigonometry, Precalculus, Calc 1, 2, 3, Linear Algebra. And I studied all of those textbooks in the summer before ninth grade.
[36:39] And then I asked my dad what the hardest math was. [36:42] He said string theory. And I was, technically it's physics, it's not math. But he said string theory and I was like, okay, I'm going to be the best string theorist in the world. And that was like, that was all I cared about for basically the next 12 years of my life. [36:56] So you have a very laid-back personality, clearly. [37:00] Wow, fascinating. Were your parents academics? They're both engineers. So my father was immigrated from the Soviet Union. [37:09] He studied chemical engineering and ended up working actually, what brought him out to the Bay Area, he worked at IBM.
My mother was a programmer at a hospital. [37:18] This episode is brought to you by Persona, the B2B identity platform helping businesses verify users, fight fraud, and build trust. Fraudsters are already using AI to spoof faces, voices, and documents, so your defenses need to adapt just as fast. [37:33] Persona helps secure some of the Internet's largest and most trusted platforms with identity verification. If you're building a product where trust matters, identity should be a priority. [37:43] you've probably already experienced Persona without realizing it. [37:47] verifying your LinkedIn profile, signing up for Etsy, [37:50] or renting a scooter with lime.
[37:52] trusted by leading companies like Square, Brex, and Twilio. [37:55] Persona gives you the building blocks to create identity flows that adapt to your customers, risk tolerance, and locales you operate in. [38:03] Whether you're verifying age, onboarding businesses, or automating KYC. It's fully configurable, so you can launch in days, not quarters. [38:12] Want to see for yourself? Generalist listeners get a free year of the starter plan. Head to com slash generalist and check it out. [38:21] Okay, so you had sort of technical conversations in your life, but not string theory conversations by the sounds of it.
Yes, that's right. So you sort of run on spite for 10 years through some very impressive places. You worked with... [38:36] what I understood from my research, you know, I'm... [38:39] Wouldn't have known myself, I'm embarrassed to say. It's like sort of the greatest living string theorist. Kwan-Waldus Hainer. Yes. What was that like? Like, what does one take from... [38:49] a mind like that i imagine that to be that person you know the sort of most numinous uh luminous you know person mind of this field of your generation you must have a very [39:01] unusual way of operating, unusual way of thinking that sort of goes beyond even that domain.
It was a huge privilege. He's one of the kindest, smartest people I know. Also, I think very fascinating. He's quite religious as well, which is typically rare for a strength theorist. And, you know, it was a huge privilege working with him while I was at Princeton as an undergrad. But I think it also, even having that happen, that was a lesson and one of the examples for me of why mentorship matters so much. [39:27] Because initially I was going to do research with this different professor. [39:31] But there was this one grad student that I knew who kind of took me under his wing and was giving me a lot of advice about, you know, what I should do.
[39:39] And he was like, no, Matan, you need to go and talk to Juan. You should try to get him as an advisor. [39:43] But Juan technically wasn't at Princeton. He was at what's called the Institute for Advanced Study, which is like a very closely affiliated but technically separate institution that had no teaching obligations. So in other words, they didn't have to deal with the pesky undergrads. Yes. They could just go do their research in peace. [39:59] And I was like, wait, I, you know, he's at IIS. He doesn't want to work with undergrads.
Like, you know, he's like, no, trust me. [40:06] Send him an email about a paper of his that you read. You know, try to have some insight and ask if you can meet with him just to talk to him. [40:13] And I was like, okay, you know, and so, you know, [40:15] spent a lot of time preparing and [40:18] you know, send him an email. [40:20] And I ended up doing that. And he ended up inviting me over to his office. And before I went to his office, this advisor, this mentor of mine was like, by the way, [40:29] The meeting will go well if at the end of the meeting, [40:32] He very calmly and subtly poses a question.
[40:35] And Pose is like, oh, this is interesting. [40:38] and basically you will have 24 hours, [40:41] to solve that problem. And if you solve that problem within 24 hours, he will likely take you on as a student. Oh my gosh, that's so great. And so I remember going into the meeting and being like super on edge, like, [40:54] ready to like find the what's like the thing that he leaves and juan is one of the most one of the most subtle and like soft smoking people soft-spoken people [41:02] And so we spend like two hours at a chalkboard.
[41:05] and [41:06] The meeting ends. [41:07] And, you know, I've been taking notes. He was at the chalkboard and like vice versa. And... [41:12] The meeting ends and I was kind of, [41:13] disappointed because at the end I'm like, wait, he didn't ask a question. Like there wasn't a thing. [41:17] And so I go and, you know, have all my notes and I take it to that mentor and I'm like kind of going through it with him. And and I'm like, hey, you know, I didn't go well. He didn't he didn't ask me the question.
He didn't plant the seed. And so I go through the notes with this mentor of mine. And we're basically like recapping the whole two hour thing. He was like there like that moment there. That was the subtle question. Like if you solve this equation and, you know, do this and send him an email, then he'll take you on. Wow. It was it was truly that subtle that you had to sort of Talmudically. [41:47] go back through this. Wow. Fascinating. That's amazing. I love that story. In some ways, I would imagine the life of [41:56] a theoretical physicist and the life of a startup CEO are almost like the total opposite ends of the spectrum.
In some sense, I could imagine both being quite lonely. But like the startup operator, you know, it's almost like the Marcus Aurelius beginning of meditations where he's writing about like, you will encounter, you know, annoying people today, you will encounter trouble, you like you must expect these things. That's sort of the fundamental nature of like building a business, right? It's like rejection, stress, annoyances. [42:25] Theoretical physics seems very monastic. How do you think about those parts of your personality? Prior to just now, I would have always described it as completely opposite.
But I actually do think you do have a point that there is something very similar. But on the surface, they're the most opposite because as a... [42:42] devoted theoretical physicist, my ideal day was alone in my room for 12 hours at a time, just like reading papers. [42:51] Like that is the ideal day, like not seeing the outside world whatsoever. Now, [42:55] I forced myself to do that for like 12 years because also... [42:58] Partly out of spite, but partly also because... [43:01] The mathematics behind string theory is just so [43:05] It's so beautiful.
It also like the chain of whys, like if you just ask why enough, it always goes to string theory or theoretical physics. [43:13] Or then beyond that, it's philosophy, but we can't prove anything on that front. Yes. But it's kind of the base of everything. And there's something very beautiful about the mathematics there. [43:22] So it did keep me going, and it wasn't purely spite, but... [43:26] It is very monastic. And I would always... [43:29] I've always been very curious just to like, [43:32] get to know people and [43:34] You know, I've had people make fun of me at times where like, [43:37] If I'll meet someone at a party, it'll be like an hour of me just like asking them questions.
And then at the end, they're like, I don't know anything about you. Like answers. Yeah. But it just, you know, it was very enjoyable for me to like learn about, you know, different things. [43:48] But I would always feel guilty doing that as a physicist because those were hours of my time not spent. [43:53] I see. Hard at work, understanding more about the universe, right? Yes. [43:58] And so I think the thing that's, you know, on the opposite side as a CEO and founder, [44:04] The amount of context switching I do is insane.
[44:06] It's like... [44:07] At the end of the day, I typically forgot what I did at the beginning of the day because it's just so varied. But... [44:13] that curiosity, especially when we're selling into the enterprise, [44:16] It's so fun talking to all these CIOs and understanding the different nuances of how they build software. [44:21] And it's very satisfying and I like engaging with people. [44:24] but [44:25] Completely opposite word. There has never been a time where I focused on one thing for 12 hours. [44:30] literally for three years like i've never done that really uh because it's just there's too many too many multivariate problems [44:36] But I think to your point about the similarity, [44:39] There are two similarities.
So on one hand, it is kind of the solitude, which is like, you know, you do feel a little bit isolated, I guess in one case, because of the responsibility, especially after going through the highs and lows. Yes. Where you see the people. [44:54] that are only there for the highs. [44:56] And then the people that are still there for the low, like that kind of really makes it obvious what the. [45:02] the kind of solitude there is like. And then in the case of physics, it's like the literal solitude.
Also, you can't talk to anyone about it. [45:08] Like no one at a dinner, like you could talk about black holes or whatever, but you can't talk nuance about it. That makes total sense. And also it's always kind of funny to people. Like if you talk at a dinner party about physics or about black holes within the, [45:20] three minutes, someone will make like a Uranus joke and then they're like, All right, I'm like, we've done it. Yeah. Yeah. On we go. Yeah, exactly. So [45:28] you know, a similar solitude there.
[45:30] But I think there's also a similarity in accepting that there are going to be so many variables that you do not know. [45:36] In the case of starting a company, [45:39] There's just a million things happening all the time, especially when we get to a certain size. I won't know everything. [45:45] but still being able to figure out how to decide what are the best [45:48] next steps how to make decisions knowing that there are things that you don't have time to fully understand similarly in physics [45:55] It would take you a full lifetime to read, more than a lifetime, to read all of the literature that has come before.
[46:00] And you just have to accept, OK, I don't know the full derivation of this equation, or maybe I don't have the deepest understanding why this is true. But in order to make progress here, I just need to accept that, get some intuition about it and just move forward. And so there is some some parallels there. At what point did your mind begin to wander? [46:18] from string theory? Was there a moment where you've noticed yourself sort of [46:22] Yeah, so when I came to Berkeley to do my PhD, that was when I first had [46:28] GSI duties or graduate student instructor duties.
[46:33] I've never been... [46:34] too fond of teaching. Not that I like it. [46:38] I enjoy it, but not as a job. Like it's enjoyable when people care. [46:41] As a GSI at Berkeley, you teach people who do not care and are there just for the kind of, uh, um, [46:48] you know, requirements. Yeah. [46:50] And it took me until then to realize, but basically, I knew everyone who I would ever work with for the rest of my life. There are very few strength theorists in the world. I'd met everyone that I would be working with basically for the rest of my life, except the potential students I would have.
And I would basically be surrounded by children. [47:07] for the rest of my life as well, because you have to be at a university to be a theoretical physicist. [47:11] And that kind of led to a cascade of weight. Oh, my God. [47:15] do I actually want to do this? Is this even the best like fit for me and my personality? And so it kind of led to this, this big cascade of kind of questioning everything. And then it [47:26] kind of slowly... [47:27] broke down the stubbornness that I had and I was like, okay, you know, maybe, maybe there are other things.
[47:32] that are out there for me. I imagine that was like extremely difficult. If you've identified, you know, you've sort of, [47:38] Attached yourself to this like vision of yourself. I certainly had versions of that where for much of my early life, I was dead set on being a lawyer and then working at a law firm disabused me of that notion. Yes. [47:50] But it was painful. It was, you know, difficult to sort of remove yourself from the dreams you had, like, [47:56] What costumes were you trying on at that point of like, I might be this person.
I might be that person. That's honestly, that's such a good way of putting it. Cause that is actually what it felt like. It felt like this costume. Like I had, if you would ask me, it was like. [48:09] I was like a physicist before I was a human. It was like that deeply associated. And so it was like, [48:17] completely earth shattering. [48:18] trying to understand what do I attach to? Like, this was everything that I... [48:22] spent all my time on. At a certain point, like kind of just leaning into the fear, it became somewhat freeing.
[48:28] Be like, I can just do it at like, [48:30] the whole, like I felt so guilty about any hour I wasn't working. [48:34] on physics that now is like, oh, okay, you know, let me explore other things. And so one thing was, you know, I hadn't explored humanities as much, and so... [48:41] set aside some time and I was like, you know what, I should become more normal, well rounded human. So I like, [48:46] watched the top 250 movies on imdb oh wow as a way like okay you know you gotta and that was so much fun like i've seen batman yeah but like i you know so many of them you know brought me to tears because they were just so beautiful or similarly like the top 100 pieces of literature i was like you know we gotta read these and looking back that was one of the most fun periods of time because it's like [49:07] These pieces of literature are famous for...
[49:10] very good reasons. [49:11] Um... [49:12] And getting to experience some of them for the far, some of the movies and books I'd read before for school, but like experiencing it for the first time is just so much fun because they're all just such legendary pieces of art. Okay. You can't bring up this topic without giving some names because I love talking about books and films. Like, yeah, what struck you? Okay. So the first piece of literature that made me cry, it's so cringe to say, but A Tale of Two Cities by Charles Dickens.
I remember I like called my dad because my dad also loves literature. I remember calling him and I was like crying and I was like, this is so beautiful. [49:42] Like Dorian Gray was fantastic. Brothers Karamazov, I think, is incredible, although there are some sections that are completely a waste and some that are just like, [49:51] absolutely beautiful. Honestly, like, the best to ever do at Shakespeare. [49:55] Again, like, [49:57] underrated even though he's the go-to yes he's [50:00] just so incredible. Yeah. As for films, um, [50:04] One that always sticks with me and I watch all the time is Harakiri.
[50:08] It's an incredible samurai film. [50:10] black and white, but it's [50:13] Oh, it's beautiful. Wow. You should really watch it. I definitely. Honestly, watch it and tell me. Like, I would watch it with you. I, anytime I get to rewatch it. [50:21] Um, it's, it's fantastic. There's, um, [50:25] There's a good Italian film called La Grande Bellezza. Oh, I love that film. It's so good. I went to the theater like three times to see that movie. What a movie. So good. Just visually so fantastic. I can't even relate to it that much.
It's like about an aging, you know, man who used to be the center of the party and now is, you know, I had nothing to relate to it, but somehow it still just like pulls your heartstrings. Yes. What are some for you? I actually, I would often say La Grande Bellezza is one of my very favorites because I think it's just like. [50:52] So, so amazing. There's a French film. You were mentioning this black and white Japanese film. So maybe that's why it came into my mind called La Haine. Oh, so good.
I love that film. Just because it tout va bien, right? Oh yeah. That's right. Yeah. Gosh, wow. I can't believe that. Uh. [51:07] Yeah, that was really great. Yeah, I'd have to think more on ones that I would really ride for. My ultimate favorite, have you ever seen The Master by Paul Thomas Anderson? I have. I really loved that. I thought that was an incredible film. [51:19] Um, okay. Amazing. So you could talk about this for hours. I also, there's like, I have a note on my notes app. I have a list that I'm, [51:25] tempted to pull out, but maybe we save this.
Yeah, we'll do that. Yeah, we'll do that after. I feel obliged to at least ask you a little more about your journey. You have this moment where you sort of explore the, you know, explore the glories of the broader world. Yes. At what point do you migrate sort of from [51:42] uh you know uh literature to to tech you know yeah so you know in my head i kind of [51:48] Funny enough, there was probably a month where I was like, you know, maybe I could be a writer. And I was just trying some stuff.
Absolutely not. Not for you. Not my strong story. And it became pretty clear. Okay. You know, I spent the last 10 years becoming pretty good at like quantitative things or math.
Want to learn more?
Ask about this episode