Back to Nick Test

Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis

NT
Nick Test
@nick-test

Dylan Patel, founder of SemiAnalysis, argues the biggest gains in AI don't come from faster chips, they come from software-hardware co-design. Optimizing the model, the kernels, and the silicon together turns a 2x here and a 2x there into 100x. He explains why DeepSeek's experts were shaped for Nvidia's Hopper (and why TPUs struggle to run it), why OpenAI's sparser models and Anthropic's denser ones pull them toward different hardware, and why the so-called CUDA moat was never really about CUDA. Dylan breaks down InferenceX, his living benchmark that runs the latest models on over $50M of donated hardware daily, tracking a roughly 60x annual drop in cost per unit of quality. He makes the case that inference will be a bigger market than oil, that the compute crunch persists because models expand the value of useful work faster than compute grows, and why Jensen Huang is bankrolling neoclouds to engineer a multipolar world.

Appears in

Uploaded
Uploaded Jun 30, 2026
File type
POD
Queried
0

Full transcript

Showing the full transcript for this episode.

[00:00] I think it's really fun inside of Semi-Analysis because we have 90 people and like a big chunk of them are technologists, engineers across the whole supply chain, and then a big chunk [00:09] is people who are formerly at hedge funds. And you see these arguments, like people are like, oh, well, that doesn't matter. And it's like, then someone's like, well, but cost. And then the engineer's like, no, no, no, but this technology is the coolest. And you see this organically fight it out. And we're pretty informal. And given the fact that I was a forum moderator, you can imagine what this internet swag looks like.

You're enjoying it. You don't wrestle with a pig because a pig enjoys it. [00:39] Bye. [00:50] - We're here in the Semi Analysis office with Dylan Patel. You know, I'm Sean from Sequoia. I have my partner Sonia Huang. Pretty insane what you've done. Semi's five years ago were not very sexy in the West. They were sexy in the East, but people here in the West had kind of forgotten about them. You did not forget about them though. You went very long. [01:11] You created probably the premier research company in the space that's been educating [01:16] the world and, you know, the state of the art from very technical details to supply chain, you know, to the bigger picture.

There's rumors that semi-analysis recently passed 100 million of revenue. I don't know how accurate those are, but whatever the numbers are, you guys are crushing. It's as accurate as the information is, you know, you never know. There's also rumors that you might start a venture fund. Like, you know, I hear all the time in the ecosystem, people wanting, you know, affiliation with semi-analysis. You've built this trusted brand. And so whatever you do, [01:46] working is clearly like just the beginning of the journey for you. Congratulations on that. But how did this happen?

Like, how did you, first question is like, what is the background? How did you kind of get to where you are now? Well, when I was a young boy and, you know, coming out of the womb, so, so, okay. So I grew up in like a small business. My parents had a motel. We lived in the motel. We laid our gas station. So, you know, uh, I was selling, you know, I joke a lot of times, the first neural network I trained, um, [02:15] was racially and visually profiling people based on when they entered the gas station, which cigarette to pick.

Basically, the cigarettes were all extrude across the top, and I was too short to actually reach them, and technically it wasn't legal to sell cigarettes at that age, but whatever. I had to move the stepstool over to the right area. I started working in my first office before it was legal, too. It was a good experience. Well, I didn't get paid, right? It's a family business. Same. We had a motel, and then across the street was our gas station. [02:45] And so if an old white lady with curly hair walked in, I'd move the ladder or the stepstool over to where the camels are.

[02:51] And if, you know, different age, demographic, profession, you know, race, et cetera, I would move the steps all over. And I joke this is the first neural network I trained because if I waited for them to tell me, I'd have to, like, move it over. And then I'd step up versus, like, just being ready. So, you know, menthols versus, you know, 100 slibs and all these things. I joke this is the first neural network I trained. But I grew up in family businesses, lived in a motel. And, you know, [03:15] So it all really goes back to when I was like, you know, it was my eighth birthday.

My birthday is in May, and it was April when the Xbox 360 was announced. For my birthday, I didn't ask for the Xbox, or I didn't ask for a birthday gift. My parents asked what I wanted. I asked for it for Christmas. We celebrated Christmas, but there was no way, at least at the time, I thought there was no way they would ask, would give me the Xbox 360 for Christmas. And so I got it for, asked for my birthday for Tab for Christmas. Anyways, Christmas comes around, I get it.

You know, fast forward a couple months. [03:45] Emma, they also lived in a motel, was going to come over for spring break, for his spring break, and we were going to hang out at my house. And he's in between me and my older brother and age. Brother's a bit more jockey, so he didn't really care too much about the Xbox. He played sometimes, but didn't really care. [03:59] But my cousin, you know, I wanted him to think I was cool, right? So I bragged many times on the phone. I was like, yeah, I got an Xbox.

And then the Xbox broke. There was something, there's a hardware defect called the red ring of death. [04:11] But long story short, I had to open it up and, you know, short the temperature sensor and it fixed it. [04:16] But there was many other tricks I tried first and then other worked. And so that's sort of how I, like, got into hardware. I was, like, open-peared-doors box. By the time I was 12, I was, like, on these forums a lot, reading, posting a lot. And this was around the time when Reddit ate all other forums.

And so I became a moderator of, you know, Android and Apple and Google, as well as, like, hardware. And was watching, you know, looking at Intel, NVIDIA, and AMD and all these other forums, right? So I was building a PC, all these forums. I was watching, reading, posting a lot. But some of them I was moderating a lot. [04:45] So, you know, smartphones, watching smartphones develop from, like, very simple to speed racing to being technologically more advanced than PCs in many ways, architecturally. And same with, like, you know, all the GPUs, like, just tracking and watching that, reading every comment, always having the economic tinge because I grew up in a small business.

So I was always looking at the economics, right? There was a time where all the, like, let's say, neckbeards on the Internet loved AMD GPUs. [05:11] And, like, I personally had bought an AMD GPU, too, because price, performance. But then when it came down to, like, what's technically better, I'd always be like, no, no, no, NVIDIA is better because they use a smaller chip to get, you know, better performance and better power efficiencies and their margins better. And so, like, I would always, like, talk about how NVIDIA's margins were better than AMD's in the GPU landscape.

And so it was, like, very fun. And you were 12 at the time? I started moderating when I was 12, but this is all through my teenage, tween age and high school years, right? [05:41] at semis. I played a ton of StarCraft at one point. I was Grand Master on the North American ladder. Oh, wow. So you've gotten just obsessively good at multiple things. Yeah, I mean, obsession is good. How are your grades? They were... [05:55] decent, [05:57] I would say like [05:59] I had mostly A's, but they're classes that I thought were really boring or just didn't enjoy.

[06:05] Spanish, I got not the greatest grades. I speak fluent Spanish, by the way, so it's really dumb. But it's just sort of- - But maybe that's why you didn't get a good grade. I didn't learn Spanish until later, to be fair. So my grades were fine. [06:24] They were fine enough for Asian parents. I was better than most of school, but, you know, it wasn't like, you know... [06:30] Try hard maxing for like, you know, all A's. [06:33] Okay, so you're very much a student of the internet then. This is how you develop this expertise.

At what point did you decide to start Semi-Analysis and what's been the biggest surprise since starting the company? [06:42] Yeah, so I went to school. I got a few degrees in stuff that wasn't related to semiconductors. Was a quant for two years at a small quant risk firm. And then basically, you know, there was a culmination of events that happened, right? One was that my, you know, sort of like I got screwed out of a bonus. I'd made my company many millions of revenue, of risk-free revenue, because I exploited like a risk thing in the market.

You know, I think well over $10 million. And then someone else took credit for my work and all this sort of stuff. [07:12] But, you know, I lost the social contract with the company I was working with. I had some, you know, my grandparents grew up in my house with us or in the motel with us. They lived with us. And so, you know, very close to them. And my grandmother got dementia and she forgot who I was. And she fell down some stairs and had like a tragic accident and passed away.

So all of that happened in early 2020. Additionally, there were some like, you know, girl things. And so, you know, there's a few things that happened that made me like kind of very sad, um, [07:39] And so all of those things sort of culminated. Then COVID happened. And my brother's like, dude, just come stay with me. He lived in Nashville. So I came and stayed with him in Nashville. We were like, oh, lockdowns will be a few weeks. You can stay with me while they happen. And then you can go back home and, you know, whatever.

Famous last words, lockdowns lasted much longer. But, you know, living with my brother for a few months, you know, it was like sort of like, okay, didn't know what I was doing. I was now at my brother's home. Everything was his rules, you know, sort of like, you know, him and his fiancee at the time, now wife, you know, were like there. [08:09] Thank you. [08:09] I basically had to tiptoe around, but I didn't care about my job, and so I was, like, posting even more than normal. I had always been posting a lot on the Internet.

I'd always been trading stocks a lot, but, like, I made a lot of money shorting COVID and long in COVID and, like, all this stuff. Semiconductor shortages happened around then, too. And anyways, I was, like, very much obsessed with posting and things like that. And eventually, around that time, I got into an argument with someone on the Internet, and they doxed me. [08:32] They publicly revealed my identity for my anonymous account. [08:35] And at the time, I was like, oh, no, I was scared. I stopped posting for, like, three weeks, and I was like, what am I doing?

Why do I care? So then I just started posting under – I had, like, blogs and stuff as well. I made a real blog. [08:45] semi-analysis. And on my 24th birthday, I posted, um, you know, two blogs. And, and then from there, it just like, it was not a newsletter, but I got so much traction because now instead of posting on anonymous name is a real name. And I put a lot more effort into those two posts than I usually did. [09:02] Instead of like shit posting on the internet, it was like real effort into the blog.

You can actually go back and read those if you want. They're not that great, but, you know, they were good for the time. They were the best stuff you could find on the internet about some of these. And I just kept posting, posting, posting. I started getting a lot of consulting business. You know, 2020, I also sort of... [09:17] I was, again, crashing out, didn't know what I wanted to do. So I packed everything up or sort of I took my truck. I bought a tent that fits on the back of the tent truck, bought an air mattress, whatever, and would like drove around all these national parks all around America.

And so like two or three or four days of the week, I'd stay in a random motel where I negotiated the price to be like $30 a night for a room. And I would work on something else's stuff. [09:47] semiconductors about ai about all the things that i cared a lot about and got way more educated over these six months where i'm just like going to every national park um and then the whole time i was alone that i was alone the whole time i was posting blogs um everyone was like dylan what the are you doing pre-star link or the very early days pre-star link pre-star link um yeah so it was like very much like what are you doing um i travel around latim again like for for a year initially with my friend

[10:17] 23, 24, end of 21, 22, 23, and 24. I'm completely, I'm still completely homeless since mid 2020. Right. Um, [10:26] But I'm traveling around to every conference in the world. I go to 40 plus conferences a year. [10:31] no matter where in the supply chain it is, I'm like, oh, that looks interesting. I guess I'll go to that. [10:36] I went to one conference like, wow, this is amazing. You get to talk to the experts, and they're going to talk to you, and then you're so excited. In the case of semiconductors, everyone's a boomer, so it's great.

They don't see young people who are excited about it, so they're really happy to tell stuff. I have to ask on this. Was there a part of the supply chain or one of these conferences that... [10:58] particularly change your view of the semi-world or that you felt then or feel now is particularly underrated? [11:05] I think the trade shows and conferences range really widely. Obviously, some of the ones I have the most fun at include NURIP. So why is that? Because it's 20,000 AI researchers, and they're generally in my distribution of age range.

So it's like a lot of fun. But they're also like leading AI researchers, and it's a lot of fun, and you learn a lot. There's also a lot of parties. And then it ranges all the way to like… [11:26] There's a random chemical conference in Japan where it's 300 Japanese dudes. It's like 20 guys from ASML, 20 guys from TSMC, 20 guys from Intel, and those are the only people who speak English. Everyone else speaks only Japanese, and you're like, I guess it's still pretty interesting and fun. I think one thing that I have a skill set of is I'm able to bond with anyone regardless of their background and who they are.

I'm able to talk to them, find something interesting to talk about. Oftentimes it's the tech stuff. [11:56] you know, the really big ones because that's where the big stuff is happening. But I think the niches that are really, really exciting is like, you know, SPIE. So there's IEEE, which is International Electrical Engineering something. And then there's SPIE, which is another ecosystem. SPIE conferences are super, super deep in details. Every single one that I went to, especially like SPIE Advanced Lithography or SPIE PhotoMask, [12:21] I went to them the first time.

I didn't even understand 90% of what I heard. [12:25] And then I read, read, read, read. I made some contacts, of course. And the next time I went, I understood like half of what I went to. Third time I went, I understood like 75% of what I went to. Even now I went and I was like, I still don't understand everything that's going on. Whereas like you go to like NeurIPS, you know, a couple times you can understand, okay, what's neural symbolic reasoning? Okay, what's this? What's that? Like you can kind of get a mapping of what everything is pretty quickly.

But some parts of the supply chain are so arcane and so deep and so technical that [12:50] It takes a lot of times for you to even understand what's happening, you know, on everything, right, for every research paper. It doesn't necessarily mean you didn't, you know. [12:57] You go to a conference for a few reasons, right? You understand the research, you understand, but like, it's all the research that's being published, but what you really care about is understanding [13:06] How does that research intersect with technology? Also, how does that research differ from what's there today?

And none of these research papers tell you what's happening today. But then you just ask people, and you build contacts, and you learn, and then you, like, learn about the supply chain, and, oh, this company supplies this company, even though it's not publicly stated anywhere. Or, like, you know, you learn that this chemical is, like, cost about this much, and a tool uses about this much, and you just learn all these things. You hear the horror stories of, like, this chemical had a shortage, and it – [13:33] totally threw off this part of the supply chain.

And then it turns out there's only three companies in the world that make that chemical. And it's like my favorite one is I learned a Japanese guy at that specific Japanese conference that I went to were no, almost no one spoke English in very broken English. He told me about how his father worked in this, in, in, in, in this industry. In the 1980s, that the only factory in the world that built this chemical burned down and that caused memory prices to like double or triple. And I was like, wow, not too different from today.

Not at all. [14:03] In France, going to be the biggest market on earth. [14:07] biggest market beyond earth. Agree or disagree? [14:11] I mean, obviously use of tokens is going to be the biggest market. Yeah. And the value that's created from tokens is going to be the biggest market. But I think tokenomics, sort of the use of tokens, adoption of AI sort of is the most important thing that's happening. And inference, whether it's open models or closed models, will be like one of the biggest markets in the world. Much bigger than oil, I think, much bigger than like, you know, many other parts like inference, [14:31] of AI will be many percentage points of the GDP.

Yeah. Right? What you've done with inference X, I think, is industry standard. Maybe say a word on why you started it, what it does, and... [14:43] What do people misunderstand about performance benchmarking on inference? Yeah, so to zoom back, right, like semi-analysis, we do a lot of stuff that's like, you know, a lot of it is like research for institutional clients and our subscription-first products. But a lot of it is also like, hey, you know, this would just be cool to figure out. Let's figure out how to figure it out and just post it publicly, and that begets, you know, more and more scale.

And so we've done this with a lot of GPU benchmarking and testing and training performance and inference performance. But, you know, ultimately we saw like, [15:11] Inference benchmarking was like point in time. You test it, and you take some time, and you release it, and it's slow and arcane and outdated. Because models change all the time. [15:22] I feel like every week there's a new model, whether it's a Chinese model or, you know, today, Mythos 5, Fable dropped. And new models are coming out all the time. On the software layer, PyTorch, VLM, SGLang, new drivers, new something drops.

You know, in fact, the update cycle for most of these libraries is twice a week. So you basically have the software updating all the time and therefore performance changing. You know, new inference optimizations are coming out, and those get updated. [15:52] driving efficiency and cost down, which is why we've seen model costs drop for equivalent quality by 60x a year. It's incredible. [16:00] But to stay on top of that, you can't have point-in-time benchmarking. You need to have benchmarks be living and breathing,, you know, constantly running on the latest hardware, on the latest models.

And so – [16:10] We embarked on a project and we got a lot of buy-in from the ecosystem. This was only possible because we had enough Aura with some of the ecosystem where we were able to get CoreWeave and Crusoe and Nebius and Oracle and Microsoft and Amazon and Google and OpenAI to contribute to us, compute. And then we were able to work with SGLang and VLM and now Radix Arc and Infraact, which are the private companies who are sort of leading those efforts, the open source efforts, to collaborate with us.

[16:40] to collaborate. Now we've got all these people collaborating. We've got over $50 million of hardware donated to us. Once we launched TPs and training, it would actually be over $100 million of hardware, you know, [16:51] maybe about like 15 different chip types, all running these benchmarks every single day on all the latest model, right? The best model for Moonshot, the best model from Alibaba, the best model from – there's about five different Chinese models, the best open source models, the best Chinese labs there. We run benchmarks on their models every day, and then also the best

S. open source models, GPT-LSS, Nemotron, et cetera. So we're running these benchmarks every day in an automated fashion, and they run on these servers that are dedicated to us for inference benchmarking. [17:21] so many different configurations and optimization types. And then what it creates is, and all the results are public and all the configurations are public, so now we have the Pareto optimal curve, because a lot of [17:31] You know, times when people are comparing inference performance, they're like taking a suboptimal curve or point for someone else and comparing it to their optimal one.

It's like, well, yeah, I can make... [17:40] I can stick, you know, if I drove a Porsche versus like some race car driver, obviously I'd drive it slower. It's the same thing with inference benchmarking. And so what we did is we created... [17:48] open source, basically containers for the optimal points across every point on the interactivity, e. how fast is it responding to me versus batch size, e. how many users am I simultaneously serving, curve. And so now anyone who wants the optimal point can just go to inferenceX, download it, and run that as the optimal point, and they can check every day if they want, or they can even auto-download the most optimal point for that model, and their inference performance will be near peak.

Is that curve like the most important curve, in your opinion, the throughput interactivity curve? [18:17] Yeah. I think... [18:20] Most things in hardware, infrastructure, model, application layer, everything is downstream of that curve, right? Is it something that needs to be super, super fast, super low latency, and I don't really care about the cost, so I make batch size very low, and I use techniques like speculative decoding or multi-token prediction heavily, and there's so many… [18:39] possible techniques there? Or is it something where actually I'm batch processing a ton of documents and I don't really care about all these things.

I don't use these techniques that actually are worse on cost efficiency to help you with speed for an individual user because I just want to pack a bunch of users. I don't care if the document takes all night to process. And right now the way we treat AI infrastructure is it's like one size fits all. But over time we're going to get to the point where there's stuff where you have batch workloads or you need instant response. And there's the whole curve that's going to matter for [19:08] And so we see this with Anthropic, right?

Cloud code fast mode costs way more than regular mode. And same with OpenAI's priority queue thing. Sorry, dumb question. How does cost factor into the start? So if I, let's say, imaginary example, I have a batch size of 100. Okay. And I can do 10 tokens per second per user. So in total, I'm doing 1,000 tokens per second. [19:30] off of that one piece of compute. That's one side of the curve, super slow, 10 tokens per second. You know, other side is I have 500 tokens per second, but I can only have one user.

[19:42] And so maybe 250 tokens per second, one user. And then there's points on the middle that are more freudal optimal, right? The average person actually wants like 50 or 100 tokens a second and maybe the number of users I can bash together. So the curve is, okay, 1,000 tokens. [19:56] total per second or 250 tokens total per second, depending on how many users I batch. And there's a curve in the middle. And so ultimately some workloads will actually want the 4X cost decrease because the same unit of hard work can do a thousand versus 250.

And some users, I'll pay 4X more because I don't care about the price. I care about time because the person using the tokens is expensive or the feedback loop that I have here is expensive. If you had to guess, you choose the [20:26] What do you think will happen in space? It can be 0%, 50%, 99%, 99%. [20:33] This is a tough one. So I think the time frame, like whatever time frame and you're [20:38] So I think the non-consensus, or at least against SpaceX thing, you know, I love SpaceX, by the way, and I totally would buy the IPO if I could buy stocks.

[20:47] Non-investment advice. Non-investment advice. Thank you. Thank you. Not-investment advice. [20:51] From Sequoia. [20:53] I don't think that space data centers will really matter in the next, you know, three to five years. With that said, I think in, you know, 20 years, I think the vast majority of compute will be going in space. Yeah. [21:06] And so the real factor there is sort of, you know, what's the cost? It's the time frame. It's the cost of building power on terrestrial land and how much power are you going to be able to do on terrestrial land.

And I think, obviously, my views of where inference, you know, how many gigawatts or terawatts are devoted to inference is a crazy curve for me personally. What's your forecast? How many gigawatts are? Yeah, I think by 2030, just open ion anthropic will have over 100 gigawatts combined. [21:32] And then you'll add, you know, meta and Google and, you know, so on and so on and so forth. It's a humongous amount of compute that will be dedicated to inference. And by like 2040, it'll be terawatts, right? The curve of like productivity that we're going to get.

And so, you know, inference deployments is going to be huge. And so if you look at like 2040, I think like, you know, probably more than half of the incremental compute will be going in space. [21:54] But if you look at 2030, I think it's sub 1%. Do you think intelligence per watt has been increasing? And then it seems like there's still a giant gap between where we are, intelligence per watt versus, like, human biology. And so, like, if we are, do you think we are to close that gap? And if so, where is that game going to come from?

[22:09] Yeah, I think it often depends on what you're doing too, right? Like a TI-84 is way more intelligence per watt in terms of doing math than us. That's like 30 years old, right? So obviously this is like a dumb, dumb, you know, sort of- General intelligence. Yeah, but general intelligence wise. So one of the things InferenceX does is we also measure the power and cost of all of this hardware. And so we offer not just, you know, throughput versus interactivity, we offer cost versus interactivity. We offer power versus interactivity.

And so- [22:35] As far as has, you know, intelligence per watt been increasing, I mentioned, you know, it's been a 60x cost decrease for same benchmark level. We've also seen the same on intelligence per watt. It's not been exactly 60x. It's been closer to like 40x. Some of the efficiencies are non-power ways. But there's been a humongous improvement in intelligence per watt on an annual basis, at least 60%. [22:58] so far this year, last year, year before, year before. And I expect that to continue. As far as where we are from the human brain, we're many orders of magnitude away.

Thankfully, [23:07] doesn't really matter. We can devote a lot of power to computers, much easier to power computers than human brains. Like, you know, we have sickness, disease, and like food preferences. Sleep. Exactly. Let me just ask one more question on the, like on the general theme. In my opinion, in terms of like, you know, intelligence per person, [23:27] water intelligence per person [23:31] dollar, like any of these metrics. I think there's kind of three levels of input. You can get hardware improvements where the hardware is more efficient. You can get low-level systems optimizations, like kernel-level improvements, major multiplication libraries, things like that.

Or you can get like [23:51] high level like model level algorithmic improvements you know at the highest level [23:57] To me, it seems like in the last... [24:00] three years, most of the gains have come from hardware level and [24:04] and some from the model level. Do you agree with that? Do you think that's what it will look like in the future? Do you think there's a bunch of juice to squeeze in the kernel level? Yeah, Sean, I completely disagree with you, by the way. Great, that's why I'm asking the question.

Okay, so I think... [24:24] So one way is to look at it as these three different layers. And in that sense, like, okay, from Hopper to Blackwell, which is all we've had over the last three years, roughly 30x improvement on DeepSeq, on the most optimized deployment, which is, you know, you can see on InferenceX, there's about a 30x improvement. [24:39] But, you know, over the last three years, we've had way more improvement intelligence per watt. A lot of that coming from the model layer, right? If you look back three years, it's GPT-4.

Now it's like, you know, maybe like QN, one of the smaller QN models that's like, you know, 27B parameters total and like 2 billion active is like way better. And so you've got this huge improvement on model layer. You've got this pretty sizable improvement on hardware, but it's that co-design layer. And I think that's what's important, right? If you look at the architecture of QNs, [25:07] Any of these models, but DeepSeek is the most famous one, at least, that's public and people have seen it. Yeah, DeepSeek got huge efficiency gains from...

[25:16] co-optimization, your kernel level, optimizing memories. Yes, I think it's like kernels, of course, but it's actually you build the hardware architecture for the chip. So if you look at the shapes of all the experts in DeepSeq, V3, they were all optimized for Hopper. And if you look at for V4, they're optimized for Blackwell and Huawei's chip. And what's interesting is despite the fact that TPUs are objectively an amazing chip, and they run all of DeepMind and they do all the training for Anthropic as well, on the pre-training side at least, [25:46] suck at running DeepSeek, but they are really, really great at running other kinds of models that don't run well on NVIDIA.

There is some level of such deep optimization that has been done, whether it be shapes, network IO, patterns, how you do the collectives, how you do things around the arithmetic intensity of the attention mechanism. All these different things are co-optimized between the model and the hardware and the infra software in between. And it's hard [26:16] I think my understanding is that China has done this a lot better than [26:22] The West? [26:23] The last few years, like, DeepSeek was one of the first models to really, like, do this. I don't necessarily think so.

I think it's more so that the West doesn't tell people what they do, right? Like, OpenAI didn't tell people that, you know, GP40 was how sparse it was, what the shape size was, all these things. But GP40 is roughly the same size, slightly smaller than DeepSeek V3. And 4.0 came out, you know, a little bit earlier, right, if I recall correctly. [26:53] the same rate and the most [26:54] the biggest gains are when you just co-optimize. Yeah, I would say there's been more gains on the model layer than on that co-op than on the sort of software infrastructure layer and the hardware layer.

But there's been innovations on every layer, and really the biggest gain and the beauty of the best labs is when they co-optimize all three. And that's what, like, you know, when Anthropic is, you know, even though they use many different kinds of hardware, [27:18] they don't really inference too much on TPUs. They mostly train on TPUs. And they inference a lot on Tranium and GPUs. And GPUs are more the jack-of-all-trades. But they've optimized their hardware. They've optimized their model. They've optimized everything so that they can do that. Whereas OpenAI, prior models were optimized for Hopper more.

Now they're more optimized for Blackwell. And you step forward through time. These... [27:40] These labs, and the same with Google, right? They've optimized, you know, Gemini 2 was really optimized for the TPU V6E, or Gemini 3 was, and then Gemini, you know, the next Gemini that's coming out is really optimized for TPU V7. And so sort of like a lot of these things are being co-optimized. And actually, when you pull that model and put it running on the old hardware, it's really not that great. And so I think a lot of this co-optimization is the most important thing.

It's called software hardware co-design. [28:10] really exciting about like [28:12] sort of what I think my day-to-day is like, great, you get to look at one layer, there's all these innovations happening here. There's all these innovations happening on every layer. But the real breakthrough innovation is when you leapfrog a few layers, you co-optimize and co-design them, and now all of a sudden you've taken what could have been a 2x here, 2x here, 2x here, and instead of being multiplicative to 8x, it's actually 100x because you've optimized it across all three layers.

And so that's what's really exciting about sort of like what you see at the labs, what you see at a company like NVIDIA who's not co-optimizing on the model layer per se, [28:42] but a little bit from the model layer all the way downstream to, you know, silicon, or you look at a company like TSMC, they're co-optimizing, not just, you know, fabrication, but all the way from the components and the consumables and the tools all the way upstream to what the designs, their chips or the customers are telling them is this co-optimization across many layers of the extraction stack.

There will always be bottlenecks somewhere in that optimization that they're like lagging behind and then need to get pulled forward. [29:08] And band-aids to, you know, a couple of minutes. [29:12] predict like what are at any level of the stack, it can be literally anywhere. [29:17] What are some of the bottlenecks you're most, like you're kind of tracking most acutely the next [29:21] And not necessarily in the supply chain, not in scale, but in terms of the actual [29:26] And. [29:27] And it can be in the supply chain too, but just like...

[29:30] you know, is it, [29:32] memory improvements? Is it that like just [29:36] It's like scaling. Memory is an easy one that everyone's talked about, but I'm not going to talk about it from a supply chain angle. I'm talking about it from a technology angle, right? Memory capacity and bandwidth have been improving very slowly. The NAND cell was invented like 25 years ago. The DRAM cell was invented like 40 years ago. And there's been no major breakthrough in cell, like, you know, what a NAND cell is. Obviously, NAND is like a very simple gate or DRAM cell.

There is stuff that could come down the pipeline that could be hugely innovative. [30:06] All we've really done is make the HBM, you know, more stacks faster. But actually there's like new innovations coming in the next few years where instead of, you know, stacking the HBM separately from the chip, you stack the memory directly on the chip and that makes your bandwidth explode. And so there's interesting companies in that space and interesting POCs that companies are trying to do there. I think like memory bandwidth is one of the biggest. [30:29] For the history of silicon, basically for the last two decades at least, how many watts a chip is can be easily predicted just by looking at it for a data center or desktop chip.

[30:40] it peaks up at one watt per millimeter squared. And so if a chip is 100 millimeters squared, generally the power consumption is around 100 or a little bit less. [30:48] And if you look at the newest NVIDIA silicon, the newest TPU silicon, it's still on that range of one watt per millimeter squared. So, you know, chips are now getting to, you know, 1,400 watts. The next generation is 2,000 watts for NVIDIA with Rubin and such. And you move forward to Rubin Ultra, it's going to be like 4,000 watts or something like that.

But really, there's increasing the amount of silicon. Okay. [31:08] What's exciting is we're now finally doing things, and it's in development right now, where you actually – [31:13] can pump the amount of power into the silicon to be way more, more than one watt per millimeter squared. And now that all of a sudden means you need less silicon. Obviously, it's running at higher power and less efficient in some cases. [31:27] But you reduce the amount of silicon and you're able to like overcome certain problems. Like thermal issues? Thermal issues. There's interference of like electrical interference issues.

There's all sorts of different issues that crop up. And that's why it's a hard engineering problem. That's why we've stuck at about one. But what's exciting is the world is trying to change these things. I think interesting, like in a different part of the supply chain is sort of like, [31:46] People talk about energy is hard, and we have energy bottlenecks. It's like, yeah, but there's actually very simple solutions one could think of. Take the millions of diesel engines for trucks that the S. has the capacity to make. [32:03] You can very trivially convert them to be using for gas in the assembly line.

[32:08] and then stick them up to an electrical motor, like back driving it so the electrical motor generates electricity, rather than the electrical motor causing the rotation of the wheel, for example, but doing it the opposite direction. [32:20] And now you've generated electricity by pumping gas into something that S. can make millions of. And then, okay, well, that sounds like a pain in the ass to service, right, because now you have to have... [32:29] hundreds of these on a data center site. Well, actually, you can just pull people out of car mechanic shops and have them run around and repair truck engines.

Actually, it's actually pretty trivial. I don't want to say it's trivial. I couldn't do it. I think you're making a really good point, which is that because... [32:46] The West wasn't really thinking about, some of those are even hardware more broadly the last 20, 30 years. We didn't have like, [32:52] much innovation. We don't have the best minds thinking about how do you improve these things. Why would you want to go work in hardware when you can sit... Yeah, exactly. Okay, I'm dying to ask. NVIDIA versus TPU. What are your thoughts?

[33:07] I think everyone wants to pick one or the other for this, but it's really a function of, look, you look two years from now, Google's going to make 10-plus million TPUs through their supply chain, and NVIDIA's going to make… [33:19] you know, many more million, tens of millions of GPUs in both [33:23] are going to be $100-plus billion. Well, Google's going to be $100-plus billion of TPU created a year, and NVIDIA will be $500-plus or whatever it is. I'm not making a specific estimate. It's not a revenue forecast. This is just a thought experiment.

Our research does that. You've been media trained so well. Absolutely. Getting ready for the SpaceX IPO. Are you guys big in SpaceX? Yes. Okay, so that makes sense. We're very lucky to be very large investors. Awesome, awesome. So I would say... [33:51] The case of sort of like [33:53] So Google TPUs versus NVIDIA GPUs, they both have, like, points that are really, like, in their favor, right? You know, NVIDIA will be like, oh, we have switches and we're general purpose. And TPUs will be like, well, we're more optimized. We're actually more energy efficient and our network is actually more optimized for certain types of network architectures.

And so you have, like, these counterpoints that both would really get into. And, you know, I could, with a straight face, argue with you, like, that GPUs are way better than TPUs or TPUs are way better than GPUs. But it comes down to hardware software co-design. [34:23] OpenAI's models are headed, it would be a terrible decision for them to use TPUs potentially. And the way that Anthropic and Google's models are headed, it's actually a terrible decision potentially for them to train with GPUs. I mean, it'd be fine for them to train with them.

So why is that? What's the fundamental difference there? There's various things, right? Like the size of the matrix multiply unit is different as a very simple thing. And therefore, the shape of the matrix multiply you do, the attention mechanism you use, the way that attention mechanism is structured, the way the experts are structured. [34:53] are converging to very different model architectures? I think they have quite different model architectures, in fact. Open IIs are much more sparse and that has benefits and then Anthropics are, they're still sparse but more dense in general and that has different benefits and there's many other things, the network topology.

NVIDIA, all of their chips are connected to switches. [35:12] NVLink switches. For Google, they have no switch. But what they've done is they've been able to, you know, NVIDIA, the NVLink can only connect 72 GPUs. For Google, their ICI can connect 8,000 chips at super high bandwidth, but you have to pass through other chips to get there because there's no switch. And so there's like, there's trade-offs there, there's positives and negatives, and that influences the model architecture. It's not necessarily that you should, you know, claim one is better than the other because, you know, you're not going to have a switch.

[35:40] At the end of the day, how do you say that this is better than that when you can't measure them in isolation because it also extends up to the model layer, right? But I remember for a long time thinking, you know, one, the programmability of NVIDIA and then just CUDA as such a big moat. It seems to me that narrative has kind of changed, at least in my mind, for the last three or six months. Like model companies no longer care about if we have to write custom kernels for, you know, this other chip, so be it.

We'll work with four or five chips if we have to. [36:10] good at doing a lot of that optimization work. And so it seems like some of the, and then it's, you know, it's not like there's 10,000 model companies that are each, you know, each need programmability. There's on the order of tens maybe model companies. And so it seems to me that like if you, the fundamental premise of like tens of thousands of big customers that need CUDA compatibility, like it seems that kind of thesis is changing. Yeah. I mean, I mean, [36:40] disentangled because, you know, models are just great at coding and all software gets commoditized in that case.

I do think there is some level of like open source and, um, [36:49] what people call the CUDA mode, [36:51] is not actually anything to do with CUDA, but it's like the fact that [36:55] DeepSeq, Kimi, and Zipu AI, and Alibaba, and Tencent, all these companies. Xiaomi had an awesome model recently. Their models are co-designed for GPUs. [37:06] And therefore, if I want to run them on TPUs, actually, in some cases, they don't run really well on TPUs. [37:10] Now Google just has to create their own open source model ecosystem or open source models themselves.

So they have the Gemma models. And so you end up with like, well, that's not really CUDA as a moat. It's that the downstream product is more optimized for NVIDIA. And in these cases, these companies are open sourcing them. Or like Nemotron is just open sourcing it. And then the users of it, for example, to open the inference API providers, the RL companies that are trying to take open models and customize them for companies' business use cases. [37:40] in video. [37:41] because the ecosystem uses NVIDIA, [37:43] Even though I don't particularly care about writing CUDA kernels because the models are great at that, but it's like the shape of, like, well, this expert, the Dmod is this, and, you know, the hidden dimension, blah, blah, blah, is this, right?

And so, therefore, it's better to run on NVIDIA GPUs than it is on TPUs and vice versa, right? If Google were to actually open source really good models, you know… [38:03] This would be the same thing, right? People would take their models and they'd be like, oh, wow, these don't run that well on NVIDIA GPUs. [38:08] I should actually just rent TPUs or buy TPUs and do it on there. For small teams, you're going to want to use all the open source software like VLMSG, PyTorch, all that stuff. But the big labs, they don't necessarily need to use all that, right?

OpenAI's forked PyTorch long ago, and Anthropic and all these other people don't necessarily rely heavily on the open source implementation of these things. They've forked things or built it on their own already. And so they don't need to rely on the open source, and therefore now it's more like, [38:34] I'll choose the best hardware and I'll co-design my model and infrastructure software through and through for that hardware that is the best and most cost-efficient. [38:43] And, you know, I'll have AI help me write all that software. What do you think of Cerebrus?

I think Cerebrus is a really innovative company. [38:50] I think in some spots of the market, they're really, really good. Very fast inference. I think that's a big market. We use fast mode almost exclusively at SemiAnalysis. I love how disciplined you've been about accounting for, I don't know if that was one exhibit you did or if you do it consistently, but accounting for the dollars spent in the ROI on each task. It's an awesome analysis. Yeah, we do it pretty diligently. So thank you. That was the dark GDP article that we wrote.

[39:20] also track everyone's token spend by day. And if someone's spiked up, I'm like, what did you do? It's like, okay, thank you for telling me that that seems worth it. Cool. On with my day. I think fast mode is obviously worth a lot for high-end tasks. I can just see so many different use cases where super fast tokens are worth it. [39:36] I can also see the flip side where there's a lot of use cases where super fast tokens aren't needed and therefore the market won't [39:43] pay for them, and they'll use GPUs and TPUs instead.

I think the big risk for Cerebrus is I mostly think the best models are the ones that you want to use fast mode on and small models you necessarily might not use fast mode on. I could see that being wrong with financial markets maybe or something like that, like a Jane Street high-frequency trading or something like that or medium-frequency trading. But ultimately, running really large models at really long context is very difficult on SRM-based chips like Cerebrus, like Grok. [40:13] What happens then if the models get too big, right? If OpenAI's model is not on the order of hundreds of billions of parameters or low trillion parameters, but it's actually 10 plus trillion parameters, now all of a sudden, I don't think that that will fit on Cerebrus, right?

And then if that doesn't – with a long contacts length, right, if you have a million contacts length, now that makes it really difficult to justify – [40:33] And so far we've seen the bulk of revenue and usage at the labs be on their best model. Even when the model price has gone up, we've seen that. There's some data that shows that even though Fable just released today, they've had incredible amounts of people switch to Fable and Mythos, sort of that next-tier model, even though it's way more expensive. And so – And that's volume by dollars totally, but without volume by tokens?

Well, I guess who cares about volume by tokens? It's about the dollars. Fair enough. Right? [41:03] I don't know, 200,000 Mini Coopers or Toyota Camry sold if, I don't know, Ford out of 150s are 5X ASP and they sell only half as much. Okay, fair enough. And therefore the most lucrative market is pickup trucks in America, right? Mostly being facetious, but like. I do think this is one of the things that you've done so well and differentiates you from almost everyone else is that you care so much about the economics in addition to the technology.

And I think very few people bridge those two things well. Yeah, thank you. [41:33] I think it's really fun inside of Semi-Analysis because we have 90 people and like a big chunk of them are technologists, engineers across the whole supply chain. And then a big chunk. [41:42] is people who are formerly at hedge funds. And you see these arguments, like people are like, oh, well, that doesn't matter. And it's like, then someone's like, well, but cost. And then the engineer's like, no, no, no, but this technology is the coolest. And you see this organically, like, fight it out.

And we're pretty informal. And, you know, given the fact that I was a forum moderator, you can imagine what this internal insight looks like. Well, you're enjoying it. You're enjoying it. You don't wrestle with the pig because the pig enjoys it very well. Exactly. Just on this topic, before I go into the next question, [42:12] here. [42:13] Like, [42:14] trigger topics in semis for you? You know, like if someone's like, which is like such a meme, you think this person must be a moron. Like if, you know, if it's like, oh, you like memory is the bottleneck.

I mean, it's true, but like, um, I think, I think moreover the one that really gets me is people are like, AI has no ROI. It infuriates me, right? Like there's like, what's the ROI or like denying model progress, right? There's these people that are like models aren't [42:44] They're going to dead end and plateau. And it's like, bro, the line has been up and to the right in terms of capabilities this entire time. And they're like, look, this benchmark didn't improve. That's because it said 90%. Look at the new benchmarks.

Yeah, you saturated. Now they're skyrocketing, right? It's like I think that's more so the issue and challenge. Like I think semis are really complex, and I don't fault people for – [43:06] lacking understanding of it. I learn stuff every day about the semiconductor supply chain from people. And I've been studying it for arguably 18 years since I started moderating the forms when I was 12. Arguably been studying it for that long. But even then, it's like live, breathe, and that's all I care about. But there's so many layers of the abstraction stack.

It's like [43:28] Like I learned about a new chemical that does like $100 million of sales like yesterday. And I'm like, whoa, didn't know this one existed and what process it did. But it's like, you know, you learn about things all the time. It's like, okay, $100 million sales in a couple hundred billion dollar industries, whatever. But like, you know, it's like. But it's essential. It's essential. And it's like actually every chip requires it. It's like, wow, I guess there are a thousand process steps. And, you know, it's like, oh, yeah, you like semiconductors name every process step.

It's like, no, come on. [43:58] and then they get the conclusion completely wrong. And that's- That happens in our job all the time too. Yeah, yeah. I mean, I think my attitude is not to be mad that you do that. It's to do it as fast as possible. I think the industry, because it's so, it's just like AI is the most important thing in the world right now. And there's so many near-term bottlenecks. We talk a lot about the near-term. Are there longer-term things that you're really excited about? Like say on a 10-year time, we talked about orbital data centers, but like Silicon Tonics,

[44:28] rated or overrated on a 10-year time frame? Are there other things that on a 10-year time frame? Yeah, I mean, I think space is like super crazy awesome in the 10-year time frame that I'm, you know, for space data centers and all these sort of mining asteroids and all these things, which is, you know, super excited about the vision of SpaceX, right? Again, not investment advice before you hop in. I think on the semiconductor side, tremendous market movements and tremendous, like, things can happen just when, like, things happen one year later or sooner.

And so that's all, like, technology that, like, you know, in terms of, like, co-packaged optics, [44:58] well, everyone knows it's going to happen by the end of the decade. The debate is like 27, 28, 29, 2030. But some point along there, it's going to happen. I think the more interesting thing is like there's companies like – did you guys invest in Naveen Rao's company? We did. Okay, yeah. So I think like he's trying to innovate on like the silicon layer, on the software abstraction layer, and the model layer simultaneously. And he fully understands that it's not like we're going to do this in a few years.

It's not a two-year timeframe. Yeah, it's not a few-year timeframe. It's a long-term bet. [45:27] And like stuff like that is like, okay, we're going to bring like potentially like analog compute with energy-based models and like all this crazy shit all at once. It's like that's exciting. [45:37] Probably won't work, but, you know, that's exciting, and I, like, really look forward to it. And it definitely won't work quickly. Yeah, it definitely won't work quickly is what I should say. I believe in Naveen, and, like, you know, I met him very, you know, I think he's one of the first people I met in the industry, funnily enough, like, in 2020 or 2021.

Actually, 2020, guys. It says something about him. I think he's someone, in my experience, he's always trying to. I baited him on the Internet. I baited him on the Internet. He's always trying to help the younger generation. [46:07] He was also so ahead of his time with Mosaic. I remember getting pitched up. No, it was 2019. I was still anonymous then, actually. I baited him on the Internet, and he started replying, and then I just took it to DMs, and then took it to a call. And, like, that was the first person who was, like, really important that I talked to in the entire semiconductor industry.

That's funny. But, yeah, sorry to interrupt. That's funny. What do you think is the end state of the ecosystem? Like, do you think every lab, every hyperscaler just has its own chips? Like, Tranium seems like it's now working, right? [46:37] we end up with every lab, every hyperscale has its own chips, at least for inference, and then maybe for training, you go to NVIDIA or whoever? What do you think is the end state? [46:44] I think everyone will try and they won't stop trying. [46:48] I think ultimately... [46:51] Supply chains matter.

What technology you can bring in matters. More and more as the industry gets bigger, supply chain diversification happens. Right now, everyone's chipped. [47:02] More or less looks the same. It's a big logic compute die in the center, and there's some HBM on the right and left and on the top and bottom. Top side is networking, and then the bottom side is PCIe and other O. And that is the exact same structure for Tranium, TPU, NVIDIA chips, and most of the startups. Not Grok and 3Bris, those are doing weird shit, but that's cool.

I think, like, [47:24] As you step forward, we're going to get more bifurcation of hardware architecture and model architecture, and therefore people are going to co-optimize them. And, you know, some of them will end up in local minimas, right? You know, as we're, you know, if it's like gradient descent, like people are trying to go to the most optimized solution, some people will race to a local minima. [47:42] And then the question is, how do you scoot back over to the absolute minima? And to some extent, NVIDIA will always be more general purpose than anyone else's chip in general, at least on a parallel AI compute basis because they have so many customers who care about different things and who will always give them feedback in the design.

The minima will always be better than them, but is that minima a local minima? Like is the TPU or Tranium or Grok or Cerebris or whoever's design – [48:10] optimize awesomely for here, but in the end state, actually, you've got to go over here, and so they're wrong. And maybe they make a great time, they're great for a little bit of time, but then they end up being wrong. It's like, that's the real question. [48:22] And so I think there will be a big market for general purpose AI compute because you talk to people at labs, they don't even know what architecture they're going to be doing in a year.

[48:31] Like, right, like they literally don't know what architecture they're going to be doing in a year. They have bets. They have many research bets, and that's this exciting thing. But they don't know where it's going. Generally, they, like, know what hardware they have, and they're trying to co-optimize. But ultimately, like, if a new breakthrough happens on model architecture, it's like, just replace the attention mechanism with something else, right? Who knows? Or, you know, all of a sudden, you know, something happens. [48:52] the best hardware will change. And therefore, like...

[48:55] Are people going to make five-year investments on hardware solely on, you know, an ASIC that is more specialized, or are they going to have some bucket of more general-purpose compute? [49:05] And so you see this with like, Google's paying $11 an hour per GPU to XAI for GPUs, right? Like that's insane, right? That's a very high amount of, obviously compute is limited and so on and so forth, but it's like very like... [49:19] But at the same – despite the fact that they have TPUs. And so there's some questions there like why do they do that?

Google actually has three different design programs for TPUs. They're making a TPU with Broadcom. That's a different architecture than a TPU with MediaTek. That's a different TPU than the architecture that is – I won't disclose by research. But they're making different architectures. It's not just like – [49:39] oh, they're making TPUs with a couple vendors, and it's the same architecture. It's different architectures. And the third one is a very different architecture from the first two. And so I think people recognize that the local minima can happen, and therefore I think everyone will have their own ASIC program.

I think everyone will deploy billions of dollars of their own ASICs, tens of billions of dollars. In the case of Google, hundreds of billions of dollars a year of their own ASICs. But ultimately, they're also going to have workloads that don't use TPUs. Some of the Google bets that are not Gemini DeepMind [50:09] TPUs. They don't use TPUs. Some of them also primarily use TPUs, right? It's a bit of a broad thing, but like, you know, maybe for drug discovery or for Waymo, you might not want to use TPUs. I don't say which one it is, but like, you know, there's different architecture bets and different paths for AI.

AI for science may have different algorithmic patterns than general intelligence AGI models. And so I think we'll see diversity continue to proliferate. Yeah. And because the market has gotten so big, niches will be carved out. And so that's [50:39] It makes it possible for companies to have their niche and actually make money, even if the majority of the pie goes to NVIDIA and TPU and Tranium. Okay, love that. Can we talk about the data center build-outs? One, it seems like, I mean, by all accounts, if you look at the charts, like dollars per compute hour, we are in the middle of a crazy compute crunch.

And it seems like it's both a demand and supply side crunch, right? Like demand for long-rise agents skyrocketing, supply, all these data center build-outs are delayed. Do you think this is a compute crunch for the foreseeable future? [51:09] alleviate at some point. Yes. Every quarter we're deploying vastly more compute than the prior quarter. And there's more data centers built in the prior quarter. [51:18] This year there's going to be 20 gigawatts, even accounting for the delays. [51:22] And next year, there's going to be more than 30 gigawatts accounting for the delays.

Of course, delays happen on everything, right? Anything hardware can have a delay. That's just the reality of life. Are we going to have a compute crunch for the rest of our lives? It depends on what happens with models. But the TAM for Mythos, Mythos 5, Fable 5, is not just like 2x that of Opus, right? The model is so much better, and it can do so many more tasks. The TAM for it is way larger than that. [51:52] Right. From, you know, Opus or maybe like seven or eight months, it's Opus 4.5 would launch to now huge, you know, 4.6, 4.7, 4.8 were improvements.

But Fable and Mythos were like a huge step function improvement. The world's compute did not double in that or quadruple or whatever in that same time frame. [52:07] But the demand for useful tasks that can be done by AI, the number of useful tasks and the value of them that can be done by AI has. [52:15] And so now the question is, [52:17] What happens? Well, obviously, Anthropic in Q2 is profitable. They're net income profitable, excluding stock-based compensation. [52:27] And I think by Q3, they may even be profitable, including stock-based compensation.

That's, like, how profitable they're getting. And their margins on an Opus token, at least Opus 4.8 token, is, like, [52:38] North of 80% for the API price. They've got a lot of deals where their total corporate gross margins gets clawed down a little bit because of how they do bedrock deals and vertex deals and things like that. But ultimately, their per token margin is so high.

Want to learn more?

Ask about this episode