SambaNova CEO on Raising $1B at $11B: "It's a Land Grab Right Now"
Rodrigo Liang is the CEO and Co-Founder of SambaNova. The company just announced a first close on a $1B round at an $11B valuation, led by General Atlantic with T. Rowe Price and Capital Group participating.
Appears in
- Uploaded
- Uploaded Jul 17, 2026
- File type
- POD
- Queried
- 0
Full transcript
Showing the full transcript for this episode.
We just did the first close of a billion dollar fund raise at an 11 billion valuation. I've been in this industry for 32 years, building high performance trips for a long time. I've never seen the interest in semiconductors higher. Now what you're seeing at scale with Anthropik and with OpenAI and with Gemini, we've got millions and millions of people using it every day. We released SN40 a couple years ago. It became incredibly popular because instead of 130, 140 kilowatt rack of NVIDIA GPU, We were outperforming it with a 10 kilowatt SM40 rack.
We could take a trillion parameter model and run it in a single rack, where we take dozens of racks of other people's equipment to run the same model. We're $2.5 billion raised in the history of the company, and there aren't really that many companies that have raised into the multiple billions. It's all about scaling. It's all about who can get to scale faster. Rodrigo, welcome to Sorcery. Thanks for having me. Well, we're here for context. We're in Paris right now for the RAISE Summit. And right now we're sitting right in front of where the conference is.
I don't really actually know what this park is called, but it's next to the Louvre. Yeah. I don't know if you know. Yeah. You've been here a bunch, right? I've been here, but I'm not sure if I know exactly the name of the park. We're right in front of the Louvre. Well, you have some big news. I think this will come out about a week after the news drops, but we'll still make a clip on that. So what is the big news? Well, we're super excited. We just did the first close of a billion dollar fundraise at an 11 billion valuation.
This is a great show of momentum for the company and great show of support. The round was led by General Atlantic with a number of incredible investors that came in. Seligman Ventures, T. Rowe Price, Capital Group. These are all significant American investors that are coming in. That shows that the company's got momentum. We're driving towards the scale and a significant amount of capital infusion to help us do that. All the energy right now is going into semiconductors. I'm sure this was a very hyped up round in some way or another.
Maybe it's been faster than others. What was the process like for you? I've been in this industry for 32 years, building high performance trips for a long time. I've never seen the interest in semiconductors higher. And I think it's a realization that chips at the center of this transformation, if you look at what what what AI is doing in the world and the build offs of the data centers, you can't do it without chips that run and run efficiently. And so with Salmanova, we're coming in and providing technology that is able to take it to scale, take it to a level of inference scaling that's just really not that practical to achieve just with traditional GPUs.
And so I think the world sees that and the excitement is coming in from some of the top investors in the world. So where we are at today is inference. Inference has really taken the stage and it's been the next evolution of computing and where everything's going with AI. So for people that don't know SambaNova, can you walk through the products and how you've evolved them for inference? Yeah, I mean, look, with AI, you've got a sophisticated audience, so they know. With AI, there was always a training and the inference.
There's no point of training a model if you aren't going to inference it, if you're not going to use it. And so the example I use with people is you don't go and invent a search algorithm if you're never going to do search. And so we're not going to train a model if you're not going to use it. And now we're in the phase of using these models. We've always used them. We've always inferenced them, but it was still research to train models better and better. And so when we started the company in 2017, we're very focused on how do
we actually lower the cost of training, right? And back at the time, we're training models for image recognition. Can we tell the difference between dogs and cats? And, you know, can we recognize voices? Can we make voices? You know, we're doing all that research. But really, in the end, inference wasn't really a problem yet because the number of people using it were very small. was people trying to test the model that they trained. Now what you're seeing at scale with Anthropik and with OpenAI and with Gemini, and you're at scale, you've got millions and millions of people using it every day.
And so now you have the problem that Sumitava was originally focused on, which is around efficiency. How do you actually deploy at scale so the whole planet can use it without burning up the planet, without kind of running out of data center space, without blowing up your infrastructure cost? because at scale, the number of chips deployed for inferencing will be orders of magnitude greater than whatever you're doing for training. And so walk through some of your chips. Yeah, so we started a company. So we're on, we've taped out six chips in the last seven years, and we'll tape out seven for the next year.
We're really excited about Generation 5 that's shipping later this year. When we started with SN10 and SN20, these are the early chips. They were really focused on training. And they're focused on, can we train those models faster with fewer chips? As you know, some of the largest AI labs are deploying thousands of chips just to train one model. And they were trained for months and months and months. And so we're focused on training. As the world moved along, it became very clear that the bigger challenge, the bigger problem was inferencing.
by that once you were at scale, how do you actually deploy those same models you trained on all of these different data centers for people to use? And now you can say, I need a gigawatt data center somewhere in West Texas. That's one way to do it. But how do you serve all of the countries, all the planets, all the different users that are worldwide? And so you need to think about power. You need to think about data center. You need to think about latency. And these are all things that we started taking on.
And so by the time we released SN40 a couple years ago, it became incredibly popular. Because instead of a 130, 140 kilowatt rack of NVIDIA GPU, we were outperforming them with a 10 kilowatt SN40 rack. But in a 10 kilowatt SN40 rack, now suddenly, and it was air-cooled, you didn't need liquid cooling upgrades. suddenly every data center that's around the world that you're using traditional CPU, traditional storage for, you could just roll into someone of a rack, air-cooled, and you got state-of-the-art inferencing faster than on an NVIDIA GPU. So that became a really, really popular way to deploy, especially if you don't want to put hundreds of racks of NVIDIA GPUs where we collapse the footprint because we could take a trillion-parameter model and run it
in a single rack, where it would take dozens of racks of other people's equipment to run the same model. And so that's kind of what SumNova is really focused and known for is really driving premium inference, the largest models that are very low cost, very low power, and delivering ultra high performance, which becomes really, really valuable if you're actually starting to deploy these incredible models at scale. How has rack composition changed? If you look at kind of for training, one of the challenges that you have is you have to aggregate all of these racks because you need thousands of gpus together and they have to work in sync right as you're training each loop you're you're working sync and the problem there was that if any one fails the whole cluster fails right and so you have to put these checkpoints you have to kind of stop you know every so often to make sure that you store the progress you made up to that point in case the next cycle fails.
And so there's all of this work. And then you need really, and you talk to tech folks all the time, you need really, really high-performance networking to connect all those things through. And so your networking equipment becomes really, really expensive. You need a lot of memory. You need a lot of software coordination and things like that. As you go into inference, the beauty of inference is it's scale-out. And so basically you're adding racks as your users grow. And so with Samanova, that minimum quantum is down to one rack. Where if you have other service providers, just to run, say, a DeepSeq model, which is now one and a half trillion parameters, just to run that, the minimum for some of the other providers might be 10 to 20 racks.
And so you're starting to think about, okay, well, if I want to be an inference provider and I want to run these large models. My minimum cluster just to start serving is 20 racks and the cost stroke is high, the power needs are high. Some of it, we can actually reduce the minimum down to a single rack. And so now you just grow as your user base grows. Significantly more efficient to deploy, significantly more flexible if you're going into environments where you don't have gigawatt. You can go into your data centers.
Here in Paris, downtown Paris, you can find an existing data center and deploy five racks, 10 racks there for ultra low latency where the users are. And so this is kind of what's changed in the racks is that we're very focused on bringing it to broad-based deployment, standard everything, standard 19-inch rack, standard air cooling, no complicated liquid cooling retrofit in the data center. We're using standard Kubernetes, standard Red Hat Linux, standard Ethernet at the top for networking. We don't have to use all this kind of really expensive networking equipment to gang these things together.
And so that allows people to go in to existing data centers, roll this thing in, pull out the old gear, and you're up and running with new services, which otherwise might take you nine months to a year, maybe as long as 18 months, to build a gigawatt data center, to bring liquid cooling in, to bring all this new power. And sometimes you have to figure out how to secure new power and new energy and nuclear power plants. And all those things that people are talking about, it just takes a lot of money, a lot of time, but you don't have to do it if you use some of the equipment.
That's pretty good. Yeah. I mean, it's operating at the velocity of inference, which is people want to stand these things up quickly. Right. It doesn't take months to inference because these models have been already trained. Whether that's an OSS model from OpenAI or Mistral in France, we've done it on Mistral, or you've got the DeepSeek and Minimax models, amazing models that are out there, or the frontier models or the closed source models that people are offering. It doesn't matter. You deploy this technology, you can bring those models in immediately, and you're up and running
with the latest and greatest AI models in the world, and then you can upgrade them as you go along. So do you buy the headlines that new data centers should be 50 or 100 billion dollars? Well, I think you're going to have some data centers like that because I think there's still going to be large scale deployments, large scale access that people want. And I think you're going to find that the world is going to be heterogeneous, that there's going to be those large data centers that need people go and secure a lot of capacity for some of the things that they want to do.
and I think you're going to see this new wave of companies that are doing distributed data centers. So these data centers are mid-sized, and it's going to be even more important as you go into this agentic world because in the world of agents, you're not dealing with a single model and a single prompt. If I go to ChatGPT, we're here in France, and what should I do if I have an extra day in France? You can talk to ChatGPT and generate an itinerary for you. That's between me and the model, and it produces the result.
In the world of agents, these agents are orchestrating within themselves without us. My problem starts in the beginning. Ten agents are all intercommunicating, and each of them taking some amount of time. So if you actually have a lead time or response time, let's say two seconds, which for a one user to one model is not very long, right? For our eyes, it takes longer than two seconds to read the output, right? And so that's okay. In the world of agents, where you have, say, 20 agents orchestrating with each other, if each of them takes two seconds to respond, now I've got 40 seconds.
The end user did an initial prompt here, like, please move some money from my bank account, my Bank of America bank account over to PayPal, and then give me a report of all my expenses over the last six months. Okay, that's my prompt. Then the orchestration of security and balances and all of that has to happen. and the output. If each of those 20 agents took two seconds, that's 40 seconds. You've already given up on that prompt, right? So the response time by the end of the user, the user's expectation is say one to two seconds divided by 20.
It's less than a second per, it's 0.1 seconds per. And so your response time is going to be really, really important. And that's why latency matters. And so you're now seeing us coming in and saying, look, we're going to deploy the hardware, where the users are in large metropolitan cities, right? Because that's where business is being run. And so latency is really important. So we're going to deploy that. Well, unfortunately, in those large metropolitan cities, you don't have those gigawatt data centers. There's no space for it. Manhattan, in Paris, right?
Where are we going to find space to drop a, what did you say, how much money did you say? 50 to 100 billion. I mean, where are you going to find even the space to build that? You're going to find some space in some places to build that, but in terms of ultra low latency in the big cities, you're going to have to find smaller quantums, smaller spaces that allow you to deploy what you need for the users there that require that really, really low latency. And thinking banking, healthcare, like there are many use cases where your latency is really, really important and you're just not going to want to wait.
On the topic of inference, what is premium inference? The way we define premium is ultimately the highest value use cases. And there's two dimensions. It's basically on one dimension is the size of the model because it's about accuracy. Okay. And so years ago, we used to talk about hallucinations, right? You talk to Jackie T. and come to us and say, well, what was this? That was a fun time. I think, wait, we should probably look back. Like hallucinations was a fun time. There's less of it now. There's less of it, right?
These models are getting pretty good and there's less of it. And yet, accuracy is still incredibly important. And so why are these models going bigger and bigger? It's not that people want to spend the hundreds of millions of dollars to train I mean they fighting for that bit of accuracy because you look at a model like Claude Anthropic right You look at that model Why did it become the most popular code generation model? When software developers go and type, it generates really good code. And you take a different model, sometimes not as much.
And this is where also the open source, the minimax model that's out there as an open source model became very popular, incredibly accurate when generating code. And so these models still are being valued significantly for the output they generate because if you can trust it to produce good output, you don't have to invest as much human energy to go double check it. If you have to go double check it, then certainly you're starting to invest more time. So on one dimension, premium is how accurate the model is. And today, that's proportional to the size of the model.
And so you have models that are very, you know, Lama 8B, for example, is an 8 billion parameter model. Very small. By 70B, pretty small. 70B used to be the big. It's tiny today, right? When you had CHAT-GPT and GPT-5 at 5 trillion, the new models are heading towards 10 trillion. Even the open source models are already 1 to 2 trillion parameter models. And so now you're starting to see these models getting very big because people are looking for accuracy, right? And they want a model that handles a broad range of things but handles it correctly.
And so we're very focused on making sure that we handle the largest models well. And then the second dimension of premium is learn it fast, right? For the reasons I just described about agents that, you know, you don't want to take a long time. We live in an impatient world. Yeah. You and I, I mean, a few seconds, we're starting to tap the phone, something happened, right? And so if the service behind it is agentic, you need to actually run really, really fast. There's not time for you to actually wait.
And so that combination of running really big models, as you know, they don't run fast. Or you look at services like RockCerebus that run fast, you can only run the small models. So how do you find the ones that run really fast on the big models? And that's where some of the strength is. you know that we take the biggest models in the world and run them in the original precision we don't you know we don't quantize we don't you know quantizing is you know you chop half the you know weights off near so so we don't chop the model down we just run original precision full precision we're in faster than anybody else this episode is brought to you by brex my favorite you become what you spend on
and i refuse to spend my time on work that shouldn't exist expense reports receipt chasing and manual closes the companies building what's next from versell open ai anthropic granola and deepgram all made the same call they all run on brex brex is the intelligent finance platform that combines cards expenses and banking into a single stack with agentic finance built in ai agents that handle expenses automatically enforce policy before spend happens and close your books in minutes. That's why Sorcery runs on Brex, so I can spend time on building and not busy work.
It's time to get Brex AF. Learn more at com slash sorcery. That's b-r-e-x dot com slash s-o-u-r-c-e-r-y. Bye. At what point do you think speed is going to bifurcate the market much more into different pricing units? I know that token costs and everything is a really big topic right now. But for consumers, we do expect speed. But should we be getting enterprise quality speed? Well, I think you're going to find that people pay for it. And it's always been true. If you look at the internet, people paid. Initially, they would pay a premium for the upper end of the packages for faster internet.
And then some people would still remain on the basic, right? And same cell phone. But I'll say this. Look, when 5G showed up, nobody's signing up for 2G, right? If your phone, if your cell phone is not transferring very fast, you're getting very frustrated, right? So I think on the curve is over time, as fast inference is broadly available, who's going to want slow, right? And so it's going to not only be the premium today, where kind of the people who need it, co-generation people, real-time banking, real-time healthcare, there's a number of industries where speed matters because it's their livelihood, they're going to pay up for it.
That's kind of the premium service that we're offering. But over time, what you're going to find is that all of us are going to want that. Like, I don't see a use case where the average population, either consumer or enterprise, is going to say, actually, I prefer this low. That wasn't the case on the internet. That's not the case on cellular, you know, on mobile data, right? It's never been the case, let me pay more for this low, right? Or even let me pay a little bit less for this low, right?
Most people over time are going to say, no, I want the fastest, right? And so as the cost of delivering fast goes down, you're going to see most people switch over to the fast. And this is why I feel like, you know, the premium inference, which is large models, which equals the most accurate, most accurate models and fast, ultimately steady state is what everybody's going to want, is whether today we can offer it at a low enough cost that everybody can afford it. And so until then, you're going to see people very quickly moving over if they have a need for that fast inference and have a need for the accuracy, which is, I think, a large part of the enterprise.
There's also a couple other macro themes that are happening that are just going to explode data and usage across the world. Yeah. One of them is Starlink. I don't know if you've been flying around on planes that have Starlink, but even having access to that in remote areas, it just increases the amount of work you can do and work for edge cases. Another part of that is edge computing. Do you guys deploy in edge remote areas as well? Well, what happens here is, again, this is tying into your question, is a $50 billion data center ubiquitous?
ubiquitous. It's going to be hard to say that when you have economies that can't afford a $50 billion data center, right? And you're going to see across the world. If you believe that AI is going to be a technology as pervasive as internet, which I do, right? That this is something that everybody on the planet should have access to, right? And so mobile service, internet, AI, everybody on this planet should have access to it. And so if that's the case, then it's not going to be equally deployed because in some countries you can't afford to put a hundred billion dollar gigawatt data center somewhere and then let tens of hundreds of millions of people come in and use it or you might be in a place where you know you don't have that and you need something significantly smaller and today because of somewhat of technology being as little as 10 kilowatts per rack we can put them inside shipping containers right so we build out these data centers clusters of 10 20 racks inside the shipping containers that you see on these ships, right?
And then you deploy them in these edge data center use cases in a much, much more cost-efficient, power-efficient footprint than having to build out this liquid-cool gigawatt data center with its own nuclear power plant, right? And then you can put, you know, a startling connection to it. You know, you can bring the internet in. You can actually have solar farms next to it, you know, and then you can actually power that. So there are lots of different ways that people are creating these data center clusters that allow their communities, whether that's in regions or in industries or in sectors, to be able to get access to the best models.
Yeah, one particular company I was thinking of is Armada. I don't know if you know Armada, but Armada, they deploy modular data centers on the edge. And so, I mean, it's not just in remote communities, but it's for critical industries, critical industries that don't have that real-time data that they've had before. So it's like oil and gas, it's mining, it's oil rigs out in the water, all that kind of stuff. And even for military and defense and that kind of thing. Yeah. Look, Armada's a partner of ours. They've been a partner for a few years now.
Dan's a good friend. Oh, yeah. Yeah. That's the same concept that if you can actually take, instead of having to put a 100 kilowatt NVIDIA rack in there, you can put a 10 kilowatt sum of the rack and generate more tokens, right? That becomes very valuable because you can actually fit a lot more output in a smaller amount of space, a small amount of power. You then deploy it into regions where you don't have the traditional data center available. You are in remote areas where you want AI to actually manage your operation out there.
And, you know, oil raids and things like that is an example, right? And so it's an incredible opportunity to actually get this type of technology deployed in different vehicles in different ways. Because, again, access to them will be quite varied, right? It won't be just in that very large $100 billion data center for them. talking back into the data center and the racks you're now working pretty much with other chip companies in a way that like you weren't before so how has that evolved is it is it weird do you think that's going to last the world of competition you've got you know partners you know that uh competing in certain
areas and collaborate in other areas look at the the core tech industry is incredibly small. But who is building tech? But if you really look at it, how many people are truly building and deploying chips? How many people are truly building and deploying systems and building and deploying racks and building and deploying data centers? There's actually not that many, right? If you really look at it, in the construct of the entire economy, right? And so we do collaborate. But here's what I'll say. In the end, we're all in service of customers, right?
And so customers are coming and telling us, Look, we have these NVIDIA racks. We have Xeon racks. We have AMD racks. We have other chips. You know, and how can you actually get my total business operating more efficiently? Right? And now that the world's going to inference, I'll give you this example. If you look at any service provider, like inference cloud provider, they've got NVIDIA racks sitting there. And it's doing all sorts of things. Right? Let's say a thousand racks. You're doing training some models from some customers. You're influencing models for other customers.
You might be doing HPC, right? You may be doing some biology, some physics. Who knows, right? Games, you're doing gaming. You're doing all sorts of different things. And you are then, you bought those racks and you're trying to monetize. So someone comes in and so we are focused on things, right? And we know that when you're influencing these models, we run it significantly faster at a fraction of the cost, right? And so if I go into that data center, the inference has arrived, and you look at the percentage, 70%, 80% of those racks are running inference.
And so why would you run inference on those racks when you can run it at a fraction of the cost at a higher performance on Suminova? And so route that traffic to Suminova, frees up all these racks for you to resell to all these other things that people want anyway. And so the economics starts getting much better because now without buying more hardware, They generate more revenue. They actually run the most popular models in a much more efficient rack at a much lower OPEX and a much lower CAPEX. And they're generating faster tokens, which then you can charge more.
Right. And so your premium service is able to charge more for faster. And so you're making more money. And in general, they're just lifting your margins on the same hardware infrastructure. And so that's usually kind of what the customers are asking us to do is how do you get our business actually generating better margins? because for them to sustain themselves, as you know, today, infant services, they're not making enough margin. You're generating lots of revenue, but you're not generating enough margin. And in order for them to sustain, they've got to be more profitable.
And this is all the investors. You mentioned a couple of investors you're talking to. They're all thinking about, well, how do I sustain this? Well, you sustain it by having every service provider make more money. If they're making more money, they can continue to invest. And what we do is we generate more margins by giving better inference service at a much lower cost. How are customers measuring that difference and how are they measuring the routing performance? Well, the routing's already happening today, right? So I'll start there. So if you look at a large data center, let's just say you're running open source models, closed source models.
You're running some HPC. when a request comes in, you are already routing those models to certain racks. And so these racks are running Enthropic or these racks are running Minimax and DeepSeek. And so that routing is already coming in. You say, hey, I want to run my service over to you, do a prompt router over to a Minimax model. It will already go over there. It's automatic. Is there like software in there? Like how does that work? Well, there's software that's on top And so if you look at kind of what these API services are, and this is one of the beauties of what the open
AIs and Enthropics really set up is they all set up these open standard interfaces. And so there's these API calls that you will run, and these are standard API calls. And some of them, we actually match that as well. And so that when you prompt, actually it will look for a particular API to a particular model running on a particular IP address. And so once you actually have that, then it's a standard interface. So when you're deploying your racks at a NeoCloud or Hyperscale Cloud, you just have the same API interfaces.
And then you can leverage kind of what's out in the open community for routing. And so you were already doing that before. So if I go and I'm a customer of GPUs, for example. The GPUs have A100s, H100s, B200s, B200s. Even the different versions of it, those are being routed too. because if I pay for the newest chip, I don't want to be routed to a no chip. And so same thing with now you can just stand up other chips. You can stand up EMD chips and some other chips and other chips next to it, and they're all just already part of that ecosystem for routing.
So that kind of the routing But most service providers and I think before we started we were talking about KPIs and economics and how do you measure And this is the way we see most providers measuring They purchase per rack. They operate per rack. And so they want to generate revenue per rack. And the revenue is generated per token. If I put a rack of hardware, I'm just seeing how many tokens I might generate in a particular model, and that model has a price per token. Multiply that by 30 days per month, 24 hours per day, number of tokens per second, and you can figure out how much money that rack is generating, and you look at how much it's costing you to operate.
So that's as simple as that. And so that's what we focus on. We're very focused on making sure that when you deploy a rack of summed over, you generate great margins relative to the model, You can change the model because different models and different pricing per token, but you're still generating significant number of tokens per month so that you're making profit on that rack that you actually spent and operating on a per month basis. Yeah, to your point, I did wonder this because with some of the conversations I've been having, a common theme is there are no standard definitions for anything.
You're like, well, it's very clear with hardware. This is how you do it. And I was like, oh, well, that is the clearest definition. It does get a little murky when you go into like, how do you determine what an AI agent is? What does that even mean? What does a jailbreak mean? We just talked about this with Dylan Field. And it's just so interesting. And even way back when, it was like maybe a couple of months ago, I interviewed Teresa Carlson. She's CEO of the General Catalyst Institute in C.
She used to
Want to learn more?
Ask about this episode