Anthropic's Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden
Katelyn Lesse and Angela Jiang lead the team building Anthropic's developer platform - the layer that both outside builders and Anthropic's own products run on top of. Angela frames the platform as a three-layer stack: knowledge, execution, and coordination. She argues the real leverage is what’s at the top: "strategies," or meta-harnesses that give each token a different job, from advising to executing to reflecting to memory. On the question of open ecosystem vs. walled garden, they say they aren't precious about owning the stack. Katelyn points to Anthropic's self-hosted sandboxes with partners like Modal, Vercel, and Cloudflare. Whether the work runs on Anthropic's infrastructure or someone else's, what really matters to them is that the architecture is sound. The deeper bet is standards: they hand skills and MCP to the whole industry, build connectors on the MCP spec, and help agents (Claude and non-Claude) work together. The one place they stay closed is model routing: they argue harnesses should be tuned to a model family, so they're designing for Claude rather than routing across models. Angela's frame for the ecosystem bet is electricity: transformative only because everyone could plug in, and no company wired it alone.Hosted by Sonya Huang and Lauren Reeder, Sequoia Capital
Appears in
- Uploaded
- Uploaded Jul 14, 2026
- File type
- POD
- Queried
- 0
Full transcript
Showing the full transcript for this episode.
[00:00] The last layer of abstraction on top of this is probably the coordination layer. So you have knowledge and you have execution and you have coordination. And at the coordination layer, we're beginning to think of these things called like strategies, where basically it's almost like a meta harness. The true low level harness is designed for execution. But the next one is about, OK, if tokens aren't really fungible and you need to give them different jobs, like maybe this token is advising versus this token is executing. You want to start composing these like these kind of orchestrated strategies that go together.
[00:30] execution still needs to know what to do. So everything in theory should kind of like ladder together. And so I think, you know, if you were to look at our roadmap and maybe kind of project forward a little bit where you kind of expect this to go, we'll move more and more from the knowledge layer to the execution layer, from the execution layer to the kind of coordination layer in terms of the abstractions that you can see us put out. [00:47] *music* [01:04] Caitlin and Angela, thank you so much for joining us today.
Lauren and I are thrilled to have you here. [01:08] You are responsible for building Anthropik's platform, and so you are responsible for building what I think is one of the most important, if not the most important, developer platform in the world. And we are really excited to interview you today to understand more about what's ahead. And so maybe just to get started, can you give us the context of, you know, what is Anthropik platform and where do you sit within Anthropik? Yeah, so platform is both our externally facing APIs, our developer platform that people build on top of when they want to build applications and systems that access Claude's intelligence.
[01:38] as well as internally we run our product infrastructure. And basically we're the layer that our apps are [01:46] build on top of internally as well. [01:49] Awesome. What's your North Star as a team? [01:51] It's a great question. We actually, because we have both internal and external, we actually kind of have like two North Stars, which is probably like, you know, you'd be like, well, you should only be one North Star. But no, we kind of planetary system. Yes, exactly. There's separate solar systems, so it's fine. But on the internal side, like we really want to provide is like literally as much leverage as possible for our internal teams to be able to ship like AGI-pilled products.
[02:21] to build on top of. But I think that key bit about speed is like really intentional for us. And we really, really care about that internally. Externally, we actually have a lot more like complicated set of things. But one of the true norths that we have there is to be able to basically give any builder the tools to be able to work with Claude to build whatever they want to build. And so it's a bit of a broad statement, but as a result, that boils itself down into, you know, being wherever that business is.
Like we really care about like bringing our platform really, [02:51] a lot of time with the hyperscalers, integrating really closely directly with them, like AWS, Google, so on and so forth. And it is a lot of primitives that we end up creating. We want people to be able to express what they think their product should be. We want them to be able to almost do custom software in their own way. In this new world with AI, what used to be probably economically impossible was that last mile of custom software now, in theory, should be
[03:21] to go and do that. And so sometimes that comes in a form of primitives and APIs and higher order abstractions. And sometimes that comes in the form of just like [03:28] standards. So for example, like skills and MCP, those are things just like Cloud needs them to be useful. And we can just give them out to the rest of the ecosystem, work with everyone to help you create those things and get the best out of Cloud. So I would say externally, you know, we really are oriented around just helping you just be able to build.
But internally, that orientation, while still existing, is probably more, you know, specified towards speed and being able to move really quickly. How do you decide what goes into the platform, [03:58] to decide what products should be available? Yeah, I mean, we generally try to have a philosophy that we try to be consistent across the board. It's actually one of the reasons why we do internal and external. There's plenty of other platform businesses and constructs where you actually bifurcate these two things. For us, we kind of try to intentionally keep it equal.
And then as a result, we try to hold this philosophy as much as we can around for any builder, internal or external, even though for internal builders might have some slightly different requirements in the same way any user would have slightly different requirements. [04:28] We want to have the same primitives that are available to everyone. And one of the maybe the overarching thesis for that is that we've just seen like the capabilities of these models just grow and such as exponential. And it's really hard to figure out like a long lasting form factor.
I think two years ago, we were all like, everything's chat. And now everyone's like, forget chat. And it's just like agents. And like, there's gonna be another form factor, another form factor. And we kind of imagine that like constantly evolving. And so the best way for us to kind of enable that for everyone, and also ourselves is to actually build a really robust platform that gives us a lot of people. [04:58] gives people those kinds of like tools to figure out what those form factors are. And I don't think we by any means feel like we're the only ones capable of figuring out that form factor, like not at all.
In fact, the more democratization we can do on that and help people and allow people to experiment, I think the more of those form factors will actually kind of naturally come out of the market. Yeah. And I think within our team, we've had moments where we're experimenting, even with just like a packaging up of our primitives in a different sort of higher order way. [05:28] products that we've built into the world. And so we can go and dog food it for ourselves, but we'd never want to fall into this trap of like, we're over-indexed on the problem as it needs to be solved for an internal user.
Because exactly what Angela said, internal users have very specific requirements. External users have very specific requirements. And so if you over-index on one or the other, you fall into a trap. So a lot of the time what we'll do is dog food something internally at the same time that we open up early access of some sort with external customers so [05:58] bring those things back into the platform. - I'd love to talk about the higher levels of abstraction that you discussed. So I guess at the base level, this is just raw access to [06:07] Clive, Opus or whatever tokens.
How do you think about the, I guess, the layer cake of abstractions above that? [06:14] Yeah, if you look back, so when I joined Anthropic around a year ago, the platform was basically just the messages API. It was a messages API. You know, we had come out with standards like MCP. We obviously have developer tooling around our SDKs and our docs and our console and things like this. But for the most part, it was a stateless API. [06:44] and over again around as the models got better at running for longer and working with more contacts at a given time, you want to build agents that can succeed in a kind of long running context and even a remote context that doesn't necessarily have a human in the loop.
And so we found that we could piece together our primitives and stand up all the same infrastructure that we're finding ourselves standing up internally to power our own products and arrive at some higher order abstractions that let you do more agentic work out of the box. [07:14] that we're solving for you are, you know, infrastructure being kind of a hard thing to deal with. Like, how do you figure out spawning sandboxes that are going to have the right governance and security and like, you know, spin them up and spin them down when you need to, or the storage around transcript sessions so that you can resume a session if you stop it and pick it back up later.
So that infrastructure is a big thing that we wanted to be able to provide more of out of the box. And we do more of that today. And then the second thing just being harnesses and [07:44] of thought and energy going into how do I do my prompt caching and how do I manage my context window as well as how do I actually just get more intelligence out of the model and how do I manage my costs and things like that. So... [07:56] We've kind of packaged up our primitives a bit more in tune with the problems that we found ourselves solving to provide more of these things out of the box for people so that they can, if they're building systems for themselves internally, if they're building products, they can just be more focused on the problems that they want to be solving.
And if they want to offload some aspects of those problems to us, they can. And that's kind of the ethos. [08:26] Managed agents offering, I guess, just take care of it all for me. [08:29] It varies by the user group. So for, I would say, really AI-native startups, like the ones who are tinkering and experimenting at a really low layer, they're just going to go for the primitives. And then for everyone else, these kind of classic, more like enterprises or areas where it's like the purpose of the startup or the philosophy behind the startup isn't necessarily to optimize on some kind of hill climbing pieces, more like stringing together a bunch of workflows and providing unique user value to that user.
[08:59] You know, it's just kind of not their core competency. It's not where they want to focus their time and resources. And they reach much more for these kind of like higher order, like package offerings. What are some examples of the primitives you've released at different layers in the last few months? We've seen a few of them. We'd love to hear. [09:13] Yeah, I think maybe one framing I would give for some of the constructs that Caitlin was talking about is like, and this is a bit of an oversimplification, but effectively, there's approximately like three layers of this cake.
At the very bottom is just kind of like, like knowledge. And so this layer, like in many ways, it's knowledge about the model, it's knowledge about the things that the model needs. And it's just like the ability to know how to actually do something with Claude is maybe the way I'd phrase that. [09:43] we still evolve them, but they tend to be a little bit more baked. For example, there's very specific shapes and parameters we put on the Messages API, and it's more like [09:51] trying to expressly, like, showcase Claude's, like, design.
[09:56] like cloud the model's actual design, the way it thinks, the way it respects certain parameters, the way it kind of like will do tool calls, like all of those different pieces. And then we started standardizing like tools and then we started standardizing bits and pieces of like context that you could put in at different moments in time, which is concretely like skills and like memory. And so those are like the kind of like knowledge layer type of abstractions that we've put out over the past, I guess, like year plus a bit.
The next layer of abstraction that we've [10:26] is like once you kind of know stuff, then you do like [10:29] execute. And so at the execution layer, that level of abstraction is the part that Caitlin was talking about around like, we're doing these like higher order pieces, but like, what are we putting higher order there? It really is because you're now getting Claude to execute work. It's not just to know something, right? I can give it a question and give me an answer. You can put string a lot of that stuff together.
But now if you need to execute, like do work, give me the output, edit files and a bunch of different systems, that becomes a lot more complicated and requires infrastructure to handle. And so that layer is basically, I would say a low level harness plus [10:59] infrastructure as like the set of abstractions. Today, we just like our high level product for that is called Cloud Managed Agents. And so that's like a piece, but we started to wrap more and more pieces in that. I think there's going to be a layer like on top of that, we have like some inklings of it, we started to build towards but the last layer of abstraction on top of this is probably the coordination layer.
So you have knowledge and execution and coordination. And at the coordination layer, we've started to expose some of these in ways that like, aren't very obvious. But we're beginning to think of these things called like strategies, where basically, [11:29] almost like a meta harness, right? The harness, the true low level harness is designed for execution. [11:34] But the next one is about, okay, if tokens aren't really fungible and you need to give them different jobs, like maybe this token is advising versus this token is executing, this token is dreaming versus this token is executing, so on and so forth, you want to start composing these kind of orchestrated strategies that go together, and they should sit on top of all these things because at the end of the day, you still need to execute and the execution still needs to know what to do.
So everything in theory should kind of ladder together. And so I think if you were to look at our roadmap and maybe kind of project forward a little bit where you kind of expect us to go... [12:03] We'll move more and more from the knowledge layer to the execution layer and from the execution layer to the kind of coordination layer in terms of the abstractions that you can see us put out. [12:11] That's a really cool thing. How do you think this all comes together into a broader ecosystem beyond just the things that you guys are building?
How do you help support people building products on top of it? And how do you help them get the most out of all these pieces? Yeah, I think this is like super top of mind for us. Like we really want to find a way to be. [12:30] to support as many people in doing this as we can. I think we're still, like, learning. Like, a lot of the industry, like, has evolved. We've seen, you know, a lot of different pieces get spun up and spun down. And I think the operative part for Caitlin and I has been in the category of, like, making sure, at least at the base layer, that we provide as many primitives across the board as possible.
So, you know, this kind of, like, yeah, like, knowledge, execution, coordination, [13:00] on top of that. And that's just from a, I think, pure builder kind of point of view. Then there's a point of view around like, how do you kind of like plug in with us, right? Like, we're also building first party products of our own. We've also created some ways to embed natively with us, like, for example, connectors, which are built on top of the MCP spec. And we try to be more open about those types of things.
And we're starting to figure out like, what are the right bits and pieces? But what we're really trying to do is get to a place where, [13:30] products that they want. They can build agents if they need to. And then those agents and those products could be things that could plug into other agents. Some of those agents could be cloud agents. Some of those agents could be other people's agents. But we want to be able to enable that kind of like transactability across the board. And then I think in order for all of that to kind of ultimately be true, there is a bit around like standard setting.
And I think there's the traditional standard setting, which is around, you know, how do systems interoperate? And that's, you know, things that you've kind of seen us do with like skills and MCP, [13:58] But they're at, again, like the builder layer. I think at a higher order layer, there's also a bit around interoperability and standard setting around how do we all kind of like treat safety together? And, you know, we've talked to a lot of these companies and this is less from, you know, philosophies aside, just more like no one really wants to have technology that's like, for example, like doing negative things on their service.
Right. [14:28] These kinds of standard settings of how can we find ways to partner with more and more people to be like, yeah, we all kind of want to make sure our critical infrastructure is good. We all want to prevent fraud or any of those things from happening. And how can we work better with each of these members? I think on the last layer, we're still kind of like... [14:45] We're still evolving. And I think we're still very much like trying to find ways that we can be better and work with the rest of the industry to bring people along and work with them.
But those are kind of like, you know, the higher order primitives or pieces that we wish to kind of like be in place so we can work with folks to ultimately solve this. I think if I were to take a step back at the end of the day on all of these things, you know, like this technology is so transformative. [15:15] and was like, you can only do so many things. But with electricity, the reason why it's such a transforming technology for all of us, and so greatly of a utility, is because you can actually like wired it into everything.
Everyone is able to actually access it. We also have like standards and ways to plug in and do all the pieces that we need. And that's not something that anybody can do by themselves. They always have to work with the ecosystem and work with partners to figure out a path forward. How do you think about the philosophy of building an open ecosystem? [15:43] versus the walled garden. And how do you think about what products are really important for you to own first party versus... [15:50] where you're perfectly happy to plug into other components of the ecosystem.
[15:54] Yeah, so maybe in using Angela's kind of layered cake that we talked about a little bit earlier, you'll see that on some pieces of this, like execution, for example, what we've done within something like Cloud Managed Agents, and I think over time you'll see us try to make this a little bit more modular. [16:11] We actually aren't precious about you should run these things on our infrastructure. Like it should be sandboxes that we control or it should be a storage layer that we control. Well, we actually like, for example, we launched self-hosted sandboxes and we partnered with Modal and Vercel and Cloudflare and a bunch of other folks, even like Amazon's new micro VMs to have a first class offering where you can go plug any of those things in.
[16:36] We launched MCP tunnels so that you can call out to your MCP servers that are behind your firewall, right, and be able to punch through there. And so for some of these things, we, you know, the weather, whether it runs on our infrastructure versus somebody else's infrastructure is actually not important to us. [17:06] and plug those things in. And we think that that generally is a thing that works really well. Yeah, I think on the kind of like verticals where we might build products, [17:17] you know, I think we kind of have like two frames here.
The first one is we are always trying to figure out a form factor, like an evolving form factor. We, by the way, don't think form factors are like static. It's like a dynamic thing. So what might be awesome for one year's worth of AI development will probably not be awesome for the next year's worth. And we just kind of try to have that mentality. [17:36] just overall around Anthropic, everyone's always trying to be like, is this AGI pilled enough? And then we always have this mentality of like, you know, we build something, it works,
[17:45] "It was cool for a year and maybe it's not the right next thing." And so throw it away, try again. And we tell platform users the same thing. [17:53] I just think that's probably just like, you know, attached to the technology. But so, yeah, one one principle is like trying to always constantly find this new form factor. So sometimes we'll like launch products in certain areas to try to showcase a new type of form factor. It's not necessarily because we think it's like the biggest ham or the most important thing to go after.
But sometimes like, OK, this is like always been a really difficult thing and people have always communicated this way or tried some things this way. And can we show that maybe there's a slightly different way? And because the model capabilities are so advanced now, can we try to express it a bit differently? [18:23] What's an example of that? Yeah, you know, like... [18:27] Claude design is a little bit of that way. I think depending on how you squint, you might see it as like a way that we kind of going into design as as as like, you know, one of the verticals.
But more often than not, it's like if you take a look at what we're trying to do with that product, there's a couple of like decisions that were made in there. The first one is that like you can actually try to offload more and more and more to Claude. [18:57] Just talk to Claude to go figure it out. [19:00] The second thing was it was really trying to express that actually like [19:04] So code is a way to solve for things that you wouldn't normally think would be the way. So a lot of people who have built kind of generative, you know, like slide decks or designs or whatever, will pick the way of like they have like some kind of design system, you integrate against design system, it's almost in the traditional, like classic Wissy Wig style of designing something.
[19:34] through some experiments early on it's like actually it looks like it can kind of do that and how can we kind of showcase that uh to the world so that's like an example we have a lot of other internal projects and this kind of falls in the category of like expressing form factor we'll all try it out internally it'll be super cool for like two weeks and then we move on to the next thing we never even ship the thing frankly but yeah we actually do a lot of product experimentation in that area that's like our labs team and then there's like the second category which is that we actually do look at tam like we're a business we do look at tam we do look [20:04] agentic operations that would happen.
[20:07] In those areas, we do tend to have an orientation towards things that are more token heavy. And by token heavy or token hungry, maybe is the way I would say that is like what we mean is like, you know, you for spending once you spend a call it like one turn. You look at the end of that turn and you say, like, am I done or am I actually so glad that I did that thing? I want to do more of that thing. We like industries where it's like the answer to that question.
You say, I want to do more of that thing. So coding is obviously the one that we all know. [20:37] once you've finished a turn, you look at that and you're like, that was incredible. I'm like unlocked. I'm going to do like more. I'm going to build more. I can do more. And there's other services where it's like, actually, when you finish that turn, you completed the job and you just move on. You know what I mean? And so we tend to like go into the ones that are a bit more like there's this kind of like iterative flow.
You're going to build more, generate more together. And then the last angle that we kind of take a look at is just sort of like, you know, there's going to be certain business functions that we're like, they are the buyer that we like to go to. [21:07] them create better products there. And I think we've been pretty transparent with some of the verticalization, like we've done like finance, we've done like legal. And we've tried to kind of like narrow on into specific areas where we feel like by having the right context and the right tools and putting it together in a good form factor is probably useful for us to be able to do.
And in each of those areas, we do, we're trying to do a bit of like showing the art of the possible across all the different ways that you would accomplish those outcomes. And so for, you know, [21:37] is a good one. You know, we, you could be a company that solves problems in finance, and you could build directly on the messages API, and you can just get some tokens, and you can build everything else on top. Or you could be someone who builds on cloud managed agents, you can get a lot more out of the box.
Or you could say, I'm going to build a plugin that are like a connector, right, that's going to sit within one of our products and within those form factors. When we did recently, we launched like cloud for financial services is like, okay, cool, we've got packages [22:07] of skills and things like this that you could choose to use within our product, within other people's products. We even launched like cookbooks on here's how you would use cloud managed agents to go and do these things. And so, [22:17] I think for us, it's all kind of an experimentation around, like, you know, we provide people all these different pieces and see kind of where they run with it.
And then sometimes we put together products that are just packaging of all of these things. Like Claude Tag, I think, is a really good example. Like, we had been seeing people in the industry go and say, like, Shopify did this with River. Square Block recently did this with BuilderBot. There's, like, a few of these examples where people said, I'm going to build, like, [22:45] an agentic platform internal to my company, and I'm going to try to give it all the right context, and I'm going to make it accessible from Slack or from various other, you know, platforms that you'd want it to be accessible at.
And I think Claude Tag was very much a... [22:59] packaging of all those same things that anybody could choose to build something similar. But this is how we're kind of like, well, this is how we're doing it internally. And if you would like to just kind of plug in and go, here's what that looks like. What do you think people misunderstood about CloudTag? Because there was all this like ruckus about, oh my gosh, it's just a Slack bot. Like, tell us what the magic of Tag is. Yeah, I think it's a great question.
And I do think it actually showcases a little bit of where maybe the future could be going. [23:25] Yeah, I think like the [23:27] I think if you look at products in the past, people are like, oh, you really attach to like the form or the UI almost. Right. Like it looks like this. So it's like super cool. And I think when you look at like tag it like, yeah, like the way you interact with it is that you like literally tag it in Slack. And so, yeah, that is like the interface.
But that's not really the important part. The important part is all the kind of like. [23:50] context engineering and like architecture that we put underneath the hood so that tag just works it really should just like just feel like a co-worker like a co you know if you go to a company and you onboard the co-worker comes into your channel and then you can chat with it it's proactive it figured out like what's like useful you could and um it just gets stuff like done for you and so if you think about you know especially like non-technical audiences this is like it's a huge unlock you just you literally create a channel and then you add claude or sometimes you don't even [24:20] you're like, hey, I want to be able to do this and do that, and I can't figure out this, and how do I actually submit an expense report again?
[24:26] And traditionally, you think about how to solve that workflow. You are going all over the place and you're talking to your manager and you're talking to your spin-up buddy. And it's really, really complicated. [24:33] And today, now you just like go talk to CloudTag. And we do a lot of the hard work on doing the context engineering, the proactivity, a lot of the harness pieces. I think Andre Kaparthi said it really well. It's like it's like an org level harness. There's a lot of like complexity baked into that. Like Kayla mentioned, like you can use our APIs to go and construct that.
You can do a lot of experimentation yourself, obviously. But this is like an opinionated take from Anthropic on like how you can have this really awesome, always on kind of agent for your entire company. [25:03] futuristic, I guess, is like [25:06] A lot of that complexity is actually like, it's like an iceberg. It's like all the stuff underneath it. That's actually becoming the harder and harder and like useful part that we're trying to like push through. And I think we'll see more and more like that kind of like tidbit that's like outside in the water.
It's just like. [25:23] the interface can actually constantly swap. Like today, right, like Slack is a place where a lot of people collaborate, a lot of business collaborate, but also a lot of people collaborate in teams. And some people collaborate by a WhatsApp group, or they text each other, or some people still email each other. And like those could be the form factors that actually completely you can imagine agents just going there and being, and they're almost taking up the same form factors as humans have taken up. It was almost like a very broad, [25:47] almost like boring take, but it's actually like, I feel like the most like forward one, because you want the agent and you want AI to basically be like another person.
And it's helping you, but it's like, you know, very intelligent, can figure out all the context, and you can always have it to be a really helpful assistant. [26:03] Totally. You talked about context and then harnesses quite a bit. And so your team is just, you know, has such an opinion point of view on like what it takes to build an exceptional agent. I imagine a lot of that comes down to the context engineering and the harnesses. Totally. Maybe like what best practices or advice would you would you share with people about what you need to get right on the harness and what you need to get right on the context?
[26:24] Yeah, I think so. It's interesting because we've kind of talked about, you know, we launched Cloud Managed Agents as this like, [26:31] Very generic but high-performing harness because we've done all the nitty-gritty work that's actually really boring and not super interesting around how do you deal with prompt caching? How do you deal with context management? You clear old stuff out of the window. Sometimes you call tools programmatically so you don't pull everything into the context window and you can keep it clean. There's a lot of those sort of details on the lower-level harness layer.
[26:56] And I think, honestly, like, best practices are just stuff like – [27:00] prom caching, do it, save a lot of money and token costs. Obviously, like try to keep your context window clear. And then putting those things together in a harness that will be performant is, you know, sometimes specific to the tasks that you're trying to accomplish, right? And then, of course, evals. I'm surprised we got this far into this thing before one of us said the word evals, but like you need evals to make sure that what you're trying to accomplish is performance.
[27:30] to go, and Angela mentioned this a little bit earlier, is more of a concept of strategies or meta-harnesses, because I do think that [27:38] Yes, you can, again, make this lower level harness is going to be performant. And maybe that's interesting for you to do yourself or maybe not. And you offload it to us. But this concept that. [27:48] You can take any given token and spend that token on just executing. Or you could take that same token and choose to actually reflect on your past agentic sessions and write learnings to memory so that the next agent does a good job.
Or you could take that token and advise with a bigger model so that a smaller model can execute and do a better job. Or you can say execute, execute. And then like a grader comes in is like, did you do a good job? No, you didn't try again. Right. [28:18] is gonna come more at that higher level on like the meta level, right? And I think optimizing within those strategies is something that our team is really excited about and we're starting to do a lot of work there. And I think a lot of other people are starting to feel really excited about this concept of strategies and like the jobs you give to tokens.
[28:36] Because again, like, yes, there's best practices on stuff like your prom caching and exactly how you clear stuff out of your context window and how you write your evals and like a lot of things like this. But I don't know that there's necessarily so much juice to squeeze in a lot of cases out of that layer as compared to a layer higher than that. Yeah. And one of the reasons for that, I think, is it has to do with the generations of the models. If you look like two years ago, a lot of the harness was like a scaffold to kind of like...
[29:01] tell the model to go from point A to point B. And you had to like, you really had to like build in a lot, you'd practically build one wall here and one wall here, so to like the thing would go in a straight line. And now the models are actually very, very steerable. And so a lot of that steering, you could just put in the prompt, right? Like go do, go from point A to point B. And the model like, [29:21] will go from point A to point B. So a lot of, if you have harnesses that are, like, designed to kind of do that kind of, like, steering, you can delete that part.
Like, that part we actually frequently encourage, we can delete part of those harnesses. I think various people have said things along those lines, and that's, I think, what people oftentimes mean when they're, like, either the model will kind of consume some of the scaffolding. And, like, in that sense, like, for sure, if your scaffolding is telling it to go in a direction that it can just intelligently figure out, like, that I think will [29:51] what the harness needs to start doing is more allow it to run longer. And so that's where like that execution bit tends to be.
I think like it sounds like a maybe somewhat silly point, but I do think it results in a lot of differences because because you can go in the direction you tell it to go, you obviously don't want it to stop at B and be like, OK, now go from B to C and then go to F and then go to Z and then come back to me on A, you know, something funky like that. In order to be able to do a lot of those things, the kinds of harnesses that you do are less the [30:21] which allows you to operate at a slightly higher level of thinking, which matches, I think, a lot of the intelligence gains in which I see with the model.
Do you think task-specific harnesses make sense? Or vertical-specific or task-specific harnesses? [30:33] I think people have different opinions on this. Like, our opinion is yes. I don't think there's, like, a general harness. I think there are some capabilities that are obviously very general, and they tend to, like... [30:42] be very useful. Like coding is a capability that like is very useful because you can use it across so many things and software has, you know, just like eaten so much of what is capable. So our ability to like write software is therefore useful.
I think when you think about like very, very specific types of domains, they were going to require like a couple of pieces of the harness to be sort of like customized. One of that I do think is how you choose to kind of like handle [31:12] model um so in like domains where you require like an extreme level of verification that logic of how you handle that very like again i think it sounds small but like i totally understand why some people feel like they really want to own the harness because tweaking that last bit will give you a ton of juice and especially domains like legal and finance where there's a lot of consequences um to you're not getting it perfectly correct like it's really going to matter and that's going to be the difference between your product and someone else's product being the thing that the user [31:42] which I would say it's not going to matter as much because you're able to compress it into a general model capability.
So the tweaks that I guess where we feel like the domain specificity is really going to matter is the specific verification logic between the model. [31:58] and your execution. And then I think it's going to be about like some of these kind of like higher order strategies on how well you're able to actually like allocate your token budget. I think the context bit is actually a little like overdone. Like, yes, you're going to like throw in context and like that's but any harness can actually handle a lot of context. And so that's just more like you have the data.
And if you have the data, then obviously you're uniquely qualified to do something useful. [32:24] Yeah. And I think when people say harnesses, they often mean a lot of different things. And I think this is why, in part, there's so many different opinions on this. Like you can think of a harness as literally just like a loop that's like, OK, cool, like user model, user model tool, you know, like that sort of thing. Then you could think of the harness as also all of the tools that are packaged up with the harness.
Right. And and there's just like a lot of different definitions of these things. [32:54] are saying to own and deal with. What I was kind of saying earlier is like, [32:57] getting your prompt caching right, right? Like maybe that is not the world's most interesting thing. Choosing to clear out old tool calls from the context window and things like that, right? Or like maybe a little bit less interesting. [33:10] And you go a layer higher into some of the stuff Angela's talking about, and then you get into, okay, yeah, these are things that I might want to own and control.
And so it's interesting with cloud managed agents. [33:20] The thing that we built today, [33:22] We call it higher order, but it's not really like that high order in the sense that you can choose to define all of the tools that you want to bring in as custom tools with the harness. Right. And like we give you a lot of knobs to control. You can define skills. You can do your system prompts. You can do a whole bunch of different things, MCP servers and things like this. And I think where we want to get to is a point where you can literally just tell an agent, here's the outcome I want and here's the budget that I want to spend.
[33:52] And so I think there's just a few different layers of this, right, that for certain things, like you might want to sit at a different layer of what you actually go and control. And you can probably get better outcomes within some of those layers by doing a little bit more optimization work. [34:07] Very cool. One of the things I'm curious about, and one that I love about infrastructure and platform teams is that you get to see what the most advanced users in the world are using and learn from them.
I'm curious, what are some things that you're seeing and learning from the people building on your platform? [34:21] There's some people that have been doing some really... [34:23] funky ways of like handling context. We ourselves explore this a lot. That's actually like one of the reasons why Tag is like such a great product is like there's a lot of really awesome like context kind of engineering that's happening. We've seen some teams be really clever about like how they do that. And they are able to kind of think through like, okay, if I have all these contexts in a bunch of different places, how can I proactively go reach out to them?
How can I try to generate enough like, [34:49] permissions across each of them and then feed that all into like an agent. And it's interesting that like, I guess like this is kind of the level of innovation that like we're actually like very excited by. It doesn't express itself as like a completely different product form factor, but what it actually does express itself as is like maximally useful to users. And we've been seeing this more and more with like actually like internal use cases instead of like external ones. So like companies who are becoming more AI native, basically, they're the ones we're seeing [35:19] had like customers try to do this for their like they've built their own like custom sdlc kind of setup in very very innovative ways we've had uh ones who do that for like their entire back office and just like the kind of nuances of how they like stream in context i think has been like actually really interesting in terms of like how they've been putting together the pieces so that's been like one category that's been like really really like fascinating uh another category that's been like really interesting has actually been with companies that are dealing with like really [35:49] that we kind of engage with.
And, you know, like they're like the systems I'm working with, they don't even have APIs. Like that's that's a dream. And so, you know, how can they use computer use and things like this to be able to start to kind of automate and create more connectivity with our systems? And that area of innovation, I think, has been really exciting. It's been really interesting to see people try all sorts of crazy stuff from like taking a laptop and trying to like run a bunch of things on it to auto generate a bunch of things that then their agents can go and use.
[36:19] area of, I think, a lot of innovation coming from a lot of our customers that we want to find ways to, like, support better and see, like, okay, maybe there are, like, how can we make this easier for you? How can we help you with some standardization? How can we get it so that, you know, like, you can just have a spec and then Claude can then respect it. And so it's much easier for you to organically connect a lot of these things. But yeah, maybe the general theme I would just give you is, like, interestingly, a lot of the innovation that's most exciting out there right now has been [36:49] Yeah, like a good one in that we were working with a customer who they built some agents on cloud managed agents.
They also have some agents they built on other models and other platforms. And they've kind of optimized each of these agents to be good at the things that they want. They want these agents to all be able to work well together. And they kind of were like, wow, galaxy brain, like, what if I expose an MCP server on top of this agent so that I can then go and like have this other agent call a tool on that agent, right? [37:19] and be able to work together. And we were like, yeah, totally.
And we sat down with them and worked through it and it worked perfectly and it was pretty cool. And so we're seeing a lot of, again, that connectivity layer that I think is one of the cooler areas where, [37:31] people are innovating, but [37:32] Outside of that, one thing that has been cool is just seeing the shift in, I guess, like industry trends of where we're seeing a lot of our usage come from. Like we talked a lot about coding, like coding as a category, like, of course, absolutely exploding. There's so much going on there.
And we're starting to see some of these emerging trends. Like more recently, we're starting to see manufacturing really pick up as just a category where people are building with AI and like one of our PMs, like getting on a flight to Detroit to go like figure out what these customers, like what they need. [38:02] and what's going on. And so I think we're going to start to see a lot more just kind of like outside of the box of what people think about today sort of use cases, which we're really excited about.
It seems like there's no there's a we went through a token maxing moments of history and now there's like the token rationalization moments of history. What are your thoughts on that? And like what should companies be doing? And then how does the platform team think about enabling that? [38:28] Yeah, I mean, it makes sense. It makes sense from the high you start to rationalize. I really like that framing. And... [38:35] I think there's like a couple of things that are like top of mind for us on this front. I think like, again, it makes sense.
And as these models get more and more capable, you're going to hit like levels of intelligence maxing that are like there that then you want to do the next kind of dimension. And the next dimension after intelligence will either be cost or it will be speed. And you just kind of go through that across all possible task complexities in the distribution. [39:05] AI usage, right? Like that's kind of the wrong move. And we do actually see some of our customers do that. So oftentimes the way that AI spend has erupted inside their company has been through some kind of like shadow IT, you know, like their employees just like want to use it.
They find a way, they end up procuring it themselves. And before you know it, like half your org has like found some way to install cloud code. And in that world, it is kind of hard to manage because these things are, again, like they're very token hungry ultimately. And so what we try to kind of encourage our customers is like, you don't want to like stop the innovation. Like if you are getting returns on [39:35] shipping faster than ever before, you can run more operationally efficient, then those are gains. And so the area that we actually try to encourage people is if there is a way for you to construct, again, a strategy that allows you to design an architecture that says, given a task, assesses level of complexity.
I mean, I'm effectively describing a router, but there are ways to do this that are, I think, a bit better now. And so this task comes in, has a certain level of complexity. For that level of complexity, you can run more efficiently. [40:02] You can define some rules, but for the most part, right, if it's, like, a hard task, you should probably route that to, like, a big, super smart model. And if it's not a hard task, you can route that to, like, cheaper models. Designing that, I think, has a little bit of, like, there's a lot of technical complexity in that, but it's, like, very, very doable.
And we actually, like, encourage people to try those kinds of things. I think ultimately... [40:21] I think within the quad space, it will like make sense. It's actually one of the strategies we imagine like designing because the way that we're kind of thinking a lot of these things is like it almost feels like every month there is a new era of something. And if we take a step back, like, OK, and it seems to be like really fast. And so what are the different ways that are recomposable so we can redesign very quickly for any new whatever the cool thing is that month kind of like bit.
[40:51] I think the bit that we do feel really strongly about on the model routing front is like we are designing our platform for Claude and we want to make sure that Claude is great at like solving all these things. So we'll like restrict to that space rather than, you know, I don't think we're that interested in saying like, okay, and then, you know, you should route to a different model or whatever. [41:10] Makes sense. Yeah. And well, some of that, too, is just like I think we have a strong belief that harnesses and just like the agentic layer should be tuned to the model family that you use it with.
And so I think there was a period where people were kind of like, yeah, cool, I can like build a harness and build an agent and then just like plug in a different model underneath. And they were excited about routers from that perspective. And I think we started to see like Vercell just did this with Harness Agent, for example. [41:40] They actually plug in the whole harness and the whole agent that's tied to a model family, which makes a lot of sense. And so what we could provide is a little bit better, smarter, like how do you mix and match the right models within the model family underneath that thing, if that makes sense.
But, yeah, on the general question of token maxing costs and these sorts of things... [41:58] I think we're just kind of going through what feels like a normal, natural cycle for companies and figuring out how to make the best use of this technology and run their businesses really well and really effectively. And it's interesting, like before working at Anthropic, I was at Stripe, and we were kind of in the very reasonable era of like, we paid a lot of attention to our AWS bill. [42:28] or whatever it is, right? Like at any given moment and causing a big increase in spend, that's not actually worth it, right?
Like we have put in place the guardrails to find that and then go ask that engineer very nicely to please turn off their background job. It's not like within the bounds of what they should be spending for the thing they're trying to accomplish. I think those are the things with AI that people are going to start to go and figure out. And I think to Angela's point, [42:54] The thing that gets dangerous is when you're kind of just like, here's a cap. [42:57] and you're stuck within your cap, like ready, set, go.
But I do think that encouraging innovation, encouraging people to create really excellent outcomes with this stuff, and then coming in from the side and looking and saying like, okay, well, there are a few different ways that we probably could have accomplished that outcome, right? And one is like you take Opus and you run it all night and you do something crazy. And another is maybe to get a little bit smarter with the strategies that you put together [43:27] And I think that's the like next layer of thinking that everyone's going to start to do.
Very cool. Is there anything that you guys are excited about building over the next two months that you can share a hint at what might come next? Yeah, I mean, I know we said this word like 20 million times. I apologize. But like we really are trying to build ways for you to compose strategies. And so that is an area that that we're like trying to move into that kind of like, yeah, coordination layer of the abstraction. [43:57] add a layer where it's like, in order to get the most return on this, you have to be a little clever about like, what is the nature of the problem that you're solving.
So to give you something like concrete, like when you try to solve for like, let's say you want to build an agent that's like trying to do bug hunting. And you could just send one off to go and do that. And it's going to give you a certain level of return, a level of return possibility. And then people kind of get stuck at that. And they're like, okay, my next options are I can like, [44:24] make a bigger I can just like swap the model for a different I'd probably bigger model um or I could like let it run like longer and that's pretty much like the only two like levers that you have to like try to make this like bug hunting agent for a lot of experimentation when we do these kinds of things there's like actually the thing like those two those two things are still true but you actually have like a third lever and tends to actually do a lot more than you think it does which is that actually if you were to like best of end the thing [44:49] it would like give you a lot more returns.
But like just to be, just saying those words are fine. And there's plenty of papers and people have published it to actually build that thing and put it into production so you can actually test it on users and see the results for yourself. That's like really, really freaking hard. And you end up building all these like custom harnesses, so on and so forth, or like, you know, all that stuff. But we're seeing like, this is where the alpha is and it's hard. And so like in the same very simple philosophy that we talked about at the beginning, like if it's like gives you the return that you want and it's hard, we're going to try to just make it easy for you.
So then you can use it to then, [45:19] run the experiments you actually need to run. [45:20] It reminds me of when people were talking about agent swarms a year ago. It's some version of that. Yeah. Has it been a whole year? Yeah. I know. We're finally there. Yes. Yeah. No, I think that that's like, that's a type of strategy. Exactly. In the same way that you have like, you know, one big one that separates a bunch. That's another type of strategy. And I think people have thought about this maybe in the way of like human organization.
I guess it could be similar, but if you take it to kind of its end state, it's actually more just like the token has a job. And I think it's this job piece that we're really indexed on. [45:50] see a lot of returns to and that's the thing that we want to spend time with users and the rest of the ecosystem on on like how can we just make that easier for folks to then experiment like we can give you like five jobs off the top of our head and we'll probably like that's what we have internally um and if we give this out to the rest of the ecosystem there's probably going to be like 100,000 200,000 who knows what other combinations that people could put together [46:10] We want to be able to keep doing this hill climbing on how do you get the most value, the most intelligence per dollar and just put that power in people's hands.
But around the edges of that, we have these personas that have things they have to work through in order to be able to really deploy AI either within their companies or within their products. And that's the sort of enterprise-ready security and compliance controls and things like this. [46:40] more modular in the right ways, like being able to plug in different pieces of the solutions that we're building. Like I want to use memory for this thing over here, right? Or whatever else it is. And having a truly excellent developer experience around that, because we spend a lot of time with enterprises who are like, okay, I have this like walled garden.
I need to figure out exactly how I can plug these solutions in. And so we're, we've got a part of our team that's innovating on things like strategies and jobs and trying to help you maximize intelligence. And they're like, that's [47:10] reasons. So I think solving those problems is really, really important to us. But then the other persona is, you know, the like weekend developer who's like, [47:18] I want to go and build something useful for myself, right? And they're often doing that on top of our platform and on top of many other just pieces of developer platforms in the community.
And I think for some of those folks, there's more that we can do to provide solutions that are maybe more open or more hackable or whatever it might be for those folks to kind of just like go wild with what we can offer them and have this really excellent developer experience. [47:48] out because I think those are the things that then unlock getting people to say, okay, yes, this thing works for me. And now I can plug in on some of the stuff that you guys are doing that's really innovative and hill-climby to get more intelligence and save costs and things like that.
[48:01] Wonderful. Caitlin, Angela, I feel, I mean, you were building one of the most important developer platforms in the world. And talking to the two of you over time, I just feel really optimistic that that platform is in very thoughtful hands that care about the ecosystem. So thank you for taking the time today to share what you're up to. And we look forward to what's ahead. Thanks for having us. Thank you, guys. [48:23] Music
Want to learn more?
Ask about this episode