AI predictions: Job markets, Codex beats Claude, and the death of org charts | Dan Shipper
Dan Shipper is the co-founder and CEO of Every, a media and software company that’s become a living laboratory for the future of work. Everyone at his company of about 30 people is an AI early adopter; from editors to ops people, they use AI to do much of their work, giving Every a unique lens into where the world is heading. A year ago on this show, Dan predicted that people were sleeping on Claude Code for nontechnical work, which proved to be remarkably prescient. Today he’s back with another set of calls: the SaaS apocalypse is dumb, CLIs are over, the forward deployed engineer is the most valuable new hire, and the only thing you need to do to stay employed is ride the models.
Appears in
- Uploaded
- Uploaded Jun 14, 2026
- File type
- YouTube
- Queried
- 0
Full transcript
Showing the full transcript for this video.
[00:00] The last time you're on this podcast, you had this hot take that people were sleeping on cloud code. You are so unbelievably right. The premise of this episode is we're going to go through what else you predict will happen. The AI jobpocalypse is not really a thing. I am super, super bullish on PMs and full stack designers. You guys are hiring doubled in people in the past year, which is not what people would have expected from a company that is so AI forward. I'm simultaneously extremely AI pilled and very bullish on humans.
Automation is a lie. [00:30] Every agent needs a human. We have so much automation, so much AI, and I also work way more. Creativity. It just feels like it's going to be more and more valuable to stand out from all the slop that people are shipping and launching constantly. What models do in general is they make yesterday's human competence cheap, and so it becomes commoditized. It's not valuable anymore. What humans do is we go in there and we're like, yeah, we have all this frozen human competence from yesterday. How do I use this to make something new and interesting?
What are some predictions for how the way we work is going to change? [01:00] is everyone's going to have at least one agent that they talk to that they can offload work to. Second is that most of the work that you do is actually going to happen on your computer in an environment like Codex or Cloud Cowork. What you're predicting here is the SaaS tools will run within Codex or Cloud Code. I think the SaaS-pocalypse is dumb. I would buy SaaS stocks right now. What agents do is increase the number of users of SaaS, not get rid of it.
A lot of people are moving to CLI and trying to work from the terminal. We speed ran the CLI era. It was nice while it lasted, [01:30] I think CLIs are over. [01:33] Today, my guest is Dan Shipper, CEO and founder of Every. Dan and his team are building maybe the most AI forward startup out there. [01:41] And as a result, [01:42] are very much living in the future of how work is going to look as AI becomes a bigger and bigger part of our day to day. [01:48] Everybody at their company, including every non-technical person, uses Codex and Cowork and Cloud Code, [01:54] to get much of their work done.
And this is why way before anybody else, Dan saw the rise of Cloud Code and what is now Co-Work, which he predicted almost a year ago when he was on the podcast last time. So I asked Dan to come back on the podcast to share his current biggest predictions for how work is going to change over the coming year for most people. We chat about what work will look like at most companies at the end of this year, how the shape of the work we do [02:24] working on right now.
Hint, hint, product managers and designers are going to do very well. Dan makes a lot of bold predictions and many quite contrarian takes that I was not expecting him to say, and we are going to revisit this conversation exactly a year from today to see how much he got right. Before we get into it, do not forget to check out Lenny's Product com for a free year of the hottest and most well-crafted AI products in the world available exclusively to Lenny's [02:51] With that, [02:51] I bring you Dan Shipper. Dan, thank you so much for being here.
Welcome back to the podcast. [03:01] Thanks for having me. Always a pleasure to be with you. [03:03] The last time you were on this podcast, [03:05] It was almost like an offhand hot take. [03:07] that people were sleeping on Claude Code [03:11] and in particular, cloud code for non-engineering work for just like [03:15] fixing files, sorting your hard drive, just all these things that people hadn't thought about. [03:19] Nobody was talking about this. This was a year ago. [03:22] You are so unbelievably right. [03:24] about this. It's just like, unreal what has happened since then.
[03:27] They built Cowork, which was this whole, they built on this very specific idea using Cloud Code for Non- [03:33] Technical work. [03:34] A codex is getting into this now. I imagine you've been seeing this. They're like leaning into this. [03:38] non-technical use of basically coding agents. [03:41] I feel like this has also been a big part of Anthropic's success over the past year, just like how do non-technical people use this stuff. [03:48] So, you were just so ahead on this stuff. I even wrote a newsletter post building on this idea.
I'm like, "Hey, this is interesting. I should dig into this." I asked people how [03:57] Cloud Code for non-engineering work. And I just had so many examples. And it's my second most popular post name. [04:02] So clearly you [04:05] Yeah, but [04:05] unique glimpse into where things are heading so the premise of this episode is we're going to go through uh [04:11] what else you predict will happen. [04:13] in the future how things will change for people building products [04:17] And I think it would be helpful to start with giving people a brief glimpse into just how you operate and how your team operates.
[04:23] That gives you this unique lens into where things are going. [04:26] So just give us a sense of how you work. [04:28] Thank you. I really appreciate the introduction. [04:33] And yeah, I think one of the things about predicting the future or the way that we [04:38] think about predicting the future at Avery is that [04:41] What you don't want to do is prognosticate. [04:44] What do you want to do instead? [04:46] is... [04:47] is just living it together. So everybody at every is an AI early adopter. We're almost 30 people now.
I think when we did our interview, we were 15. So we've doubled in size in the last year. [04:59] we're all early adopters and we have engineers, we have designers, we have writers, we have editors, we have [05:06] sales people, we have customer service people [05:09] And, [05:10] Everybody has a little bit of that. [05:13] whatever that thing is where you're just like, I like to explore. I like to experiment. I'm very curious and I'm like super all in on AI. And what I what that does, I think, is it creates this like little pocket of the future where we're all living in it together.
And we get to be a little bit further ahead because any other company, there's like a mix of people. There's really adopters. There's like there's sort of like the middle of the pack people and there's people who are like very anti-racial. [05:37] And another thing that happens, which is really cool, is we get to, because of our role, [05:42] Um... [05:43] reviewing models and [05:45] being a little bit of a tastemaker in AI, we get access to stuff before it comes out. [05:50] So, [05:51] we get to beta test and alpha test and kind of help steer the direction of where things are going a little bit, which is very, very cool.
And so when I think about predicting the future, [06:00] It's actually when you create an environment like that, it's actually just about noticing what's going on. [06:07] And I think a core part of it too is writing about it. [06:11] articulating what you're noticing, articulating the future, kind of brings it about in this way that [06:18] makes it real for you and your team and then anybody else who's like on the internet who's reading it. [06:24] And so the Cloud Code thing, it's this very organic thing where [06:30] for us, [06:31] Um, [06:33] We tried Cloud Code when it came out.
That's sort of our job. We try all the new stuff from all the new model, we try all the new stuff from the model companies. [06:42] And. [06:43] At the time, it was like a little bit early. [06:45] But right around, I think, like Sonnet 3.5 or Sonnet 3.7, we were testing that to do our vibe check on it. [06:51] And we were like, holy shit, this is crazy. This is like really, you can, they got rid of the code editor. [06:58] And so from that point on, we just basically we run at this point now we run like six products, software products internally.
At that time, we ran like maybe two or three. [07:06] And from that point on, we just started shifting to a world where everybody was – [07:11] No one was looking at the code. [07:13] Everybody was talking to their computer in English using Cloud Code in the terminal. [07:18] And so I was able to see like, ooh, this is starting to happen. [07:23] Um... [07:24] And then because my job is a little bit to just like push and play with stuff, I was like, I wonder if I could use this for like my writing.
Like, how could I do that? And then it just like starts to unfold and you're like, OK. [07:35] This is... [07:36] Not ready yet. [07:38] But it's obviously useful for me. You know, like one of the things that we talk about internally is what I call the reach test, which is like, do you just like when you wake up in the morning, do you like reach for it organically? [07:48] I love this combination of [07:49] You are using the latest stuff and I think this is, as you said, maybe an underrated skills.
[07:55] You're good at [07:56] being self-aware of here's what's weird and new and different and interesting. So that's a really cool combination, partly because you have to write about it and you write about it. So I think that's like the perfect recipe for someone having a sense of where things are going. [08:07] This episode is brought to you by our season's presenting sponsor, WorkOS. What do OpenAI, Anthropic, Cursor, Vercel, Replit, Sierra, Clay, and hundreds of other winning companies all have in common? They are all powered by WorkOS. If you're building a product for the enterprise, you've felt the pain of integrating single sign-on, SCIM, RBAC, audit logs, and other features required by large companies.
WorkOS turns those deal blockers into drop-in APIs with a modern developer platform built [08:37] specifically for B2B SaaS. Literally every startup that I'm an investor in that starts to expand upmarket ends up working with WorkOS. And that's because they are the best. Whether you are a seed stage startup trying to land your first enterprise customer or a unicorn expanding globally, WorkOS is the fastest path to becoming enterprise ready and unblocking growth. It's essentially Stripe for enterprise features. Visit com to get started or just hit up their Slack where they have
[09:07] WorkOS allows you to build faster with delightful APIs, comprehensive docs, and a smooth developer experience. Go to com to make your app enterprise-ready today. [09:17] So the way that I'm going to structure this conversation, there's going to be basically three buckets of predictions. [09:22] One is how the way we work is going to change in the coming years. [09:26] Two is what the shape of the work we're going to be doing is going to [09:30] look like and change. [09:31] And then three is who is going to be most successful in this future?
[09:35] slash what should you be doing and working on now to be successful in this future? [09:41] is we come on a year from now and then you score it. I want to score. Okay, so this is a year from now. Okay, okay. So is this, let's actually, uh, [09:51] Is this like... [09:52] your predictions for in a year this is what it's going to look like or this is like the emerging future? I will probably say I don't have like an exact timeline. I think most of the stuff that I'm going to talk about will be pretty apparent within a year, but it probably it may it may take longer than that.
But I think it should within at least a year be like. [10:10] not obviously wrong. Like it should seem like it's moving in that direction to count. [10:15] Okay. May of 2027, we will review your predictions. [10:21] Meck. [10:22] I love this. So let's dive in. [10:25] What are some predictions for how the way we work is going to change in the coming year? [10:29] One of my favorite questions, because I think if you look at the benchmarks, you're just looking at, OK, like, yeah, AI is going to just take all of our jobs, basically, you know, meter has this really cool benchmark where it's like it measures how long it can like the newest models can do tasks autonomously.
And it's like, oh, it's like. [10:47] It can. What's it called? Oh, like mythos preview, the like big anthropic model that everyone's like so worried about. It can do tasks of 17 hours at 50 percent accuracy. It's like, holy shit, that's crazy. And I think it is real. It's true. And and the progress like model progress is going up exponentially. [11:08] Thank you. [11:09] And my experience and my feeling is that we will look back in a year and say, [11:15] We actually have a lot more work to do. Humans have a lot more work to do.
Um, [11:19] even as models get better at doing work. And there's like a really interesting paradox there. And my prediction for the, like how work will, my big prediction of how work will change or how you will be doing work in a year [11:34] is it's going to bifurcate in this in two main ways how you how you use agents [11:40] One is you're going to be doing, I think like what we, [11:45] figured you would be doing like five years ago when we thought about how [11:49] work with AI works, which is [11:52] Everyone's going to have at least in their company, at least one agent.
[11:56] that they talk to that can do work, that they can offload work to. And we'll talk about what that looks like, but it's essentially like OpenClaw. [12:05] Second, [12:06] is that most of the work that you do [12:10] is actually going to happen on your computer [12:13] in an environment like Codex, [12:15] or a Cloud Cowork. That becomes the sort of operating system for [12:21] It becomes a sort of operating system for how you do all of your work whether that's your email the documents you create like all that kind of stuff it's going to be on [12:30] that kind of a surface.
That's becoming the clear competitive landscape. So I want to go in order of those two. [12:39] So the first one is you're going to have agents you delegate to probably in Slack, but [12:43] you know anywhere. [12:44] First thing that's interesting about that one [12:47] is it's not clear what the architecture is going to be like for that. Is everyone going to have an agent? [12:53] Is every team going to have an agent? Is it going to be like just one agent? Is it like do agents specialize? Is there this like parallel shadow org chart?
[13:04] And, you know, [13:06] When OpenClaw first came out, everyone internally and every adopted it, and I was very convinced that [13:12] it would be [13:13] a everyone has their own agent and there's like some real really interesting [13:18] things about that world. [13:20] of a parallel org chart [13:25] agents, [13:25] in that world sort of become little reflections of you which is like really cool and really interesting it's like [13:30] Did you ever read The Golden Compass? [13:33] It's like having a little demon on your shoulder. You know that's a little part of your soul.
I really think like that's sort of what it looked like was happening. [13:41] And so I was very into personal agents, and I have completely flipped. [13:47] And I really think that. [13:50] The model for now. [13:52] is going to be a super agent, like one agent for the entire company. [13:56] And you're starting to see this in some companies. So like Shopify very famously has one. Ramp has one now. And I think there's some like really interesting reasons for that. [14:07] I actually still think that the personal agent thing is coming.
[14:11] Thank you. [14:12] But... [14:13] What we found, [14:14] is there's all this hype with OpenClaw, everyone's like, "I'm gonna set it up, it's so cool," or whatever, and then everyone realizes it's way too much work. This thing breaks all the time, I gotta fumble around with it, I gotta be able to SSH into my server, and blah, blah, blah. [14:28] And most people [14:31] to do work at least, just don't want to spend that time or can't. [14:35] Um... [14:37] and [14:38] The fundamental underlying thing that drives that is...
[14:43] Whether it's OpenClaw or any other harness, [14:45] in order for [14:47] an AI agent to be useful right now, it really needs a human who cares about it. [14:52] It really needs a human personal connection with someone who's like, [14:56] Watching what it does and make sure that it's doing the right thing and that it's useful for people and the minute you like sever that connection so the minute someone's like ah like I don't. [15:05] I don't want to maintain this dumb open claw is the minute the agent is not really that useful anymore.
[15:10] And [15:11] That's why I think it has started to shift to a more... [15:15] uh, [15:16] one agent per company model because for now, the ideal is [15:22] You basically set up a forward deployed engineer or someone with that sort of profile who's responsible for making sure that that agent is working for the whole company. [15:31] and then maybe you have some little team agents. [15:34] And I think as the models get better at being more independent, that will like shift down and you'll it'll be more likely that we'll have more personal agents because we don't have to fuck around with all the internals.
But yeah. [15:45] The model that I see working for us and for a lot of other companies, including the model companies, the model companies themselves are starting to see this. [15:53] is... [15:54] When it comes to the sort of like async agents, it's really a... [15:58] you know, you have one agent at the top that's like doing sometimes it's everything. A lot of times it's, um, [16:04] a particular kind of job that you've decided that everyone in the company is an agent for like data requests. [16:10] and then I think it will start to, it starts top [16:15] at the top and then it sort of starts to trickle down where you may get more specialized agents and teams and all that kind of stuff.
[16:20] And the mechanism is agents need people who care about them. [16:24] That is so interesting, that point about you need to like garden your agent. [16:28] Because there's context you have to keep adding to it. There's like, it breaks, as you said. And it's just like, once it's just too much work, you're like, okay, forget this thing. I'm going to go back to it. [16:36] Codex or Collider or something like that. [16:38] Exactly. Okay, cool. So this is a cool opportunity. So the idea, so what you're predicting here is companies will have this super agent that everyone can talk to.
[16:46] I said Shopify's got river, I think it's called. What's the ramp one called? [16:49] I can't remember. Okay. It's probably got a fun name. [16:52] Okay, so that's the prediction. Okay. [16:55] That's the first prediction. We will start with agents at the top. [17:02] that [17:03] that are more general and are used by more people in the company and then it will start to kind of [17:09] grow down as people get more used to these use cases. [17:13] They get more specialized and agents become... [17:17] less less fiddly like they just work better and is this mostly gonna be in slack do you predict [17:25] For work, yeah, it seems to make sense.
I think people love having the green bubbles on OpenClaw. Sorry, the blue bubbles on OpenClaw, like if you can use it with your iPhone. But... [17:34] I think there's this little thing in people's heads where they really like to keep their personal and work agents separate. [17:40] And I think there's a whole territory. Our COO, Brandon Gell, calls this computer errands. [17:49] There's like this whole territory of [17:51] using personal agents for your computer errands it's like [17:55] order my groceries or whatever. And it's like, there's so much of that, that I think this is going to, it's going to be huge for, but yeah, [18:00] Um, I focus, we focus mostly on the work stuff.
Um, [18:04] And I think that's going to happen mostly in Slack. [18:07] Sweet. Go Slack. [18:08] Should we do want to talk about the the other work surface codex co-work? Okay, this is the let's do it. I'm so excited about this. I think it's the coolest thing. So basically, [18:19] What happened was [18:21] anthropic [18:22] realized. [18:24] at some point that [18:26] with Cloud Code if you put an agent on your computer. [18:29] and it runs on your computer. It has everything, it has access to everything that you have access to.
It uses the terminal, so it has, like, basically super-powered access to it, and [18:38] Not only that, these agents really understand how to use the terminal because there's so much [18:44] content online about about that. [18:47] And it created this like super powerful coding paradigm, which is, you know, Anthropic was really doing it first. OpenAI for a while was, in my opinion, like very, very behind on this. And then in my opinion, has surpassed them recently. It's really interesting. [19:03] But they were [19:05] Very early on this, when people were still thinking about [19:09] coding agents or coding models as being really pair programmers.
They were among the first to be like, [19:15] No? [19:15] and do it successfully. There are people before them like Devin who [19:19] I think had the big cloud environment and OpenAI tried this too, but [19:24] But the real adoption seems to have happened when you put it on your computer. [19:31] So they figured that out. And then I think they figured out along with our community, [19:36] that [19:37] Once you have a coding agent on your computer that can build anything, it's actually really good for any kind of work you want to do.
[19:43] and people started [19:44] just hacking cloud code essentially to do all their work. So Anthropic then built co-work. [19:50] which is a little bit of a nicer wrapping around cloud code, but it's fundamentally the same thing. [19:56] And then I think... [19:59] You know, I think. [20:00] OpenAI made a couple of different bets, but [20:03] Their main bet on a programming agent was the earlier versions of Codex were, like, very technical and they were, like, super smart, but they were, like, a little bit autistic. Like, it was a little hard to...
[20:14] They didn't quite get what you meant. They got exactly what you said. [20:19] And I think maybe like three or four months ago, around the time that they launched 5.3, [20:25] they, [20:27] started to move in this direction of, oh, no, we get it. Like, this model is fast. It's, like, really good for general purpose, knowledge work type tasks. And then they launched the Codex Desktop app. And I think the Codex Desktop app takes... [20:40] if you look at all the lessons that like, [20:42] Anthropic learned. They went from Claude Code to Cowork.
[20:46] And you can kind of see that in the tabs on the Anthropic desktop app UI. [20:52] I think OpenAI was just like, we see where this is going. Like, let's just skip to that. [20:56] And so I think Codex right now, this is a horse race. Like they're going to have different positions. Um, [21:02] But I think Codex right now is my daily driver. I spend all my time in it. Basically, I flip the cloud every once in a while. But I think they're getting the paradigm right. And it's clear to me that whoever is in the lead, because again, I think it'll change.
[21:17] Whoever's in the lead, [21:18] it feels very obvious to me that [21:21] all of the work that you do is going to be in one of those surfaces where [21:26] For example, when I'm writing a document, [21:28] Codex has a browser. [21:30] in the app. It has an in-app browser. [21:34] And when I'm writing a document, I just go into one of my codex threads, which I have one thread for every project. [21:40] And I just open the in-app browser. I go to the document. I usually do it in proof, which is this online Markdown editor I built.
[21:48] And then I just have codecs. [21:51] running and watching me in proof and Codex can see what I'm doing. I can see what Codex is doing. It's all kind of in one place, which is the an extension of the same [22:00] thing that made Cloud Code work really well originally. [22:03] And I basically feel like I have this parallel... [22:06] work buddy that not only can it like respond and write in the document but then it can go do research it can go it can use my computer to basically do anything that I can do on my computer and that's like [22:17] incredibly powerful.
[22:19] And I do this with everything like I've been in. [22:22] I've been at Inbox Zero for like [22:24] 10 days straight now, which [22:26] If you know me, [22:27] is crazy. I never like this. And that's because I literally just have [22:33] Codex. [22:35] gather all my emails with Cora, which is our email agent, and then it renders a little page. And I think I showed you this at the Anthropic event. It renders a little page and I just like monologue into it and just talk at each email. I'm like, okay, go research this.
[22:51] Here's a question from our lawyers. Can you go collect all of the [22:55] you know, documents for the last like four years and then put them into a report and send them and it just does it. And so all the stuff that I would procrastinate on, I don't really procrastinate on anymore. And so I feel like there's this. [23:07] For a long time, we thought, I thought too, that the optimal experience of AI was going to be [23:13] Take AI. [23:15] and put it in a browser. [23:16] Okay. [23:17] And I think the reverse is actually starting to happen and be like really, really valuable in a way that I did not expect, which is.
[23:24] Take the AI agent that you use all the time on your computer and put a browser in it so it can see everything you're doing. [23:30] And that is just like a magical combination that I think will be is very uncommon now. You can't even do this in cloud in cloud code because they don't let you browse external websites inside of cloud code. So it's very uncommon now, but I think it will be super common in a year. [23:45] This is more profound than it may even sound. What I'm hearing is...
[23:49] instead of AI being baked into... [23:52] SaaS tools, which are predicting here is [23:55] uh, [23:55] the SaaS tools will run within... [23:58] codecs or cloud code. [24:01] That is one really important second-order effect of this is – [24:07] Um... [24:09] Okay, so... [24:11] Yeah, like I'm using proof or really any website, maybe post hog or whatever. [24:16] Thank you. [24:17] And I'm doing it inside of my agent. And the agent has access to the website. So it has access to everything that I have access to. [24:23] and it has access to my whole computer.
[24:24] When I run the agent on that website, I'm using my tokens. [24:28] I'm not using the vendors tokens. I'm not using the apps tokens. And so it puts SAS back in this place where [24:35] Yeah, you want to make it friendly for an agent. And everyone's got a CLI now. You want to make the HTML really usable. You want to make sure that anything that happens in the CLI shows up for the user immediately, all that kind of stuff. There are a lot of issues to... [24:49] to deal with.
[24:50] But once you do that, you actually don't really need to [24:54] think about having a [24:56] an AI surface that's primarily going to be the thing that users use. [25:00] in the sense that you don't need to build an agent necessarily into your product. I think you can, and there's another really interesting bifurcation of this that we should talk about, which is that having two agents is better than one. [25:12] Um, but, [25:13] I think for now, there's this really cool thing where, [25:18] With Proof, for example, anyone who uses it, I don't pay for tokens because they bring their AI to Proof.
[25:26] And so it changes what you build as a SaaS company, and you build it now for both humans and agents to use at the same time. [25:33] And it changes your margins back to, well, I don't really have to pay for tokens anymore because the user is going to bring the AI. [25:39] So I think this is a huge deal. So what you're describing here is, [25:42] More and more work that we do more and more professional work is just going to happen within [25:46] codecs or cloud code [25:48] Where does Cursor fit into this?
Is there potential there? That's a good question. I think that Cursor sees a lot of the same stuff. [25:57] And in some ways they have some of the same stuff, but it's better. Like I think that Cursor's cloud implementation is better than either OpenAI or Anthropix and is more advanced. [26:08] and [26:09] I think that Cursor has... [26:12] at least so far, [26:13] more distinctly chosen a lane, like they're more distinctly choosing to be for programmers, [26:20] and [26:21] that may limit how far they get in here. Like I think the definition programmer is expanding enough that they'll have a big market, but I don't know that they're going to jump into like, okay, use this to make a slide deck or whatever.
[26:32] But it is really clear that [26:34] that [26:35] Every model company is starting to realize how important it is to have a harness to get the most out of the model. [26:43] And so where all the platforms are moving, [26:46] is to a world where [26:48] you're not just doing prompt and response when you call the model on the OpenAI platform, the Anthropik platform. [26:56] they're literally running the model on a computer that is in the cloud that they run, and then giving you the result out of it. [27:03] And they know that they, in order to get the best results of the model, they need to offer that.
And so you see, [27:09] Anthropics got cloud managed agents, [27:12] um open eye does not have a have a response yet but i assume that that's going to happen and now cursor was just essentially acquired by spacex it's not like a full acquisition but it's close so i think people are starting to realize like i can't just do the like model part of it i have to have this like [27:29] harness above it and [27:31] I think the ultimate form of that harness is like, I can do any kind of knowledge work.
Cursor itself is feels like one of the things that it's going to be a hard decision for them whether to stay just for coders or not. [27:42] So people building products that aren't open AI or anthropic, [27:46] if this proves to be true, [27:48] The prediction here is they're going to be using your product [27:51] over time inside of [27:53] one of these agents. [27:54] Is there something you would do if you were one of those companies to prepare for that future? [27:58] I would just prepare for that. So like, you know, for example, um,
[28:03] every more classic piece of productivity software, whether it's Slack or, [28:09] Word docs or PowerPoints or whatever it's [28:12] really mostly meant for [28:15] A human to use? [28:17] Um, [28:18] And now people are doing CLI, so it's meant for... [28:21] an agent to use independently of a human. [28:24] And we're moving into this new paradigm, I think, where [28:27] the human and the agent are on the same level. [28:30] piece of work. [28:31] together and they're both doing things and you need to have [28:34] I need to have visibility into what the agent is doing.
The agent has to have visibility into what I'm doing. We have to go back and forth in this sort of like seamless way. [28:42] And the kind of software that you make for that is going to be very different. So for example, [28:48] Um, [28:49] Like there's a lot of stuff that proof doesn't have. I don't have to have a lot of the like word document kind of like formatting or page breaks or like, you know, making tables or whatever because the agent just does it. I don't need to worry about that.
It can do all the formatting for me. So you can make the products a lot simpler and faster to start than the legacy products are. [29:07] And then there's all these other affordances that you need to start to have because the way agents interact with software is very different. So, for example... [29:17] Agents can do a lot. [29:18] At once. They can just do like a billion different things to your document or your slide deck or your code base or whatever. And... [29:25] how you display that to the user is gonna be very different than the way you might display [29:30] a human being concurrent in your document and doing stuff you need um you need like approval
[29:36] You need a sort of inbox that sort of summarizes, here's all the stuff that's going to happen or has happened. [29:42] You need logs and the ability to roll it back real quick. [29:45] So there's all those kinds of considerations that [29:48] um, [29:49] that change the actual product and then the underlying UX of it or the underlying infrastructure you need is different too because [29:55] you know agents can make a billion requests in like three seconds so how are you going to deal with that right um this is exactly why you know github is having problems right now because because the the number of people using github has skyrocketing exponentially and it's really just people's agencies in github so i i think it's a this whole new world that is just starting you're just starting to see like a little peak of it but there's so many cool things about it so for example [30:18] in proof [30:21] and some of our other products too.
[30:23] When someone has a problem, [30:25] They don't email support. [30:27] their agent... [30:29] sends a bug report. [30:31] And an aging bug report is way better than a human bug report. [30:35] It has like here's exactly what I did. Here's the exact repro steps. Here's like [30:40] Proof is open source. So here's what I think is going on in the code base. And then we just get that. It becomes a GitHub issue. And then we can just like send off an agent to fix it. [30:48] And [30:49] Um, [30:49] You can't do that with everything, but it's so much better.
[30:53] And you can see the glimmers of this very fast [30:58] closed loop. [31:00] between [31:01] I ran into something, a paper cut, a little feature I want, a little bug, and my agent just goes off and talks to the company agent, and then the company agent just goes and... [31:10] fixes it that I think is incredibly cool. [31:13] So is a part of this that a lot of people are moving to a CLI and trying to work from the terminal. Is part of this prediction that people will shift away from that and back to...
[31:21] actual UX with agents kind of running alongside them. CLIs are over. [31:25] We speed ran the CLI era. [31:29] It was nice while it lasted. [31:31] But I think it's pretty clear that [31:33] It's not that CLIs, sorry. It's not that CLIs are going to completely go away. Obviously they've been around for the last like 30 years or 40 years or 50 years or whatever. [31:42] they will continue to be around. [31:43] And I think there is this moment when cloud code was like, so, [31:47] popular [31:48] And [31:48] or when Cloud Code was really starting to gain in popularity, that people were like, the thing that's working is the fact that it's the CLI.
[31:56] And I don't think that's what it is. And when you move into an actual UI for this, you start to realize... [32:03] Um, [32:03] We made GUIs for a reason. [32:06] and [32:07] It's just nicer to be in a GUI. [32:09] And you can get all the same benefits inside of GUI, especially for non-programmer work. But I would... [32:17] I would estimate that [32:19] Definitely... [32:21] The majority of the technical people inside of Every are not using CLIs anymore as their main work surface. [32:26] I think. [32:27] A lot of programmers are still flipping into it every once in a while, but it's more or less they're using...
[32:32] codecs, cloud code, cursor, that kind of thing. [32:35] Awesome. Okay. I definitely wanted to make that part clear. So coming back to kind of the big picture of the prediction here. [32:42] there's kind of these two modes of work that you're anticipating. One is this kind of super agent within a company that you chat with through Slack, most likely that can go off and do work and answer questions. [32:52] And then there's on your computer running codecs or cloud code. [32:56] And within that, all the work... [32:58] that you normally do kind of on your computers now are going to be living within codex or clock code, or maybe some third party that emerges that we're not even aware of yet.
[33:05] Yes, and you're going to use apps inside of the internal browser of those tools. [33:14] Wow. Okay. Like, [33:16] Listening to you talk about it, it's like it may not feel as profound as it is because this is a big change to how we work. We don't currently have an AI that we talk to. [33:25] regularly in Slack. [33:26] And we also don't work currently mostly in Codex or Clark code. [33:30] So this is actually a pretty massive shift. [33:32] I think so. [33:33] Thank you. [33:34] Is there anything else along these lines before we get into our next prediction?
[33:38] Well, a few things. I'm definitely not an agent maximalist. Like, I really think we're going to have a lot of different agents that we use. [33:44] Seems pretty clear to me. [33:45] And I really do think that two agents are better than one. [33:49] So, [33:51] That's a good example. [33:53] When I have codecs, [33:55] interact with another agent. [33:57] It can... [33:58] give so much more context about me and what I want, than I would be able to type. [34:05] And it can go back and forth talking about things that would take a long time for me to express directly to an agent.
[34:11] that you get this like speed up effect. [34:13] when [34:15] you assume that your users are using Codex or Cloud Code or Cowork as their basic way they access your app. [34:24] And [34:25] A really simple example, [34:27] We have this... [34:29] hosted open claw product, which we had on waitlist. We actually had to pause it because [34:34] We started to do all the waitlist and OpenClaw is just a very hard agent harness to... [34:39] to make work. It's like [34:41] It's moving so incredibly fast. [34:44] And if you're like a platform for it, it just it's like when things break, you can't fix it.
It's very hard. [34:51] But one of the things that we learned in that process [34:54] is if you're, let's say you're building an agent product or any new software experience, [35:00] what you would assume, let's say to set up an agent, is you need to build like a little like web interface or a little slack workflow that asks people about, OK, like, [35:11] Who are you and what are you going to use this for? And like, what's your, what's your ideal? [35:17] you know, dream outcome or whatever the things you are that you would put on an onboarding checklist.
[35:23] Thank you. [35:23] If instead you just make a hard line of, [35:28] we are only going to service users. [35:31] who use codecs or co-work. [35:34] So, [35:34] what happens is you just paste something into, you just paste a prompt into Codex or Cowork, [35:41] It goes and talks to the app and the app can be either just a regular server or it can be its own agent. [35:47] and Codex has so much information about you that it can just give it, here's all the stuff I've been working on with Dan, here's all the ways that he might [35:58] he might want to use this app and then bring it back to me.
And it's this very custom experience. [36:03] And also for a technical product like an agent, when something goes wrong, [36:07] I can just tell Codex, go fix it. [36:09] and Codex will go talk to the app and figure out what's going on for me. [36:13] And so I think the whole paradigm starts to change when you assume that everyone's got an agent and those agents are talking to other agents in this like really magical and important way. [36:22] There's a couple more things I want to touch on before we get to it.
There's so much to talk about. One is you made this point about [36:28] SaaS tools not using... You can use tokens from [36:32] the model companies basically when using a SaaS tool. Talk a bit more about that because that may change [36:37] the business model for SaaS companies in the future, that feels like a big deal. [36:40] Well, I think it actually... [36:42] may save their margins. [36:46] Because right now all these companies are rushing to like add an agent to their offering. [36:52] And thinking, oh, the agent is going to be the main way that people interact with me.
[36:57] And [36:58] I think that [36:59] And that costs tokens, obviously. And I actually think, [37:02] Once I have codex or co-work as my main work surface, [37:07] I still want to use SaaS. So this is another good prediction. I would buy SaaS stocks. [37:11] right now. [37:11] I think the Sasspocalypse is dumb. [37:15] and SaaS stocks will be up majorly in the next couple of years. [37:19] Non-investment advice, but [37:21] You know, I would buy SaaS stocks. Um, [37:25] So – [37:27] So... [37:29] So I think it saves your margin because now what you're what the way that you're thinking then is not I have to build AI into this is it's more like.
[37:38] I have to make a piece of software that humans and AI want to collaborate on together. [37:44] And [37:45] That's hard, but once you build it, it's a lot cheaper than assuming everyone's spending tokens. [37:52] And, um... [37:54] I think it's a good business. And part of the reason I'm so bullish on SaaS is [38:00] A. [38:01] Everybody internally here is, like I said, we've all got agents and we're all using codecs and whatever. And we still pay for a ton of SaaS. And our SaaS spend is up year over year.
[38:12] So [38:12] And we're not like vibe coding every single like little thing, you know? [38:17] And, [38:19] I think that what agents do is increase the number of users of SaaS. [38:24] not get rid of it. [38:26] And so I think SaaS companies are going to see like an insane spike in the amount of demand that they have because there's going to be tons of agents using these products at like a very high volume. [38:36] And like I said, that's a huge infrastructure challenge. [38:39] There's a lot of interesting pricing challenges.
[38:42] But it makes me very bullish on SaaS. [38:47] I love that. If anything else comes out of this conversation, Dan Shipper, SaaS is the future of AI. [38:54] B2B SaaS. Hashtag send tweet. [38:58] I love just, yeah, this is quite contrarian and [39:01] The other interesting piece is that the fact that you guys are hiring, that you doubled in people in the past year. [39:07] which is not what people would have expected from a company that is so AI forward, [39:11] Talk about your experience there of just, okay, we still actually need humans.
Automation is a lie. [39:16] um [39:17] In the sense that every time you automate something in order to make sure the automation is working well, you need a human on top of it, like making sure that it's working well. And so. [39:27] You know, I wrote this piece a couple of years ago called the allocation about the allocation economy, like the idea that. [39:32] the way that humans are gonna work with AI is gonna be [39:36] like being a manager. [39:39] And [39:40] The thing that you have to remember about managers is like managers actually spend a lot of time working.
[39:45] most managers are not like on the beach. They're like, [39:49] checking in with their employees all the time and trying to figure out, okay, how do we make this work good? How do we make it better? How's it doing? How's this person doing? All that kind of stuff. And I think there's some differences between being a human manager and being a model manager, but [40:02] fundamentally it still requires a lot of time and attention [40:06] and [40:08] I think that we kind of miss that in the model discourse. And one of the reasons is [40:15] benchmarks make it look like AI is more autonomous than it is.
[40:20] And by autonomy... [40:22] I mean something specific by autonomy and I'm going to try to express it. It's like a little hard to express, but [40:27] I learned this for myself because I've been feeling this paradox a little bit. I've been feeling that like we have so much automation, so much AI, and I also work way more. [40:36] And I think part of the paradox [40:38] Part of the paradox started to resolve for me a little bit when I made my own benchmark. [40:43] So I made this senior, it's called the senior engineer benchmark.
And it's like, how good is AI versus a human engineer? [40:51] And the way that I built it [40:53] is, again... [40:55] I had this app proof, I just vibe-coded it on the side, and like while running the rest of Every. And when we launched it, because it was completely vibe-coded, it just started going down and I couldn't fix it. And it was very embarrassing, I had a lot of egg on my face. And like the product worked, we tested it internally, [41:12] We had... [41:14] A lot of beta testers, but like the day after launch, it was like, [41:17] And just every like 10 minutes, the servers would go down and people were looking at me and I'd be like, I don't know what's going on.
Like Codex, fix it. And Codex is like, I don't know what's going on. [41:25] Or really Codex is like, I do know what's going on. I fixed it. And then [41:30] it would cause four other errors. And then you're just going around in a circle and I wasn't sleeping. And I vibe coded so hard I got bursitis on my elbow. [41:38] So there's a life lesson in there. Vive Coder Elbow. [41:43] So anyway, I got actually two different senior engineers to fix it independently. [41:51] So I have two different rewrites of the code base that, [41:56] tells me how they did it.
[41:58] And so what I get to do is when we get new models, [42:02] I just give the new model a prompt. I say like, [42:05] This is Vibe Coded Slop. [42:07] If you wanted to rewrite it from first principles, how would you write it? Go do it. [42:11] Thank you. [42:12] And [42:14] All the models until GPT 5.5 got like a 30 out of 100. [42:19] And senior like a human senior engineer gets like high 80s, low 90s out of 100. [42:24] So there's a lot to go. And then I tried GPT 5.5 and it got like a 62.
[42:29] And mind you, the 60 score was GPT 5.5 using an Opus 4.7 plan. Opus 4.7 plans are very good. [42:39] GPT 5.5 is the only model, though, that has the sense of agency and confidence to just rip out old code and actually rewrite from first principles. [42:48] Other coding models, they kind of like try, they like end up papering over the edges or around the edges and they're like, oh, this is a big job. Like, I'll just do a little patch. And you're like, no, I like specifically told you not to.
[42:59] So GPT 5.5, there's like a [43:01] 30 point bump in the score 60 out of 100 it's like [43:05] very it's very clear that in [43:09] a year or less. [43:10] It's going to be. [43:12] senior engineer level. [43:13] Right. And that gives you a certain picture in your mind, especially based on how I named the benchmark, which I think a lot of benchmarks do. [43:19] And I can tell you that when we get to that, [43:22] point. [43:23] it would be very easy for me to change the benchmark to zero out the current model.
[43:28] So that gets a zero out of 100. [43:30] Yeah. [43:30] And so, for example, [43:33] it seems like there's no skill or no thought into the prompt, which is, this is vibe-coded slap, like, [43:40] fix it from first principles, but actually it took me a while to get to a prompt that [43:45] didn't give away the answer. [43:47] But, uh... [43:48] but got the model to reveal [43:51] what it's capable of. [43:53] and the original prompt I gave it was the original prompt that I gave it when, uh, when I was trying to fix the issue and production was going down, which is like, I'd woken up, I'd, I'd woken up in the morning, [44:05] Thank you.
[44:06] And I was like... [44:08] Okay, we had four or five reported issues yesterday. I want you to go through all the issues and then come to like a... [44:14] make a plan for how to resolve all of them, [44:16] and go do it, right? And every coding model on the market, and I'm pretty sure, [44:23] Here's a prediction, I'm pretty sure every coding model on the market will still do this in a year. [44:29] Every coding model in the market will take that instruction seriously. And if I tell it, here's a bunch of issues, go fix it.
[44:35] They will just go try to fix the issues. [44:37] What an actual human senior engineer does is they go look at the code base and they're like, this is a piece of shit. [44:43] This guy doesn't know what he's doing. And then they say, we're going to have to like actually rewrite a lot of this and it's going to be hard and risky. I know you don't want to hear that, but like we're going to have to do that. [44:53] And [44:54] If you asked the model, "Hey, should we do that?"
It'll probably get there. [45:00] but it's not gonna do it on its own. [45:02] And there's a lot of incentives pushing against it doing that. And even if it does that, there's always a higher frame for us to go. [45:11] And so I think it's really important [45:13] When we think about benchmark progress to, um, [45:17] Think about it from that perspective, which is benchmarks, [45:20] Rise on problems that we've framed that we can articulate that we can score and there's a lot of work. That's human work that is [45:28] It.
[45:29] it can't be scored until you write it down, but the act of thinking to prompt it or write it down, um, [45:34] is something that you can't measure, but like kind of means that even if the benchmarks get saturated, it doesn't mean the same thing as you totally replace all senior engineers. And I think it's why [45:47] even though the models are getting better at automation, [45:50] I still hire engineers. I am so excited to tell you about this season's supporting sponsor, Vanta. Vanta helps over 15,000 companies like Cursor, Ramp, Duolingo, Snowflake, and Atlassian earn and prove trust with their customers.
Teams are building and shipping products faster than ever thanks to AI. But as a result, the amount of risk being introduced into your product and your business is higher than it's ever been. Every security leader that I [46:20] of protecting their organization, their business, and not to mention their customer data. Because things are moving so fast, they are constantly reacting, having to guess at priorities, and having to make do with outdated solutions. [46:33] Vanta automates compliance and risk management with over 35 security and privacy frameworks, including SOC 2, ISO 27001, and HIPAA.
This helps companies get compliant fast and stay compliant. [46:47] More than ever before, trust has the power to make or break your business. [46:51] Learn more at com slash Lenny. And as a listener of this podcast, you get $1,000 off Vanta. [46:57] That's com slash Lenny. [47:00] One thing I mentioned recently on the podcast, I heard that speaking of the code that you have of like humans writing code. [47:06] uh, [47:07] data labeling companies are buying code that was written before 2021, 2022, before AI became a thing is like, [47:13] very valuable data.
Artisanal human code. Yeah, exactly. That's exactly right. And it's so interesting that that's exactly the kind of code used to build this model. [47:23] Well, what's interesting, so I want to clarify there. So I did not have a human... [47:29] write the code all by hand because I actually think that that's sort of it feels silly to me [47:34] Like, I don't really care because I know if an engineer is not using AI, like, I'm not going to work with them. [47:41] I don't really care. It's like, it's sort of like, am I going to race...
[47:45] a human against the car like I probably wouldn't do that. [47:48] But I would race a human in a car versus another human in a car and say which one's better. [47:54] And in this case, the way the benchmark is structured is, yeah, like these human engineers used AI, but. [48:00] They used it in a way that I could not because I didn't understand it and I didn't have time and I didn't really want to go in and try to understand the code base to be honest. [48:07] And I think that's a really important thing when we think about benchmarks is, you know, [48:12] AI is a broadly distributed technology that any human can use.
And when we... [48:17] are benchmarking against humans [48:19] AI against humans are actually really always talking about one human using AI versus another human using AI because AI doesn't use itself. [48:27] it may be able to in this like slightly somewhat recursive way, but there's, [48:32] In any real use case, there's always a human pretty close to it, making sure that it's working. [48:36] Okay, I want to try to wrap up our first bucket. There's so much to talk about. I made a little list of things that I think people should do based on your predictions to be successful.
We'll talk about this at the end too, but just a few things. One is start using codex or clock code more and more for the work you're doing. [48:53] And... [48:54] especially the browser, use tools inside of it. [48:57] uh two is allow your age allow agents to be to use your products if you're building a SaaS tool [49:03] make it easy for agents to be a [49:05] a user. [49:06] essentially. [49:07] Three is [49:08] start thinking about some Slack bot that you can work with, like try out tools. [49:13] I know Slack has their own Slack bot that I think is really good too, and I haven't played with it.
[49:17] People really like it. [49:18] So look for, I guess, a tool that could become the AI agent within your company. [49:23] uh, [49:24] buy SaaS stock ASAP. Not investment advice. I think that's totally right. My slight tweak is [49:34] when you're thinking about
Want to learn more?
Ask about this video