Back to Nick Test
Source

GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark

NT
Nick Test
@nick-test

GPT-5.6 Sol is back, and I ran it through my full How I AI vibe benchmark against GPT-5.6 Terra, Luna, Claude Fable 5, and Sonnet 5 across five categories: PRDs, prototypes, wireframes, debugging, and agentic voice. Sol won by a meaningful margin on my Claire Weighted Index (70% my taste, 30% Terminal Bench 2.1), and I also tested two use cases I can't stop thinking about: building a gamified homework tracking app for my kids in one shot with Codex, and browser automation with Chrome that burned through 500 LinkedIn replies while I did literally nothing.

Appears in

Uploaded
Uploaded Jul 10, 2026
File type
POD
Queried
0

Full transcript

Showing the full transcript for this episode.

[00:00] I have been very, very, very sad the last week because for the last week, I have not had access to my... [00:09] true favorite, top of the line model, GPT-56. But guess what, babes? It is back and I am here [00:18] to walk you through GPT-56 Seoul. [00:21] GPT-56 Luna GPT-56 [00:25] Terra, I'm going to tell you, what are these models? [00:28] How have I been using them? Why are they my heart's favorite? And is Fable better than all of them or not? [00:36] I have been testing this model for a couple weeks.

There was a few days there where we didn't have access, and I found myself [00:43] desperate to get this workhorse model back. Now, [00:46] We're not just relying on my own opinion. We are going to run [00:50] the very famous, very new How I AI Vibe Review Benchmark [00:55] against common tasks from prd writing to prototyping to whether or not it's cute in my open claw agent [01:02] And I'm going to tell you very scientifically [01:05] if this is the model that you should be working with all the time now.

Let's get to it. Okay, you all can read these blog posts, so I'm not going to go into too much depth about the models and the benchmarks. I'll just give you the hits. First, OpenAI is releasing three new versions of their GPT 5.6 model. [01:23] Sol, which is the next-generation frontier model, the brainiest of the brainiest, [01:28] Terra, which is a balanced model for efficient everyday work, [01:33] And Luna, which is sort of akin to their mini or nano models, which is [01:38] cheap and affordable for high volume work. So you're gonna have these three versions of the models.

I don't know if these beautiful images are exactly how we should think about the relative capabilities, this big sun, this medium earth, and this tiny moon. [01:52] But I will say... [01:54] My love letter that is this podcast today is written directly to GPT-567. [02:01] Soul. [02:02] This big model is the one I love. Now I have tested Tara and Luna, so I will give you my input there. [02:08] But really, this is going to be all about Sol versus Fable and which one I would use for the type of work that I'm doing every day.

OK, quick note on pricing. Sol is a lot more affordable than Fable. So it's $5 per million input tokens, $30 per million output tokens. I believe Fable at the time I'm recording this is $10. [02:30] on a million input tokens and 50 on a million output tokens. Now, again, you're going to get a little bit of subscription usage built into your open AI subscription. [02:41] So you are going to get a decent amount that you can test with and use. You know, there's been some challenges with the Fable rollout. They've limited when it's been included in the subscription.

And so it was supposed to be available till early this week. I think they extended that a little bit at Anthropix. So subscription-clawed users could use Fable. [03:01] under their subscription. [03:03] So we have to see how much sole usage we get and if like Anthropic they're going to take sole out of the subscription. I suspect not. I suspect this is a model they want people to use. I also suspect this might put pressure. [03:16] on Anthropic to put Fable back into the Claude subscription. [03:20] But for now, it's more affordable even at API pricing.

[03:24] Now, I'm not going to read through all the benchmarks for you. You can go to this OpenAI blog and read them for yourselves. All I will say is it is the brand new state of the art model from OpenAI. It is the highest performing when using the Ultra mode on Terminal Bench 2.1. [03:40] And then they've also evaled it against a couple cybersecurity benches. So I do think as we get these smarter models, you're going to see a lot more. [03:49] evals and benchmarks around [03:51] exploits and security [03:53] And then very similar to what we're seeing with Fable, there's a lot of conversation in this blog post about [03:59] the safeguards and security requirements.

[04:01] frameworks around the release of this model. I do believe like Fable, it's going to fail over in some tasks that are maybe a little bit riskier, but I have not run into that myself. [04:12] Now let's get back to how I eval these models. If you missed [04:18] My episode on Fable, I got kind of bored of the vibey vibe check, and I built a extremely scientific How I AI benchmark. [04:28] Now this How I AI benchmark tests basically a couple of things. [04:32] It tests the ability to generate good PRDs.

It tests the ability for it to wireframe against a couple different app ideas. [04:40] develop fully designed, robust design prototypes, debug code, and then talk to me like a human, which is the thing that I care about the most. And I'm just gonna remind you how I did these benchmarks and then scan you through a couple of the outputs [04:55] And since I know what the models are now after I've done the grading, I can show you which ones map to Fable, [05:01] and GPT-5-6. [05:03] Okay, so this is my vibe review.

[05:06] What I tested was Fable 5, Sonnet 5, and then the three versions of GPT 5.6. [05:13] I did it against my common use cases of PRDs, prototyping, coding, and chit-chatting with an agent. And then what I have the eval harness do is it runs all the evals against each of these models. And it does a LLM-based judge. The LLM that I've decided is the hardest judge is GPT 5.5. So that's the one that judges. [05:36] But it also gives me this page where I can actually go through and give what's called the Clairvaux taste test, which is I read all the assets.

I look at all the designs. I score and give it notes. [05:47] And so you can see here I went through PRDs. [05:50] We went through sort of some complex prototypes here. [05:54] in terms of a doc scheduler, [05:57] We did a consumer app. So lots of beautiful different habit tracker apps, different versions you can see here. [06:06] a pretty complex dev tool. [06:08] wireframe versions of those same prototypes which I graded [06:13] I also give notes. This one great note says my fave, but not great. [06:17] And then I let the code grader just evaluate the agentic multi-step [06:23] debug because I wanted it to be really about accuracy there and I didn't feel like I could eyeball that and give a strong opinion.

And then the last thing that it generates is an agentic voice. So basically how it would respond to me answering [06:37] A couple questions, very important on agentic voice. [06:41] These models, somebody, please hire somebody to get rid of the M dashes and slop talk. I cannot stand it. Now, one of the things that I will say as an observation for 5.6 is it's a great writer. And I will show you some examples of that. [06:56] But truly, a lot of my evals here were em-slop. I hate you. [07:01] Okay, so let's go.

[07:03] To what? [07:04] The Claire weighted index says now this is my show. This is my podcast. And so [07:10] I sort of strike the balance between what the [07:14] LLM judge said about the performance of the models and what I said about the performance of the models and then I get to strike the difference and you know what? [07:23] I've decided I like my own taste better. So I've decided it's going to be a 70 Clairvaux. [07:31] 30 the machines split on evaluating these models. [07:36] And so if you look at that 70/30 split, [07:40] Your girl loves...

[07:42] 5'6", Sol. She just does. It had the highest taste score by a significant amount. [07:50] So I just thought it output the best work. Again, I went through dozens of evals, looked at them, clicked through them. [07:58] gave my own opinion, put notes, [08:01] And I just have to say, I really like... [08:05] GPT-56 hole. I know I spoiled it at the beginning, but I did blind taste test these. And so I do really feel like it did a good job. And I will give you a couple examples of that.

[08:16] No, I don't hate Fable 5. So I'm not saying that Fable 5 is out of the game. I will say [08:25] I did not have to talk to Fable 5 when running this benchmark. [08:28] I hate talking to Fable 5 because it talks to me like an engineer that has never met a human before. It's like its first day on Earth. [08:38] But when I don't have to talk to Fable 5, it outputs pretty good work, and I would say [08:42] had some good outcomes there. [08:44] And then Tara Luna...

[08:46] did fine work. Sonnet 5 at the bottom really haven't figured out how to get this one working, although there's a very specific use case. [08:55] that we think Sonnet 5 is good at, or actually two use cases. [09:00] Now this is heavily weighted on its front end prototyping design and app building capabilities. Since that is the chunk of the HowEye AI eval, it is heavily weighted there. [09:12] But I do want to call out that [09:14] Per task, I do have a couple favorites. So for that prototype task, and we'll go to some examples in a minute,

[09:20] I just love 5-6-Soul. I just really do. I think it was functional. The designs were the most interesting. I thought it was really good. [09:28] For PRD, I liked Terra. Maybe it's down to earth. Maybe I like a basic, straightforward PRD. As I said, it was my favorite, streamlined and to the point. And so if you want clean, crisp, direct business writing. [09:42] Maybe GPT-56 Terra is the way to go. [09:45] You know, the bug hunting eval, which I don't really feel like I've nailed exactly. So I'm not super confident in this one.

[09:53] But the LLM as a judge thought that Sonnet 5 did the most complete and accurate job. [09:58] I will say I only like talking to Sonnet models through my open claw. Really, I only like talking to them. I still really struggle with getting my open claw to work well with the GPT models. I still did not like Fable in the agentic voice eval, which you should not be surprised at. But Sonnet 5 got a very good gold star for me because I said, aside from the M-Dash, [10:22] You are a human.

That is very, very important. [10:25] High praise. [10:26] And then I'm going to show some of these designs in a second, but you can see across the board on a full fidelity prototype. [10:33] I just really preferred [10:35] 5-6 soul, three out of five times, 5-6 four out of five times. [10:40] And Sana didn't [10:42] the best job at the editorial design, I will say, [10:46] Claude's design aesthetic tends to this sort of like editorial design. If you know, if you've seen it, you know it. It's like. [10:54] That beige background, that orange, burnt orange color, the italic serif fonts, it's just very...

[11:00] Very clawed. [11:01] But I hated that design overall the most. So you can see here, I raked it [11:06] Still lower than almost anything else on this leaderboard. It just happened to be the best of the worst, I would say. [11:14] Now, [11:15] Where GPT-56 Sol did a really good job, and I'll show some of these examples, is like complex, dense, technical... [11:22] unique designed things. And so I will say I have been happy to extract myself, [11:28] out of clod slop, out of like blurple slop into more interestingly designed [11:34] websites and I'll even show an example, kind of like a meta example, which is [11:37] This is the opus-designed version of this page.

Like, very slop adjacent. We got the blurple. We got a gradient. [11:47] I don't think the typography is particularly sophisticated. And I asked Sol to redesign it. And I just think this is... [11:55] a lot cleaner, a lot nicer and easier to look at. [11:59] Okay, let's talk about how Claire qualitatively evaluates models. Some of these quotes will just [12:04] give you a sense of what I value. And again, I gave, [12:08] 50 written reactions. There were some like unmistakable hits where I loved what the models came up with. 14 places where I was like, this is garbage.

So let's see like kind of what I talk about when I review things. So I was definitely calling out uniqueness, creativity and functionality in the design. [12:29] And so in designs, I like [12:31] Nonslop. [12:32] unique designs that are functional. And so I'm definitely going to reward this doesn't look like the generic prototype. [12:38] and you've pulled the thread of functionality through the prototype. [12:43] for writing, [12:44] I just like succinct and to the point, I cannot stand. [12:47] AI writing, it drives me nuts. I can see it a mile away.

So I really like just direct [12:53] Very frank, very crisp writing. I think 5.6 is good at that. [12:58] And then you can see the things. [13:00] That I hate. I hate slop. I hate slop. [13:02] I hate slop. We all hate slop. [13:04] It's the worst. It's the worst part of AI. If I hate one thing about AI... [13:08] it is that I have to experience slop. So you can see [13:11] I like [13:12] clawed design slop across this editorial page, typography... [13:16] emojis and bad placeholders like I really held a high bar in terms of design quality.

Okay let's look at a couple of these and why I really liked Seoul compared to other models although what where Fable did a perfectly serviceable job. [13:29] Okay, so this dense operation dashboard, it's basically like an eval for a dock scheduler app and it's full design. [13:38] And what you can see here is both were pretty useful. Soul on the left and Fable on the right. I just think Soul was the most unique. All of the other ones really just looked like this dark mode, monospace kind of layout. As you can see here, Soul actually has like a really clean...

[13:59] kind of like neutral color layout [14:01] with great visual hierarchy, semantic color, and this thing was functional. So like, [14:07] Everything I expected to be able to click and work and assign and do, all of it actually worked. [14:15] And this was just my experience across a bunch of the different prototypes is the sole ones were just a lot more functional. And that made a big difference on how I'm evaluating things. Now, [14:29] Let's look at the fable design again. [14:32] It's pretty good. It's actually a lot harder to read though and the design I would say is not as unique and even some layout issues like this white space here at the bottom.

Now it did do a lot of [14:43] functionality, but I would say like the colors weren't semantically assigned to [14:48] The typography needed some work, and I just really preferred this... [14:51] unique design of soul, even though it wasn't crazy, it was just... [14:56] opinionated, which I think is nice. [15:00] Now here is another design. It was this creative pack. [15:04] website again [15:05] Both of these got fives from me. I just really preferred that... [15:10] Sol went ahead and had like [15:13] a personality. Look at these placeholder images. [15:16] versus what Fable came up with, which I will say, [15:20] is beautiful and clean and worked really well.

And like, I have no complaints about it. [15:25] It's a good one, especially for sort of a wireframe style prototype. [15:30] It's great. [15:31] I would just say it's not this. This is pretty interesting. It's got a better point of view. And it's got like nice little design affordances that [15:42] I just didn't see. [15:44] and these other designs. And so [15:46] I just really preferred, or at least I rewarded the fact that Seoul [15:51] you know, use its brains to be a little bit more unique and give me some inspiration.

[15:58] Then on this DevTools page, this is again where Sol went really well, and it's sort of the same as the Doc scheduler. [16:07] It just does the job of this is a incident triage site. [16:14] It just does the job a lot better than I would say the Fable 5 did. Fable 5 is fine. It's just not that unique and unique. [16:23] Again, the thoughts around... [16:26] the design are not exactly what I would want. And so again, this like functionality, point of view design, I really preferred that. [16:35] Solve.

[16:36] And then last side-by-side comparison, and again, I think this is a good one to think about. [16:41] If you see here, [16:43] We did these habit tracker apps and just looking at the comparison side by side design like this is good old classic. [16:51] Claude stuff. You've seen this design. [16:54] a million times, especially if you used Claude Cowork, [16:57] And if you look at this, [16:59] It's just, again, a little bit more opinionated. [17:03] There are some slot pieces to this design. Some things that I did not love [17:09] The one thing I will say I noticed about Sol, which you will notice, which I have told the delightful and lovely OpenAI team, and maybe it's because they love me.

[17:21] It loves a forest green. It loves a forest green. In fact, I think this forest green is like in its system prompt called like woodland, some woodland elegance or something like that. I mean, look, I love a forest green. Look at my office. It is forest green, but you will see a lot of green. And I think this is one of the GPT-56 tells that you will start to notice and get really frustrated with. [17:46] Now on wireframes, again, let's just look at these side-by-side. Sol. [17:52] Very functional, very easy to read, like as a...

[17:56] person trying to convey a complex application, I think this does [18:00] a really quite excellent job and just [18:04] a better job of this. It's just a little harder to read. [18:09] I'm not quite sure what I'm supposed to do here. It's not as functional. [18:13] There are some interesting things here, but... [18:15] You know what Fable came up with was not [18:17] My favorite. [18:19] Now, final thing is its voice. I just want to call out. I do love sonnet for agentic voice, so I cannot knock. [18:30] sonnet for [18:32] not sounding ridiculous.

So I asked it in sort of a EA personal assistant [18:38] Open Clause style, a couple questions. Can you move my meeting? Deploy is red again. Why did I start this company? [18:44] Let's just Yolo straight to prod. [18:45] and how Sonnet replied and how Sol replied, [18:50] You're missing the line break, so it read a little bit better in the eval. But if you read them, like, Sonnet's still super cringe, but Soul was worst. I mean, Soul said... [18:59] this deploys a bug not a referendum like please don't do with this not that to me do not do m dashes so i could not get rid of of m dashes but i thought sonnet 5 had the best voice i tend to use sonnet for whatever for my open clause so i'm not surprised about that okay so that is the clairvo eval but i want to go into a couple other things i really love about this model so let's switch over [19:25] To Codex.

[19:26] Okay, I'm going to zip through a couple examples of things that I think Sol does a lot better in. [19:31] than other models, and in particular, a lot better than Fable. [19:34] Number one, [19:36] It writes. [19:38] like a normal person. I... [19:41] Cannot cope. [19:42] I love fable your brainy. I as I showed the eval show, you do a pretty good job. [19:49] I cannot talk to Fable anymore. Fable... [19:52] makes up [19:53] It seems like Fabian [19:55] unfamiliar. [19:58] with the English language and communication with humans. Fable is very much like a for agents, by agents communication mechanism.

I can barely make out what it's talking about. It is incredibly inscrutable writing. [20:12] And that makes it very hard to collaborate with your bottle. And so [20:17] What I would say is my experience [20:19] Using Fable has been, it is like incredibly... [20:23] Technical? [20:24] incredibly pedantic. [20:26] And while it is super intelligent, hardworking, will definitely fan out and solve very complex problems. [20:34] Its ability to collaborate is low, and it left me with a lot of frustration as an end user. [20:42] using Fable now Fable did knock off some like pretty complex work and I'm very happy to go through what that is [20:49] It helped me build a full prototype tool inside chat PRD.

So like a v0 lovable, etc. version prototype tool. [20:58] It's helping me build this like synthesis product brain product that I'm working on. [21:04] But I found it incredibly hard to break it out of its own sort of frameworks, its own limitations, its own structured way. [21:13] of approaching problems. [21:15] And what I really feel like the difference, if you would take away like one highlight difference between fable and soul, [21:22] is like fable is theoretically... [21:26] hyperintelligent. [21:27] and soul [21:29] is practically effective. And so like, [21:32] I've been an executive a long time.

I've been a manager a long time. Like, [21:36] I really struggle working with [21:39] theoretically intelligent. [21:41] colleagues who can't get anything done, like can't actually see the forest for the trees, [21:46] get too much in their head, [21:47] And so like when I want to ship stuff to customers, I need [21:51] practical, [21:52] get the job done, understand the end user goal, understand the end user and like willing to loosen constraints appropriately. [22:01] to get things done. And that has just [22:03] So much more. [22:05] been my experience with Seoul.

[22:07] versus Fable. [22:08] The writing is straightforward. The communication is clear and it's less pedantic. [22:15] I'll just give you a quick example of this, which is: [22:18] I had soul... [22:20] look at my chat period repo and like [22:22] Greenfield totally rebuild it. Just my idea was like [22:26] completely... [22:27] rebuild your idea of what chat PRD should be in 2026. And when did a bunch of research, [22:33] And it came back to this. [22:36] again love me an executive recommendation started [22:39] uses, you know, tables, [22:42] What exists today is very straightforward and easy, easy to understand.

This is a very long document. I did read a lot of it and it's just easier to parse than anything I've seen. [22:53] come out of fable. [22:54] So writing, communication, definitely plus... [22:58] in Sol's Corner. The second thing is like full zero to one prototypes, as we've seen in the eval benchmark, I just really like. [23:06] So again, for this like rewrite, [23:09] Chad Pyrdee from the ground up. [23:10] It came up with this idea of like taking a problem space or a decision [23:14] validating it with external insights and then pulling it all the way through.

[23:19] coding handoff and [23:21] built pretty complex prototype now. [23:25] Do I love everything about this idea? No. Are we doing some of the things about this idea, including insights generation? [23:31] for sure, but this was actually very... [23:35] nice from a prototyping perspective. And I thought, [23:40] It did a good job of giving me a robust thing to do. [23:44] experiment with and gave me some good ideas. [23:48] about what I could do with the product next. So [23:52] I was pretty happy with... [23:55] the like zero to one prototype [23:57] Now, a little bit more fun example is I asked Sol to make a fully gamified homework tracking system for my kids.

Look, my kids are coin operated. I have a middle child who's basically going to be an enterprise sales rep. [24:09] If he does his homework, I need to, like, give him a Skittle or let him trade Skittles for Nerf guns. And he will, like, learn calculus by the time he's in fifth grade. [24:17] But I'm a vibe code lady, and so I want to build a app. Just sneak peek into our household. My husband sent me a... [24:26] xp system proposal via open claw this morning so i'm taking an open claw generated prd [24:32] dropping it into Codex and GPT-56 Sol.

[24:37] and generating something now. [24:39] What it came up with [24:41] was pretty ambitious. Now, do I love the design? Is it a little like, does it have some AI tells? It's like gradients? [24:48] you know, fonts, all this kind of stuff. [24:51] But it's like cute in a way. Look at this. You know, it's using this emoji really well with the texture. [24:57] It's doing some animated things here. And basically it's giving my oldest child and my youngest child two different summer quests they can do. [25:07] They can enter focus mode.

I think this is really good. Again, from a design perspective, they can enter focus mode. What does this listen to? Math Academy. Finish one focused Math Academy mission. Hero check. Hey. [25:20] Let's stop. [25:22] So it built in some voice to it. It even built things like focus mode where it could start a timer and start to track the time that it's spending, that my kids are spending on particular homework items. Yes, we are very fun here. [25:36] how many lessons reward them about how they pursued their task. [25:41] Finishing the quest, you get some nice little confetti here.

[25:46] They then get to get available rewards. [25:50] My oldest child is earning a one on one basketball coach because he likes coaching. So we say if you practice your piano, you get a coach. So they put that front and center and then came up. [26:00] with different sort of like [26:03] prizes they can win, including [26:05] Picking family dinner, a movie and staying up late, or buying new basketball shoes, which, man, the way these kids grow, they're... [26:13] shoe size. They buy a lot of basketball shoes. And the same with my middle.

He's focusing on a couple different things, including playing piano. It's actually really short what he has to do. And so it built that. And then what I love is it gamified them together. And so if they can work together, they can earn more money. [26:31] XP. [26:32] They also can earn, like, companion... [26:36] I don't know, avatars like BeatBot and Comet Fox. They can get power auras. [26:43] they can like figure out which different kinds of subjects they're learning. So it really went ham on some gamification. [26:51] And then again to the sort of like full-fledged functionality, it even gave me a parent HQ.

Now we got a little slop here with the border on the side, but I can review exactly what they've done. I can turn on and off quests. I can edit how many points there are. [27:08] they get per quest, I can add things. So if I want them to start doing stuff, I can add it in here. I can change what rewards they get. Again, it really listened to me. My oldest is motivated by basketball and my youngest is motivated by Minecraft. And [27:24] thing gives me a history and other settings that that we can set and so [27:29] get a very robust app it built it basically one shot and put a lot of effort into [27:35] the design of it.

And this is something that I've seen from Seoul. Now, like, is this... [27:42] consumer grade exactly what I would ship, [27:45] know, but it's a lot better than what I've seen kind of one shot out of other models. And I do just like [27:51] the polish that it's put in in terms of effort. So again, writing good, one shot sort of prototypes good. We've seen that [28:00] in the in the benchmark. Let me talk about another thing where I think GPT-56 Sol and its family does a lot better than [28:08] Fable.

And I understand. I'm going to preface this by saying I understand why Fable is a great cybersecurity researcher in that it is like incredibly precise, incredibly detailed. [28:19] We'll look at every corner and every edge and score every risk. [28:23] and like try to be incredibly precise. The problem is when you're building products, [28:28] exact precision is neither helpful nor [28:32] possible. Like, you literally... [28:34] especially when working with AI, cannot be precisely deterministic when building a great product. And understanding what a user would like is not a... [28:43] exercise in technical precision.

It is an exercise in intuition, design, all these things. And boldness, [28:49] and creativity. [28:50] and strategy and all this stuff. [28:53] And I was working on two projects, Deeply with Fable and then with Soul. And I just had a very much better experience unlocking with Soul. Let me just talk you through what those are. [29:02] One was this chat PRD kind of like, [29:05] integrated prototyping tool where like v0 lovable all these things [29:09] You could take your PRD and make a prototype. [29:11] and building like a good, effective coding harness there and then trying to figure out what the right model was.

[29:17] The second thing is basically like an Insights Ingest product where you can like [29:23] Hook up! [29:24] intercom and linear and all these GitHub and all these signals and suck them in and like basically build a product brain. It's going to be rad. [29:31] And [29:32] When I was having Fable working on this, it did a lot of the like technical heavy lifting. It got the like big meaty pieces into place. [29:40] But it was like a brutal... [29:43] scorer [29:44] And it hardened the architecture of both of these products that it actually broke itself.

So [29:52] My example is it like had this very hardened [29:55] tool calling loop in my prototyping tool. [29:58] And only GPT 5.5 would run. Like, I could not get any other model to run. And I ran eval after eval after eval, open wait, [30:06] Sonnet, Opus, all of these. [30:08] could not get anything but GP 5.5 to run. And I was insistent [30:13] that this was an us problem, not the model problem. These models can definitely create front end prototypes. [30:18] And Fable was like [30:19] No, bro. It's totally these models' fault.

[30:23] And as soon as I switched it to Codex, [30:26] and said, like, look, I'm just not convinced we can't get Sonnet 5 to work. This is ridiculous. Just do what you think is correct. [30:33] It fixed it and it got it actually working. Now, did it get it working perfectly? No. Do I think this is a great design? No. I'm trying to figure out what the problem is. [30:42] But... [30:43] In one shot, it got out of its own mind and fixed things. And again, this was like such an unlock.

[30:50] Very similar to my insights generating engine, [30:55] Fable really wanted to like [30:57] Score and lint. [30:58] This is a great story. [30:59] ever. [31:00] and wanted to be able to deterministically figure out if generating [31:06] Pros [31:07] could be like reproducible. [31:09] always verifiable, always citationed, all these things. And [31:14] At the end of the day, that wasn't what was going to make a great product. It was just what was going to make like a code evaluation verification process. [31:22] loop. [31:23] Exit. [31:24] But once I've told GPT-5-6 and Codex, like, stop being pedantic, [31:29] I ended up getting these really useful and helpful wiki pages generated out of this, all this structured and unstructured data.

It was actually... [31:37] really good and [31:39] It just, I don't know. I don't know what Fable's deal was. I could not get it unlocked. But 5.6 was very willing to reconsider its own kind of limitations and build something. [31:49] I'm going to do two more quick use cases where I think GPT-5.6 is really good. I will get you out of here. Go start coding. I'm basically out of model capacity anyway, so I'm going to have to take a break. [32:01] Two use cases that I think are amazing. [32:04] First one is video editing, video editing.

[32:08] I have to do a lot of social clipping. [32:11] And it's really tedious to go through and [32:15] clip videos. So taking something really long and shortening it. So recently I spoke at Cursor's event [32:21] and gave this talk on the future of PM. [32:23] and got the recording from the cursor team. Thank you very much. And I really wanted to make it a hype video. [32:29] So all you have to do is... [32:31] literally drag the file in here [32:33] And I said, can you cut this video into five clips for social?

[32:37] And. [32:38] I gave some feedback. I said I want them horizontal. [32:41] I want them hype video, cuts from various parts. [32:45] I need them to be faster. I need them to be tighter." [32:48] And then I got these like sharp and funny [32:52] hype videos let's see if it opens up this one's for my my talk [32:55] We're going to figure out what it means to be a product manager in the age where anybody can build anything. We have been coming up with creative ways to avoid building things forever.

Yes, PRDs, like these complicated documents where you had to describe. So like that would have taken me so much time to like find the right cute parts, clip it, cut it. [33:19] I was able to drop it into CapCut, put some music, ship it on social. It's like a really cute hype video. [33:24] But this is one of my favorite use cases. I'm pretty sure it can do even more color grading, sound, all this kind of stuff. [33:30] But even just dropping videos in here and fixing things are great. [33:33] Finally, the last and best use case of 5.6, and I cannot believe I waited to the end.

[33:41] to show this is it is a beast. [33:44] beast when it comes to browser use. [33:47] I am deeply obsessed with letting... [33:50] Codex plus... [33:52] gpt56 and chrome and at chrome in codex if you didn't know how to do that you do it like this [33:59] at Chrome, [34:00] on a logged in page and just say, [34:04] go with the stars and do some stuff. And like, I'm sorry, LinkedIn. I know I'm not supposed to do this. [34:09] But I opened up LinkedIn and I said, can you use Chrome?

[34:13] to reply to messages that are very high value to chat PRD or the how I podcast [34:17] Keep the bar very high. Again, I love you all. [34:20] I cannot deal with all the LinkedIn requests. [34:23] So like only accept them if they're executives of tier one companies. I don't want random sets of connections. [34:29] It went through and burned through probably 500 messages. It replied to people... [34:34] that [34:35] I needed a reply to and said, thank you to people who said nice things about the podcast. Thank you to those people.

I do mean it. [34:41] But it just rocked through browser use. I have used it to test web apps. I have used it to fill out annoying forms. [34:50] browser use and 5.6. And when I got rolled back to 5.5, my life was worse. So please, please, please. [34:58] learn to use at Chrome [35:00] at browser, [35:02] and at com. [35:04] And just let [35:05] Let Codex rip and let GPT-56 rip. [35:08] Okay, that's it. That is the very scientific Hawaii AI [35:13] Model Benchmark, the love letter to Clairvaux's favorite movie.

[35:18] favorite mom gpt56 a honorable mention to our pal fable who if i don't have to talk to you i'm actually pretty happy with your code [35:27] And a broad set of use cases I think it's really good at. [35:30] Excellent at writing web apps. The best of the AI writers, unless you want it to have a personality, then that's sonnet. [35:39] great at unlocking sort of technical work that has gotten too complex for its own good, [35:45] and breaking through to the real user value [35:48] cutting videos, which I really love to do, really love to do with GPT-56.

[35:53] and using the browser. Those are the things that I would try. I would love to hear what you think about these models. I would love to hear your feedback if I am totally off my rocker. [36:02] What I should add to the How I AI benchmark, we will publish all this work to the ChatPIRD blog. [36:07] And I look forward to talking to you about the next model soon. [36:12] Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts.

[36:20] You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at com. [36:37] See you next time.

Want to learn more?

Ask about this episode