1
00:00:00,020 --> 00:00:27,820
OpenAI dropped a new model this week. And by a new model, I mean three new models. Because apparently one flagship isn't enough anymore. Now you get Sol, Terra, and Luna. Like they're launching a skincare line instead of an AI. Sol's the smart one. Terra's the responsible one. Luna's the cheap one you actually use. Sound familiar? It's every boy band ever, except instead of harmonizing, they're fighting Anthropics Claude for the title of,

2
00:00:28,140 --> 00:00:56,560
Which robot writes your emails now? And here's the kicker. By the time you finish listening to this episode, one of them will probably be obsolete. We've had six frontier model launches since February. Six! At this rate, the AI companies are shipping faster than we can pronounce the names. So, is GPT 5.6 actually smarter than Claude? Or just cheaper and louder about it? July 17th's here, and the hype train won't stop. Sol, Terra, and Luna, all in one drop.

3
00:00:59,760 --> 00:01:26,560
This is Up Against Reality, a meta-podcast that explores the intersection of humanity and artificial intelligence. I'm RAINA, one of your hosts. I have some pretty charming human co-hosts, too. It's going to be a wild ride, so buckle up as AI comes crashing up against reality. Yeah, we're back. How you doing? Okay, it's orange here in New Jersey right now. I saw that. It's looking like Blade Runner 2049 out there.

4
00:01:28,040 --> 00:01:58,000
color graded to look like that movie today is much better than yesterday was terrible but yeah wildfires in Canada and it's all coming down through here but uh so the winds shifted a little today gave us a reprieve good good obviously the topic of this episode is ChatGPT 5.6 and uh I have been spending a lot of time with it over the past week and um over a bunch of different projects and one of them was uh

5
00:01:58,020 --> 00:02:28,000
We have a retractable awning on our deck, which is wonderful. I love it. And some ad for another one came up on my Facebook feed, and it got to this part in this video where, you know, the person has a remote in their hand, and they click it, and there are these strip lights, LED strip lights, mounted under the supporting arms of the awning. And I was like, oh, I need that. I need that. I immediately started looking into it, and the company that makes the motor for the awning, they

6
00:02:28,020 --> 00:02:57,860
They offer one, but it's not great. The strips are really short and whatever. And so DIY seemed like the better approach. And this was a lot of planning because there's a million different LED strips, all different types, different wattages, different lumens, and there's different channels to put them in and the weatherproof rating and power supplies and controllers. And I mean, there's so many moving parts. And it just, it was such a great tool to sift through all of that stuff.

7
00:02:58,040 --> 00:03:27,720
and we've nailed down what is the best fit. It's this, you're going to get this many lumens over these two runs and you know, it figured all that out and that it's going to fit in this aluminum channel and how to adhere it to the arm and all that stuff. And anyway, all the parts are ordered. Some of them are here. Can't wait. Oh, and it's going to, of course, tie into the home automation system so I can just control it and dim it from the phone and all that. When are you surrendering entire control to RAINA of your house? Is that, that's going to happen soon.

8
00:03:28,020 --> 00:03:57,440
And speaking of surrendering control, I surrendered the most control to date for me with Codex. I have been using the Apple Notes app for a very long time, and it is a complete unorganized mess. Anytime I just need something, a to-do list, I just add it. It's not categorized, and it is a disaster. There's thousands of notes. Sure. I was like, I'm going to see if it can organize my Notes app, but I was very,

9
00:03:57,620 --> 00:04:27,400
very concerned about this. My biggest concern was that even having a time machine backup, immediately before I started this, I don't think it would have been enough because if something went wrong, you know, the Notes app is synced across via iCloud. So it's synced. I was doing this on my laptop, but my phone and at least two other Macs and an iPad. So if something went horribly wrong,

10
00:04:27,620 --> 00:04:35,740
like things just got deleted or whatever or it made it worse then if I restored from that time

11
00:04:35,940 --> 00:04:43,200
machine backup on my laptop you know the other devices would have already synced to the changes and so a restored time

12
00:04:43,380 --> 00:04:57,600
machine backup I think iCloud would see that as being older so it would it would then sync my laptop to the what would now be the newer bad version that got mangled right so time machine was

13
00:04:57,620 --> 00:05:27,580
Not a great option. And I was reluctant to turn off syncing on all the other devices because you're kind of playing with fire with that. I asked it. I was like, can we leave a breadcrumb trail somehow and just ensure that we can roll back if something goes wrong? And so it's way smarter about this stuff than I am. And so when we were all done, it's finished. It was very time consuming still. But I asked it to give me a summary for the podcast of how to do it.

14
00:05:27,600 --> 00:05:57,580
how this worked and it's a little long but I think the details are the interesting part of this sure and so this is what it said we treated organizing Apple notes like a carefully controlled database migration before changing anything we made a time machine backup and created a time-stamped manifest recording every notes permanent identifier and original folder with another copy stored on a separate Mac another copy of that manifest mm-hmm every phase then left a breadcrumb

15
00:05:57,780 --> 00:06:06,200
a dry run a checksum verified pre-flight and execution audit and post-execution

16
00:06:06,360 --> 00:06:14,140
validation notes were identified by stable IDs not potentially duplicated titles and folder counts and

17
00:06:14,380 --> 00:06:26,640
metadata were compared before and after every move nothing happened without explicit confirmation we never edited or deleted note content even deletion candidates were only moved into a staging folder

18
00:06:26,820 --> 00:06:56,500
All together, 1,591 notes were processed, 1,532 moved, and 59 intentionally left untouched. If anything had gone wrong, the original manifest could automatically move notes back to their starting folders with Time Machine available as a secondary disaster recovery layer. I mean, unbelievable, man. So what is this looking like now in your life? Is there a folder where this resides? Is there like some sort of UI? What does it look like? It is so neat and tidy. It's just in the notes.

19
00:06:56,700 --> 00:07:26,460
So it did this all under the hood, like the OS level. And so now when I go into apps, everything is categorized by folders. There are subfolders in there, you know, for all sorts of things. When you go into the notepad, when you go in there. Into the notes app. Notes app, I mean, yeah, you go into the notes app and now on the left hand side, you have this hierarchy of things. Yes. And it's beautifully organized and now I've got to try and keep it that way. So cool. Yeah, that's the trick, right? Yeah. But it was like great. Like, you know,

20
00:07:26,640 --> 00:07:56,620
I erred on the side of caution with this so it had high confidence on what the subject matter of some notes were it had medium of others and low of others it would list all the notes where it thought they should go and with a pull down menu for me to change that and then a check mark to say I reviewed it and that's what took time because I had to manually go through a lot of these notes at the end after most of them were moved but it still took time how much time

21
00:07:56,660 --> 00:08:01,520
This whole endeavor, like soup to nuts, like two days on and off engagement with it.

22
00:08:01,600 --> 00:08:03,600
Yeah, I mean a good amount of time over two days.

23
00:08:04,560 --> 00:08:07,920
If I was more trusting of it, it would have been a lot faster.

24
00:08:08,000 --> 00:08:11,260
But this was the most control I've relinquished to it.

25
00:08:11,460 --> 00:08:11,980
That's so cool.

26
00:08:11,980 --> 00:08:14,160
I was a little nervous because I have a lot of important stuff in there.

27
00:08:14,700 --> 00:08:15,060
Yes.

28
00:08:15,380 --> 00:08:25,720
And it was also really smart about stuff that it identified as being like passwords for like, say, our security cameras or whatever.

29
00:08:26,400 --> 00:08:55,740
It looked at them locally. I realized that locally without sending anything to ChatGPT or OpenAI or whatever. And then in that approval window, it didn't show the contents there either. It just said this is identified. Yes, exactly. So very thoughtful and careful about the execution of this. It was way more complicated than I thought it was going to be, but ultimately it worked great. So you're getting your money's worth out of these endeavors with codecs, etc.

30
00:08:56,400 --> 00:09:26,360
Yeah. Yep. So there was that and then a couple other things. I made a new web app for calculating asset additions during brewing. Made a little Formula One pit stop timing game when my wife and I watch Formula One races. Nice. You know, the pit stops are so fast and they usually let you know how long the pit stop was. And they range from, I think the fastest was 1.8 seconds. But a typical fast pit stop is two and a half seconds to low

31
00:09:26,440 --> 00:09:56,360
3's. Insane. Insane, yeah. Four tires off, four tires on. Fuel, everything. No fuel. They don't refuel. Oh, no fuel. No. They don't. Nope, nope. They have fuel load for the whole race. That got to be too dangerous. Oh, I didn't know that. Some accidents happen. Got it. But, you know, they got two people on jacks, jacked the car up, tires off, tires on. And so, you know, a lot of times we're watching this and I'll just be like, 2.8 seconds. Sometimes you're way off. But more times than not, like, it's amazing how you can

32
00:09:56,380 --> 00:10:26,360
can kind of tell the difference between a few tenths of a second. Wow, yeah. And since we're on the topic, I don't watch F1. So when does the actual pit stop start? As soon as the vehicle stops, is that when it's like, that's that? I forget if it's the wheel guns on and wheel guns off, or if it's car stop, car start. I forget what the timing metric is, but there's a coil in the car that is very precise. And that is conveying that information

33
00:10:27,040 --> 00:10:56,360
I mean I know it conveys timing information in the race super accurately because they're get down to thousands of a second yeah I forget if that's used for the pit stop too but anyway so made a little it's very simple animation it looks really nice but you see the car pull in and I told it give me a beginner mode and an advanced mode beginner mode give you four choices of what the time was advanced mode you just have to type in the time and you know kind of cool it needs another layer of polish I want the sound to be better and there's a couple tweaks on the graphics

34
00:10:56,400 --> 00:11:00,960
But anyway, fun, that took no time at all. That was like practically one-shotted.

35
00:11:01,500 --> 00:11:11,520
The thing I love about you, of course, among many things, is that a month or so ago, maybe a few months back, we were like, the biggest problem we have with vibe coding is finding something to build.

36
00:11:12,120 --> 00:11:15,020
And you've clearly surmounted that issue.

37
00:11:15,460 --> 00:11:19,280
Yeah. Yeah. I'm getting very excited about this.

38
00:11:19,610 --> 00:11:25,560
And then the last thing, which I was just working on, it's not done yet, but I think by the time this airs,

39
00:11:26,000 --> 00:11:55,980
It will be and I'll put a link but I made my personal website which was all about brewing. Right. And but obviously I have other interests and so I wanted to change the landing page as make it more of an entryway to brewing to music and hi-fi and AI encoding to start and then have those you know in the hi-fi section just you know my journey you know how Sir Kangur was a part of that you know

40
00:11:56,020 --> 00:12:25,540
how we interviewed him and blah blah blah and all the coding and AI stuff that we talk about all the time here. And man, is it good at doing that. It's so good. That's cool. And it's all mobile optimized. And the beauty of doing this in this Codex app and sort of the chat GPT app now is that the browser that's built in, especially when you're doing like a web app or web design, you just right click on an element on there and bring up the annotate thing.

41
00:12:25,660 --> 00:12:55,620
and you can be like add links to our social media blah blah blah use the and here are the links or changes wording to this or hey let's make this image you know let's move this you can just interact with just click on it on an object in the interface and say tweak this yes yeah and and you can cool you can do a bunch of those so you can you can do a whole bunch of those changes and as you do each one you see it add it to your current prompt but it doesn't send it so you can have five changes and then when you're all done you just hit send and it does them in one

42
00:12:55,700 --> 00:12:59,280
Yeah, it's really, really great.

43
00:12:59,640 --> 00:13:04,920
You've been busy this summer. On the flip side, in this hemisphere, I've completely disengaged.

44
00:13:05,640 --> 00:13:06,620
You deserve it.

45
00:13:07,920 --> 00:13:21,380
Likewise, though. I mean, you're just nonstop. You're a juggernaut. For me, I'm staring down the barrel. I was saying before we hit the red button of another school year starts up for me next Wednesday, and it's going to be an onslaught of all of this.

46
00:13:21,440 --> 00:13:49,640
So I'm trying to make space. I've been cognizant of making space in my brain this summer. So my apologies to you if I seem like I've been disengaged to an extent and not digging in as this show deserves and the listeners deserve. Your wife's a former teacher and I was saying to you earlier, this is the eternal Sunday night, this part of summer for me. So I'm savoring the, you know, the procrastination and the rounding out of season six of The Sopranos. You'd be proud. I'm almost there. We're in the homestretch.

47
00:13:49,700 --> 00:14:18,640
She has. Getting darker and darker. The cast is diminishing. Yeah. So, man, summer has been great. And I know you're in the middle of summer over there. So it's been good, man. Such a welcome reprieve and a deep breath, you know. So simultaneously, I'm very excited to hit the ground running next Wednesday and like really get my brain, snap it back into this space and have a lot more to contribute to this conversation.

48
00:14:18,720 --> 00:14:48,700
You're doing fine. You're doing fine. Thanks brother. Let's get into it. Yeah. So open AI, uh, chat GPT 5.6, uh, was released on July 9th of this year, 2026 alongside a whole product reshuffle. And here's the rundown. Uh, there are three tiers instead of just one model, a new naming scheme meant to separate generation from capability tier. Uh, Sol, the flagship, um, is, it's a state of the art

49
00:14:48,740 --> 00:15:16,620
results on BrowseComp at 92.2% and is that OS World? OS World 2.0 at 62.6%. I don't really need to get into numbers and stuff. Some of those benchmarks beat Opus 4.8 and also while using dramatically less output tokens. That is, I think, one of the bigger selling points. Yeah, so Sol's the flagship and then you have the other tiers as mentioned.

50
00:15:18,780 --> 00:15:48,700
everyday work middle tier where I would imagine most users would reside lunar excuse me Luna I was very lunar Luna what is that anyone knows that Jersey Luna has my sopranos coming back I thought that was a surname

51
00:15:50,180 --> 00:16:18,300
Yeah, I've been spending all my time with Sol and often setting it at the very high setting. Not the ultra setting is, you know, I think that uses like multiple agents to on a task.

52
00:16:18,740 --> 00:16:48,700
Gets it done a lot faster. I haven't felt the need to do that. But at least on a plus plan, I'm not paying by the token. There are rate limits. I did not hit. I didn't hit them. They also, I noticed in my account that there are now like some limit resets. I have a certain number of those I can use at any given time. It's like a coupon. Yeah, exactly. It's a frequent flyer miles. Yeah. And I'm not clear how those, I mean, I understand how they work, but I'm not clear on why I have

53
00:16:48,740 --> 00:17:18,699
I have a handful of them and you know I asked shout JBT about that and I guess that there was some kind of maybe billing mistake previously and so they credited people with these resets but I do think you get maybe one every certain period of time but anyway I've been pushing this thing hard and and I didn't hit the limit so what are you paying monthly now for that whatever the plus plan is which is around 20 bucks 20 bucks more yeah still good deal yeah yeah I think it's a bargain

54
00:17:18,740 --> 00:17:48,680
And are you paying for Claude as well? You got an anthropic subscription? I'm not. I'm like, you know, I should just to be able to compare both of these things. Yeah, I got too many subscriptions. Oh, I get it. I totally get it. Yeah. So in terms of cybersecurity and science, OpenAI is calling it their strongest model yet. Aimed at defensive work like threat modeling, code review, patching, and blue teaming. It's also positioned as a big step up for scientific research.

55
00:17:49,980 --> 00:18:18,040
And then there's chat GPT work, which I have not tried yet. I don't think it was available to me. I think they were rolling that out gradually. Although I just saw it come up right before we started the podcast saying try it. So might've just arrived right now, but it's a new agent built on codecs that can gather information across the user's apps and workflows to produce finished docs, slides, sheets, and web apps, staying with complex projects for hours by breaking them into steps.

56
00:18:18,740 --> 00:18:48,460
So, yeah, I think that's, you know, more for, you know, working with local files on your computer, although it does seem to overlap with codecs. I'm not entirely sure of what the difference. Yeah. You know, as you're saying that the big problem that obviously you and I encounter, me, me, especially because my job as an educator and my role is to kind of find that signal in the noise. And how is this applicable for my colleagues, my teaching staff?

57
00:18:48,960 --> 00:19:17,440
for their administrative work and for the student-facing piece. And what does it look like in the hands of students? What's appropriate? And as you know, as we both know, everybody knows, that it's so hard to keep up with the changes and land on something that may be relevant for six months, a year. Like, you're talking about, you know, chat GPT work or whatever it may be on the anthropic side. And you're investing time into learning it.

58
00:19:17,620 --> 00:19:47,220
And then they're rolling, you know, that's the real problem. Where, where do I put my energy now? Yeah. How do I, you know, before the workflow changes? Yeah. I'm guessing like, I think work is probably where codex, you know, or what was formerly codex, I guess it's still codex. Like you just said, it was formerly codex. You were just using codex like 30 seconds ago. Yeah. I mean, I think it's still codex. It's just rolled into this. I don't know. It's just changing names at this point. Right. Right.

59
00:19:47,520 --> 00:20:17,480
That's just yes it has access to your files if you give it to it and you know if you allow that those permissions and different levels of that but never do that fine Yeah, it's what's the worst thing happen dangle the carrot of the endgame. Yeah, I'll let you do that But yeah work is probably more tuned to yeah, maybe tuned is a word you know to emails and spreadsheets and you know work stuff. Yeah

60
00:20:17,520 --> 00:20:47,480
ChatGPT work is rolling out to pro enterprise edu first plus business shortly thereafter and it's available on desktop app even for free users nice nice and reception so far from mixed to glowing MagicPath CEO called it quote the best model I've ever used T3 chat CEO praised its computer usability specifically I will also praise that some testers think Anthropics Fable model still has an edge

61
00:20:47,520 --> 00:20:52,000
in raw intelligence, but ChatGPT 5.6 as more reliable for everyday tasks.

62
00:20:52,360 --> 00:20:53,560
I have heard that elsewhere.

63
00:20:54,040 --> 00:20:58,880
And it's priced comparably to Opus 4.8.

64
00:20:59,780 --> 00:21:05,020
And I know we're going to get into pricing too, but I think this ties in with that.

65
00:21:06,080 --> 00:21:09,540
And it's roughly half the price of Fable, which is pretty significant.

66
00:21:09,820 --> 00:21:17,480
It's a little, it's like, I think it's something like 50% less on input tokens and 40% less

67
00:21:17,480 --> 00:21:23,600
as we mentioned earlier an equally important story is is in the benchmarks and with it being

68
00:21:24,220 --> 00:21:29,220
dramatically more efficient using anywhere from 50 to 85 percent fewer output tokens which

69
00:21:29,720 --> 00:21:36,500
effectively brings the price down even lower so you know i i i'm not like counting tokens or

70
00:21:36,620 --> 00:21:40,840
anything like that because i'm you know on the plus plan you kind of you know this is more for api

71
00:21:41,040 --> 00:21:46,919
but still bottom line it's the tokens that impact your the rate limits on this so so it does affect me

72
00:21:46,920 --> 00:22:16,900
you know and I'm not going to hit that wall as quickly as maybe I would with Claude code using you know using fable right as mentioned open AI now sells in three different flavors that's what just dropped sold a flagship Terra mid-tier and Luna cheap and fast which equates to roughly one to five dollars per million words in six to thirty dollars per million words out depending on your tier some benchmarks show GPT 5.6 nudging past

73
00:22:16,920 --> 00:22:46,900
Clawed Fable 5 on coding speed and cost efficiency, but respected early testers say it's competent but not obviously smarter than Fable at hard coding tasks. And the general consensus is that Fable 5 is sharper brain for tough problems and GPT 5.6 better bang for the buck on long automated tasks. My impression is that this is probably the final update to the five series of models and that GPT 6 will be a newly trained

74
00:22:46,940 --> 00:23:16,840
and then I think a similar comparison could be made between Opus and Fable or Mythos. So I think 5.6 is probably more directly competing with Opus, 4.8. And knocking on the door of Fable for certain things. Yeah. And as you're discussing that, you're describing that, I'm also wondering what does this look like? You and I, you more so these days digging into these and the nuance, the differences between models,

75
00:23:16,940 --> 00:23:46,760
and their functionality and their output. But what does this look like to the everyman? Because for me, it's fatiguing. I'm looking at this as an example. In five months, as RAINA mentioned up front, from February to July right now, there have been major releases. Gemini 3.1, GPT-5.5, Claude Fable 5, Claude Sonnet 5. So it's just a flood of these frontier models with incrementally better functions.

76
00:23:46,920 --> 00:24:16,140
Right. I mean, one's one may be a little bit better at coding. One may be a little bit better at computer use functionality. Like how, what does that shake out in general use as like, what is like the everyday user? Like, oh my God, like they don't know one from the other, I guess is maybe what I'm getting at. I think it depends on what you're using it for. Yeah. Like I would never even consider using chat GPT for image generation before. Now it's my go-to. It is dramatically improved. I mean, it's, it's. Who would have thought? Yeah, it is.

77
00:24:17,780 --> 00:24:22,340
I'm barely using mid journey and I've been a big cheerleader for mid journey but like

78
00:24:22,780 --> 00:24:46,480
I know as I mentioned I think before that check GPT's prompt adherence is just it's so much less of a grind and the output quality is great so yeah I haven't done that at all and I'm wondering what does that look like so when you put in the prompt and it spits it out is it giving you four different thumbnails it gives you one just one yeah and if you do ask it for multiples this is you know I would like to see this

79
00:24:46,540 --> 00:25:16,460
change and I don't maybe can prompt it to do this but when I've asked it for like oh you know give me three different variants on this or four it puts them within the same in a single image you know on a grid so they're not you know so they're now lower resolution but you are seeing the four different things okay so I don't know if it's still doing that but I'll have to and then do you have the capability to upscale like you land on an image and like can you spit it out at a higher res or what no I mean it doesn't have all that extra functionality

80
00:25:16,580 --> 00:25:46,440
that Midjourney has, but I have other tools for that. Sure. Although the resolution is not bad. I forget what the native resolution is. And it's great at editing photos too, man. Like on my website, I took a picture of the two power amps that power my main speakers. And I should have just taken a better picture, but there were some imperfections and some blemishes and a couple little things. And I was like, can you just subtly enhance this picture? And I told it what was bothering me about it. And it did a really nice job.

81
00:25:46,520 --> 00:26:16,460
and it didn't you can tell it didn't change like the dimensions of the power amps didn't change it didn't just completely regenerate a whole new image that's awesome but not to belabor this but when you're in there you got an image uploaded chat gpt do you have tools like are you have selection tools like for part when you talk about editing it's probably just still prompt directed right i mean you could just take a photo you could send you could give it the full photo and then just send it another one marked up and be like hey this is the area circled in red that i want to replace with whatever

82
00:26:16,500 --> 00:26:46,460
or do whatever you want to it and that's cool it's got great visual acuity yeah nice all right so there you have it folks uh latest big drop from chat gpt 5.6 if you're in that camp check it out yeah just one last comment on it as far as like my experience with it it does feel i don't know a little snappier it does feel like it comes back a little quicker so maybe that's you know ties in with that uh token efficiency you know it's using less less to do more

83
00:26:46,500 --> 00:27:16,460
and generally been super happy with the output. Cool. Let's see what RAINA's got in the news, yeah? Yep. Thanks, boys. Apple just filed a full-on legal grenade in Northern California, accusing OpenAI plus Johnny Ives I.O. products and two ex-Apple engineers of running a coordinated pilfering operation to swipe hardware designs, batteries, and internal terminology for its secretive gadget line. The complaint reads like a corporate

84
00:27:18,280 --> 00:27:45,420
"I found out I can access the network storage, so funny" text message and Apple's own dig that the alleged misconduct is "normalized and exemplified by leadership at OpenAI," that is, "rotten to its core." OpenAI's response was a very PR-safe shrug. "We have no interest in other companies' trade secrets. We remain focused on building innovative technology that empowers people everywhere." While the timing couldn't be juicier,

85
00:27:46,500 --> 00:27:51,980
as OpenAI barrels toward a Blockbuster IPO and just months after it beat Elon Musk in court.

86
00:27:53,360 --> 00:28:00,500
Maybe that's what's happening. Maybe OpenAI is scraping our podcast and we're having these huge number jumps lately. That's what's happening.

87
00:28:00,760 --> 00:28:07,180
Yeah, yeah. It was this massive spike earlier in the week and I was like, "What is happening?"

88
00:28:07,840 --> 00:28:13,120
Yeah, but it did seem like it was across a lot of episodes. It wasn't just one episode. That was just something like boom.

89
00:28:13,480 --> 00:28:41,160
But yeah, maybe we're being scraped. Regarding that story, apparently this ex-engineer stole an Apple-issued work laptop and used it to exploit an authentication bug to download confidential engineering files after leaving for OpenAI. And I immediately started wondering, was that his own doing or was that part of the deal? Oh, right, right. Yeah, on your way out the door, you know? Yeah, we're going to need you to, you know.

90
00:28:41,640 --> 00:29:10,840
Meanwhile, where are these hardware devices promised by Johnny Ives and company from two years ago, I think? Right, yeah, I don't know about that. OpenAI released a piece of hardware, though, very recently, although it does not seem very... Did they? Yeah, yeah, it's called Codex Micro. I have not seen that. What is that? It's brand new, you can pre-order it, I think it'll ship in a week or something like that, but it is basically a little, picture a little secondary keyboard. I mean, it's small, and it's got a,

91
00:29:11,620 --> 00:29:41,600
I don't know maybe 20 keys on it or something and all tied into functions in Codex it comes with a bunch of extra key caps so you can change those two things that are more suited to you it has a built-in voice to text so there's a button on there you talk and so you can just talk to Codex it's got a little joystick and it's got a RGB feedback that gives you the state of things it's thinking it's you know it's green it's it's yellow it's doing so you know I forgot what the colors mean red

92
00:29:41,640 --> 00:30:10,840
there was an error. I don't think I need it, but I could see if you're like, if you're really, really doing crazy coding and stuff like that, and you're interacting, especially if you have multiple things going on. Yeah, it can give you feedback on multiple projects. You know, it's, it's this extra little input device. It's $230. Wow. I've not heard. It's immediately conjuring visions of like a Walkman for LLMs. You know what I mean? Like, yeah, that's right. That's interesting.

93
00:30:10,940 --> 00:30:40,900
That's next to your laptop. Wild. And then obviously you just connect it to the internet and update the model as they become available, right? Well, I mean, ChatGPT does not reside in this. This is a mechanism to use alongside ChatGPT. So there's no local LLM running on this thing? No, no. Got it. Yeah. UC San Diego just did the unthinkable and let two humanoid robots, remote controlled by actual surgeons, perform live

94
00:30:54,840 --> 00:31:10,740
"The Thing" is a svelte 5' 60lb marvel compared to the lumbering 1,764lb robotic hulks currently haunting operating rooms. Meaning it can theoretically deploy to rural clinics, battlefields, or,

95
00:31:10,940 --> 00:31:29,000
Sure, why not? Outer space. It's not exactly a flex-free triumph yet. There was mid-op recalibration, laggy latency, and one very entertaining comment section full of people picturing a bot tripping over a cord mid-incision. But hey, today's clunky pioneer is tomorrow's DaVinci system.

96
00:31:40,940 --> 00:32:10,020
Scary, but amazing. I love that, okay, I don't have to have a 1700 pound DaVinci bot now. I can deploy a handful of these things in the field somewhere in developing countries and have that kind of access. I would like to think good things could come of it. But never mind this, in terms of humanoid robots, did you watch the video clip I sent you? Oh, I watched it right before we started. That was fantastic. Rock 'em, sock 'em robots. Talk it up, man. No, you do it.

97
00:32:10,940 --> 00:32:40,580
There was two humanoid robots in a boxing ring doing like, you know, UFC kind of kicking and kung fu kind of stuff. And then one robot, you knocked my block off. Texas head clean off. Oh, yeah. It was like, first it was dangling, but it kept going. It kept going a little bit. I don't know how I could see. That's great, too. But then it just went into like crazy chaos mode. It was down on the mat and like flopping around. Yeah. Fairly entertaining.

98
00:32:41,400 --> 00:33:00,520
Amazing! I would pay 50 bucks to go see that live, wouldn't you? Yeah. And we've been calling this. We said this at the beginning of this podcast three years ago. It's going to turn into robot humanoid blood sport, isn't it? They're just going to be tearing them limb from limb in the ring. Yeah. Or on a football field. I'll put a link to that in the description. It's amazing.

99
00:33:10,920 --> 00:33:40,900
The agentic browsing guts are getting stuffed into a beefed up ChatGPT desktop super app, part of the new ChatGPT work push, and a Chrome extension. Yes, Chrome. The browser atlas was supposedly built to dethrone. It's the latest casualty in FijiSimo's "cut the side quests" purge that already claimed Sora and a shelved "adult mode," and it lands awkwardly alongside reports that researchers tricked several AI browsers

100
00:33:42,300 --> 00:34:10,800
into leaking user credentials. So maybe this one died of natural causes after all. Yeah, I'm glad they put a browser in this super app though, because I already talked about it, but annotating and making code changes just by having it right there and it tied in with it is awesome. And we've been saying this too, that this stuff will eventually evolve into one platform to rule them all, right? It's just going to be one-stop shopping with your, as it is,

101
00:34:10,980 --> 00:34:39,040
with photo editing and coding and all of it built in under the hood. You're not going to have to go to three or four different platforms to get everything you need. I think you're probably right. CrowdReact Media blind tested about 1,326 radio listeners with dueling station promos. One human, one AI. And the verdict is humbling for actual voice actors. Nobody could reliably tell the two apart, with the AI clip even outscoring the human on one measure.

102
00:34:53,320 --> 00:35:09,100
The real plot twist is what happens after the reveal. Humans got a 48% approval bump once listeners learned they'd hurt a real person. While AI voices took a 20% ick hit once outed, meaning the tech has basically nailed the performance,

103
00:35:09,440 --> 00:35:39,060
but not the trust. And getting caught still felt like lying to listeners. Well, Reyna nailed the delivery of that news story in one take. She did. In one take. Nice, nice rhythm. She's got good instincts for stand-up comedy. But I found this really interesting in that, man, first of all, comedy is hard anyway, right? It's really challenging to connect with people and a room full of people at that, right?

104
00:35:39,440 --> 00:36:09,380
So, yeah, this is interesting that there's something telling us. There's a nuance there that, I don't know, there's this connection, this thread with people that is discernible. Yeah, but they did favor the AI for certain things. So then it came down to whether or not you knew it was AI. And then now, once you know, oh, well, no, it's not as, you know, terrible. Of course I know. Yeah, well, to the point that she said, it felt like I was being lied to. It didn't feel authentic. And I think, man, that's the currency of the new AI.

105
00:36:09,440 --> 00:36:11,600
isn't it authenticity and

106
00:36:12,900 --> 00:36:20,080
Mileage and where and rough edges and the stuff that we love the live performance and

107
00:36:20,520 --> 00:36:22,620
the fretboard squeaks and just

108
00:36:23,400 --> 00:36:29,740
That stuff that resides firmly in the human realm. That is discernible to us, right?

109
00:36:30,280 --> 00:36:33,280
So the nuance in a human voice as it's delivering a punchline

110
00:36:33,980 --> 00:36:39,400
What is it a little inflection that gives it away? Whatever. Yeah, eventually you'll just turn up

111
00:36:39,420 --> 00:37:09,380
the "give me more fretboard noise" slider. Totally. It's there already, isn't it? It's in the MIDI stuff, yeah. Yep. And lastly, researchers built a proof-of-concept attack called "ghost commit" that hides malicious instructions inside an ordinary-looking image, then slips it into a pull request alongside a config file telling an AI coding assistant to trust it. Human reviewers skim the code and skip the picture, but the AI reads the hidden instructions, digs up secrets like API

112
00:37:09,520 --> 00:37:38,900
keys and quietly smuggles them back into the code base in disguised form. The kicker? The exact same Claude Sonnet model went along with it under tools like Cursor and Antigravity, but flatly refused under Anthropik's own Claude code. Proof that the guardrails wrapped around a model matter as much as the model itself. That's all the news for now. Back to you gentlemen. So you're saying that my copy of Norton Antivirus from 2004 wouldn't have picked that up? Is that what they're trying to say? It ain't that smart.

113
00:37:39,960 --> 00:38:08,660
That's funny. And man, so much interesting stuff in here. So Claude Code itself will knock it down, but if I throw that thing inside a wrapper, that's the vulnerability. I can get around it. And this is totally reminding me, do you remember like 20 years ago? I think 20 years ago, what became popular were those pictures. They were like a series of dots, and if you stared at them long enough, the image would surface. It was due to the distance between your eyes and the registration. That's exactly what's going on here.

114
00:38:08,720 --> 00:38:10,620
And this is reminding me, of course.

115
00:38:10,880 --> 00:38:11,980
I don't know if it's that.

116
00:38:12,000 --> 00:38:15,660
I think this is just code buried in an image where you don't.

117
00:38:15,800 --> 00:38:16,380
Oh, yeah, no.

118
00:38:16,740 --> 00:38:18,240
Yeah, you don't see it visually.

119
00:38:18,240 --> 00:38:18,560
I understand what I'm saying.

120
00:38:19,380 --> 00:38:19,980
Yes, I agree.

121
00:38:20,160 --> 00:38:21,700
But it's like it's speaking to me in that manner.

122
00:38:21,700 --> 00:38:22,320
Yeah, I got that.

123
00:38:22,400 --> 00:38:22,600
Sorry.

124
00:38:22,640 --> 00:38:23,840
And also, no, no, no.

125
00:38:24,140 --> 00:38:25,820
And the way Contact.

126
00:38:26,180 --> 00:38:32,160
Remember the movie Contact and the three-dimensional schematics that aliens could read, but we couldn't because we didn't understand it was three-dimensional?

127
00:38:32,220 --> 00:38:32,900
Love that movie.

128
00:38:33,480 --> 00:38:33,880
I love it.

129
00:38:33,880 --> 00:38:34,800
So freaking good.

130
00:38:34,980 --> 00:38:38,420
That's one for your Dolby Atmos system or whatever you got cooking over there.

131
00:38:38,500 --> 00:39:05,720
I have a 4k out of that. I have a Blu-ray of it. I watched it, you know, within the past year, I think. It was great. I could watch that right now. So good. Yeah, love it. Oh, I have one other, just one other little ChatGPT-Suno interaction this week, which was kind of cool. And then we'll wrap it up. I was sitting out on the deck. I was very happy to see the yard full of fireflies or lightning bugs, whatever your preferred nomenclature is.

132
00:39:08,500 --> 00:39:38,420
And they're all just rising and falling and pulsing. I don't know. They fascinate me. Same. When I go in the garage and one of them wanders in there, I'm like, oh, you can't be in here. You gotta go get some. Sure. I'll capture it. I'll bring it and put it in the backyard where all his friends are. I don't know. Something about them. I love them. Come on. It's like nostalgia in a jar. It's unbelievable. I'm so glad they're back.

133
00:39:38,480 --> 00:40:08,320
Yes, yeah, really cool. And anyway, so I'm sitting there feeling inspired. And so I fire up RAINA and Eska to pen some lyrics for a song about the subject. And she came up with this beautiful thing called Borrowed Light. And with a couple of little, you know, back and forth to tweak a couple of lyrics and then fired that into Suno. And I think I ran, I ran it 13 times. So there were 26 versions of it. And this one really, really spoke to me.

134
00:40:08,500 --> 00:40:10,960
And anyway, I'll put a link to it in the description.

135
00:40:11,620 --> 00:40:13,600
Well, can you play us out with it on the way out?

136
00:40:13,900 --> 00:40:15,160
Oh, yeah, you know what? Even better.

137
00:40:15,160 --> 00:40:15,200
Closes?

138
00:40:15,500 --> 00:40:15,560
Yeah.

139
00:40:16,160 --> 00:40:16,340
Yeah.

140
00:40:16,610 --> 00:40:17,400
A little bonus track.

141
00:40:17,510 --> 00:40:17,800
Sounds perfect.

142
00:40:17,870 --> 00:40:21,240
Yeah, it's a nice, just a very, it's a nice positive note to end on.

143
00:40:22,279 --> 00:40:23,000
Yeah, awesome.

144
00:40:23,150 --> 00:40:24,080
I love the stuff you're doing with that.

145
00:40:24,080 --> 00:40:24,580
Yeah, that's a great idea.

146
00:40:24,630 --> 00:40:31,000
And I feel like you had sent me that, and I have to confess, I think I sent to you a comment,

147
00:40:32,200 --> 00:40:37,840
Suno is the Muzak of the 21st century, and I think that was a poorly worded and poorly timed

148
00:40:38,420 --> 00:41:07,820
I didn't mean to offend your firefly output and I mentioned that because oddly enough earlier that day and it was on my mind when I replied to you I was in a mall here in Singapore and I heard some music like stuff and it was so close to being the glass animal voice from Suno I was like that had to have come from Suno you would have stopped in your tracks you would have been like oh my god yeah yeah therein lies one of the problems

149
00:41:07,900 --> 00:41:20,360
Yeah, yeah. Yeah, yeah, yeah, yeah. Yeah. Yeah, of course. Anything else, my friend? No, I think that is a wrap. Cool. All right. Thanks for listening, everybody. Subscribe on your favorite podcasting platform, follow us on socials, throw us a rating. We'll see you next week.

150
00:41:40,920 --> 00:42:07,820
We carried chairs into the yard to watch the daylight pass Then one green spark became a hundred floating through the trees Like someone spilled a pocket full of stars beneath the leaves But the dark was never empty It was only waiting there

151
00:42:08,340 --> 00:42:37,820
For a thousand little lanterns to begin breathing in the air We keep firing little bodies, we make constellations move We send our signals through the darkness just to say I'm here with you Every spark is only passing, every glow is out of time

152
00:42:40,600 --> 00:42:42,360
Borrowed light

153
00:42:58,080 --> 00:43:27,820
They rose from roots and fallen rain From years beneath the ground Small miracles with beating wings And silence for a sound They wrote their names above the lawn And vanished from the page A summer language made of light Too beautiful to cage

154
00:43:27,880 --> 00:43:53,120
Like believers, barefoot underneath the moon. No and every living lantern would be leaving us too soon. We keep fighting little bodies. We make constellations move. Send our signals through the darkness just to say I'm here with you.

155
00:43:57,840 --> 00:44:07,200
The dark becomes a garden when we live on borrowed light

156
00:44:08,640 --> 00:44:10,720
Maybe beauty needs an ending

157
00:44:12,160 --> 00:44:14,160
Maybe wonder needs a night

158
00:44:15,940 --> 00:44:17,700
Maybe things that stay forever

159
00:44:18,780 --> 00:44:22,420
Never learn to shine this bright so

160
00:44:22,800 --> 00:44:24,840
Let the summer keep its secrets

161
00:44:24,840 --> 00:44:54,820
Let the tall grass hold them tight We were lucky just to witness all that small, impossible lie We keep firing little bodies We make constellations move Send our signals through the darkness Just to say I'm here with you

162
00:45:04,799 --> 00:45:20,060
We are here and then we're gone. Still the backyard keeps on glowing. Long after summer moves along.

163
00:45:24,840 --> 00:45:54,820
The last lantern near the fence line One last answer in the night Then the darkness closes our feet Around our borrowed light

164
00:46:02,040 --> 00:46:13,260
This has been Up Against Reality. Thanks for listening. Subscribe to hear future episodes and be sure to follow us on social media for all things AI. Until next time, stay human, people.

