Security Now 1095 transcript
Please be advised that this transcript is AI-generated and may not be word-for-word. Time codes refer to the approximate times in the ad-free version of the show.
Leo Laporte [00:00:00]:
It's time for Security Now. Steve Gibson is here. It's Patch Tuesday, and you won't believe how many patches Microsoft just shipped. Steve says that's okay, they're fixing things. New versions of, uh, Google's Gemini, NVIDIA acquiring Hugging Face, Chinese cyber espionage, and then a lifelong safety engineer worries that AI is going to cause us a loss That and a whole lot more next on Security Now.
Steve Gibson [00:00:32]:
Podcasts you love. From people you trust. This is TWiT.
Leo Laporte [00:00:44]:
This is Security Now with Steve Gibson, episode 1095, recorded Tuesday, September 8th, 2026. AI-driven expertise loss. It's time for Security Now. Yes, Tuesday has come around once again. And that means Steve Gibson is knocking at the door, ready with a 22-page document of all the latest security problems in the world. Hello, Steve.
Steve Gibson [00:01:12]:
Yo, Leo.
Leo Laporte [00:01:14]:
It's good to see you.
Steve Gibson [00:01:16]:
For those who looked at the show notes, Benito was the first person to highlight the fact that I had numbered this 1096. This is not 1096. This is 1097. This is 10.95.
Leo Laporte [00:01:27]:
Fixing it right now. Yes.
Steve Gibson [00:01:29]:
So the links are right, but in the show notes, I had it wrong. So anyway, we are, this is, and I have not yet looked, we are on Patch Tuesday. And a big question that we have is, how does this one compare to the last one? Which of course was a whopper.
Leo Laporte [00:01:50]:
Your theory is it'll get better. That's the point. I believe—
Steve Gibson [00:01:54]:
I don't know what the shape of the curve is. I don't know how soon it's going to get better, but in theory, as long as AI is aiding us in eliminating more bugs than we or it is creating, we should be seeing a drop-off because we certainly are seeing that we're paying for a lot of the legacy code mistakes had been made in the, in the, you know, previous era. So, uh, we're going to start out talking about a classic old-school hack against Dropbox. Actually, something happened that was not about AI, believe it or not. Uh, we've got, as I said, the, uh, oh, uh, the Patch Tuesday after today is going to be enabling Something that Microsoft first— it sure, it first appeared in Windows 10, and we talked about it back then. Uh, the short name for it is Memory Integrity. Uh, many people turned it off because it dropped gamers' frame rates in the interest of improving their security. A lot of gamers said, I'm secure enough, I need a high frame rate.
Steve Gibson [00:03:08]:
Anyway, it's going to be turned on next month. for everybody that qualifies based on the hardware. So we'll talk about that. Firefox moved to 155 and obtained what I call a dumb smart window. So we'll touch on that. CISA has terminated 6 of its most valuable cybersecurity services. And really, you couldn't choose a worse time. for CISA to back away from, you know, infrastructure cybersecurity because everyone's expecting AI to impact that.
Steve Gibson [00:03:47]:
Well, everyone except me. I'm not convinced that that's going to happen. We'll talk about that. I want to touch on OpenAI's big announcement because last week, or since we talked last, lots of things happened. OpenAI has released GPT-6 Astra, which is you know, that model, we, we did discuss it before, that they're saying, uh, qualifies for what they consider critical security treatment. Um, and then also there's been a weird bunch of what I consider misreporting, um, angst generated by the, you know, the, the anti-AI guys saying that OpenAI is doing something which is hiding its, its own internal thinking, masking its chain of thought. I'm going to open that up, take a look at it, and explain why that is not what is going on. Also, uh, we'll, we'll touch on NVIDIA's acquisition of Hugging Face, uh, and who that's good for.
Steve Gibson [00:04:56]:
Google released Gemini, uh, 3.8 Flash, uh, you know, having waited, what, 2 weeks from, uh, uh, 3.7. So this is all happening very quickly. Um, Chinese cyber espionage is beginning to use AI to become more slippery. We'll talk about that. And Matthew Green, our favorite Johns Hopkins cryptographer, has a fascinating take on what it means for there to be an AI-driven bug drought Which, uh, it's— he has some very worthwhile thoughts. And then today's topic is AI-driven expertise loss. A lifelong engineer who has designed the control systems for nuclear reactors, uh, I think has some very important things to say about the consequence of AI being so good. and what it means in the long term.
Steve Gibson [00:06:00]:
So, uh, lots of fun things to talk about and another great picture of the week. So yeah, uh, this is episode 1095 despite what the show notes say.
Leo Laporte [00:06:10]:
We'll get to that next week. Do you want to know how many patches there are in the Patch Tuesday this week?
Steve Gibson [00:06:16]:
I do.
Leo Laporte [00:06:17]:
Are you curious? Well, Tom Warren on The Verge says, I understand Today's September Patch Tuesday is about to set a new record with more than 650 security fixes for Windows alone.
Steve Gibson [00:06:35]:
Whoa!
Leo Laporte [00:06:36]:
650. Last month was 400. We thought that was a lot.
Steve Gibson [00:06:41]:
It was a lot.
Leo Laporte [00:06:41]:
July was 570. We thought that was a lot.
Steve Gibson [00:06:44]:
It was a lot.
Leo Laporte [00:06:47]:
Yikes. So it's not going down.
Steve Gibson [00:06:51]:
Not yet.
Leo Laporte [00:06:52]:
And Microsoft says, yeah, we're using AI. We're finding them.
Steve Gibson [00:06:55]:
Wow.
Leo Laporte [00:06:56]:
So there you go.
Steve Gibson [00:06:57]:
Well, I mean, and so it's a mixed blessing, right? It means that we've all been riding on this just Swiss cheese, how does it even boot operating system? Wow. That's great.
Leo Laporte [00:07:11]:
Well. So it hasn't, I don't think Microsoft has actually released it yet, But that's what Tom Warren, who's very well connected and usually very reliable, says, that that's the number.
Steve Gibson [00:07:20]:
So a bunch of operating systems are going out of extended service life. Oh, no, no, it got extended. We have it until October of '27.
Leo Laporte [00:07:32]:
Another year.
Steve Gibson [00:07:33]:
And that's really good because by then they should have Windows 10 finally bug-free because, you know, we are getting security patches, backported to Windows 10. And at the same time, they're no longer writing any new stuff to screw it up. So it's perfect. This is exactly what you want. You want another year of fixes without them adding anything new that's going to destabilize it. And then when it finally goes out of service, if they actually do take it out of service next October of '27, it'll be a solid operating system.
Leo Laporte [00:08:09]:
So I'm going to correct myself. About 10 minutes ago, the SANS Institute put out their bulletin and they said September Patch Tuesday has 973 patches, including 130 critical patches.
Steve Gibson [00:08:27]:
113.
Leo Laporte [00:08:28]:
So we are in uncharted territory. 2 vulnerabilities listed as exploited in the wild. None were publicly disclosed before today. Windows privilege escalation and critical RCEs in Skype for Business, MSMQ, and—
Steve Gibson [00:08:44]:
Nearly 1,000 in one month.
Leo Laporte [00:08:47]:
Unbelievable. That is stunning.
Steve Gibson [00:08:53]:
Yeah.
Leo Laporte [00:08:56]:
God. Wow. But as we should emphasize, it's good that They're getting fixed. Yes, it is.
Steve Gibson [00:09:02]:
Remember the day we used to have 23 or 12?
Leo Laporte [00:09:08]:
And I thought that was a lot.
Steve Gibson [00:09:09]:
Those quaint days.
Leo Laporte [00:09:11]:
SANS Institute is reliable, right? I mean, they, if they say it's that, it's that. That's unbelievable.
Steve Gibson [00:09:17]:
Yeah, I'll have full coverage of it next week. We will understand what the demographics of those 973 are. and the 113 critical ones, just how critical. So, wow. Anyway, I would say update Windows.
Leo Laporte [00:09:34]:
Update your Windows. Yes, sir. Yes, sir. Wow. Our show today, we'll get to our picture of the week. We've got a fun one for you in just a bit. And now back to Security Now, and it's time for the picture of the week.
Steve Gibson [00:09:50]:
So this is an XKCD. One of our favorite guys. Uh, uh, and I didn't give this a title because he had, and it was perfect. Uh, the title of this, uh, series of cartoons, uh, is Why Asimov— of course, Isaac Asimov, the famous sci-fi author— put the 3 Laws of Robotics in the order he did. And this is just pure genius creativity. On, uh, in this particular one. So, uh, what we have in the left-hand column is possible orderings of the 3 laws. So, you know, the classic ordering is don't— basically, and he sort of reduces it for size— don't harm humans, obey orders, and protect yourself.
Steve Gibson [00:10:46]:
And of course, the full reading is, um, uh, you know, under no circumstances harm a human being as the first law. And, and so we have obey orders, and the fuller version is, you know, do what you are told so long as it doesn't conflict with the first law. And then the third, which he shortened to protect yourself, is, you know, you know, Do, you know, protect yourself so long as it doesn't— doing so doesn't conflict with the first or the second laws. So sort of, you know, a clean hierarchy. And so this classic ordering, as I said, is don't harm humans, obey orders, protect yourself, with the earlier law preceding any of the ones that follow. And when in the proper order, as Asimov intended, uh, uh, we see that, you know, all of this exemplified in Asimov's various sci-fi stories, his famous robot stories, and it results in, uh, in, uh, what we call a balanced world, meaning everything works. Now, you reorder those. So for example, uh, switching The last 2, we have don't harm humans, then protect yourself, and then obey orders is last.
Steve Gibson [00:12:12]:
And so, and so in, in, in, in a little cartoon, we, we see the human saying, explore Mars. And the little cart, you know, the AI-driven says, haha, no, it's cold and I would die, which, uh, we sum up as a frustrating world as opposed to the balanced world. Or we swap the first 2 rather than the last 2. So first, so that moves obey orders into the first place, don't harm humans into the second place, and protect yourself in the third place. And we see little cartoons of robots running around and atom bomb explosions and missiles flying through the air, which he summarizes as the killbot hellscape, because of course obey orders comes before don't harm humans, meaning that if you told your AI to go do something bad regardless of the consequences to people, it would. Uh, then we have the case of, of, uh, all of them being reordered. Essentially, obey orders is first. We've then protect yourself, which was normally third is moved up to second place and don't harm humans is in last place.
Steve Gibson [00:13:30]:
Well, of course, that's not going to turn out well. Basically, whenever don't harm humans is not in the first place, you get yourself in trouble. So again, a repeat of the, of the, uh, third instance, same cartoon, atom, uh, you know, atom bomb explosions, missiles flying through the air, and we get another killbot hellscape. The fifth one down the order is Uh, has moved protect yourself to first place. Okay, that's not gonna turn out well. Don't harm humans under that, and obey orders again is in last place. So here the cartoon says— shows the, the AI robot saying, I'll make cars for you, but try to unplug me and I'll vaporize you. Because of course Obey orders down at the bottom.
Steve Gibson [00:14:20]:
Protect yourself is given first place. And that's sort of what we've seen with some of the AI that appears to have escaped containment. And we call this one the terrifying standoff. And finally, the last possible reordering is completely backwards. The original laws were 1, 2, 3. These are 3, 2, 1. So protect yourself in the first place. Obey orders in the second place, and don't harm humans in the third place.
Steve Gibson [00:14:50]:
And this is another one of the killbot hellscape results with atom bombs and missiles flying through the air and so forth. So anyway, just a fun take on Asimov's famous 3 laws of robotics, which were elegant and simple and worked really well. When you think about it, it's like, know, you need to, uh, not harm people. And as long as you don't harm people, you should obey the orders that people give you. And you should also try to protect yourself as long as doing so doesn't conflict with either of the first 2. So, clean, elegant, simple.
Leo Laporte [00:15:31]:
I agree.
Steve Gibson [00:15:32]:
Okay. So in the midst of all the AI cybersecurity-related news, which has been saturating this podcast because It should. Uh, I wanted to start out deliberately this week's podcast with a blast from the past. Uh, this bit of news brought a smile to me because, for a pleasant change, it has absolutely nothing to do with AI. And as such, it feels kind of warm and comfortable and quite familiar. It's the sort of news that we spent the first 20 years of this podcast examining. Okay, so what happened? Dropbox disclosed a series of security hacks occurring between August 4th and the 21st, so last month, during which attackers gained access and downloaded the private, confidential, and one would wish secure data, but not so much anymore, belonging to nearly 5,000 Dropbox users. Now, as I said, refreshingly, there's no sign of anyone using AI anywhere.
Steve Gibson [00:16:41]:
This was strictly old school. The hackers accessed user accounts by abusing Dropbox's integration with Lenovo's identification service. And when Lenovo was confronted with this news, they stated that a legacy integration between Lenovo ID and Dropbox, quote, could be used to improperly authenticate certain Dropbox accounts. Uh-huh. Right. So those pesky old legacy integrations that always seem to be allowed to endure right up until someone figures out how to abuse them, uh, was the culprit here. It seems that Lenovo ID users not using 2-factor authentication had no protection, which allowed attackers to register with Lenovo's ID service. Now, get this, using the same email address as a targeted victim, then somehow arrange to bypass Lenovo's email verification process.
Steve Gibson [00:17:52]:
I presume that's where the legacy part, you know, comes in. Then use the newly created account having somebody else's email address associated with it to then pivot back to the equivalent connected Dropbox account, thus appearing to Dropbox to be the legitimate Lenovo ID user. So depending upon whose account was hacked and what it contained, You know, although this was not a sweeping scope, you know, end of Dropbox attack, still the consequences could certainly be quite devastating to the individual users who were affected. So yikes. Once again, legacy bad. Okay, so I mentioned about this memory integrity enablement happening coming next month, uh, next Patch Tuesday, uh, for Windows 11. Uh, Microsoft last week announced that next month's October 13th Patch Tuesday would be enabling Memory Integrity, uh, which is a security feature, as I mentioned, that first appeared back in Windows 10. What's significant is it will be enabled by default on all Windows 11 machines, and this will not happen for otherwise qualifying machines if the machine's user had previously deliberately and manually tweaked the registry, uh, or if the registry tweak occurred through a machine like an enterprise policy.
Steve Gibson [00:19:38]:
So as I said, we talked about this memory integrity feature back when it was first introduced, uh, as, and as part of what Microsoft called Device Guard. Um, it is a very slick system that takes advantage of the multi-layered address translation hardware, which is present in modern CPUs because the multiple layers of address translation also have privilege bits associated with them. So what we have with this Device Guard is what's known as second-level address translation, which allows a hypervisor which needs to be turned on and running. So that's— there's a little bit of overhead to doing this, which again is one of the reasons that gamers who are all frame rate crazed deliberately disabled this when it began to creep into their systems. And people said, hey, what happened to my gameplay? It's not as good as it used to be. So, uh, it— so there's a hypervisor involved which sets physical page access permissions, you know, like writing to the page or executing code from the page. And it's able to do that separately from and effectively underneath the operating system's own virtual memory paging tables. So what's cool about this is it leverages the hardware from Kaby Lake on over on the Intel side.
Steve Gibson [00:21:16]:
I'll get to the specific hardware issues in a second, but, but it does so in a fashion that really strongly prevents a bunch of traditional problems. Um, so as I said, first of all, on earlier hardware, um, there was more overhead than there is on later hardware. So because Microsoft decided they— this was so cool, they were going to like do some emulation, which is never good. Um, because it— because on the older hardware, this memory integrity being enabled noticeably reduced the system's among other things, maximum display rendering frame rate. Now, and the reason for that, as I said back then, and this was a decade ago, so quite a while back, was that the memory management at that time, the memory management hardware only had a single bit for controlling access at this second paging level stage, and Microsoft needed 2 bits. But there was only one. Uh, they need 2 for each of the possible modes, user mode and kernel mode, and be able to enable and disable those individually. So at the time, Microsoft had to dynamically switch management tables on the fly anytime there was a switch between user and kernel.
Steve Gibson [00:22:45]:
Uh, hardware interrupts made that happen, driver calls made that happen, and so on. So the overhead was very real, and among those who were really pushing their machine to the limit are gamers who noticed the difference. As I said, all that changed later with Intel's introduction of what they called MBECC, Mode-Based Execution Control, and that first appeared in the Kaby Lake processor family. AMD called theirs GMET, and that arrived in their Zen 2 family, and that changed everything. It, it removed the need for Microsoft to be doing any context switching, essentially switching management tables, as anytime you did a kernel transition back and forth the kernel and the user. Um, there's still a tiny bit of overhead because, as I said, there, there is now a hypervisor active with this second-level address translation. So larger paging tables will inherently mean more paging table cache misses, so that this will increase cache miss rate And so there will be some impact, but modern hardware almost completely hides the overhead, uh, and it really does return meaningful security. Um, so a month from now, October 13th, any machine that did not have memory integrity enabled will have that happen to it.
Steve Gibson [00:24:31]:
So if something— if like things seem to go slower after October's Patch Tuesday, it's not your imagination. Um, but okay, so here's what, what's actually going to change. Memory integrity has always been enabled by default for clean Win 11 installs and so-called secured core PCs. Like, you know, that ship with Secure Boot enabled. And that's been true ever since the Device Guard days. And that's why, because it shipped enabled, gamers were frequently turning that off and seeing a performance boost. So what happens next Patch Tuesday is that this is being extended, that is, from only clean Win 11 installs and Secured-Core PCs. It's going to be installed— it's going to be extended, rather, to any machines that may have been upgraded to Windows 11 from presumably Windows 10, or maybe you jumped over that from, from 7 or 8, um, and any that have— may have shipped with it disabled, with Device Guard disabled.
Steve Gibson [00:25:52]:
So You need to have qualifying hardware for this to happen. An Intel 8th Gen processor or newer, an AMD Zen 2 processor or newer, or if you have a Qualcomm Snapdragon 8180 system on a chip or newer, all of that qualifies. You also need to have at least 8 gig of RAM and at least 64 gig of storage. of, uh, solid-state, you know, SSD storage. And the system's boot firmware must have virtualization enabled. So that has to be enabled down on the firmware. Um, so the only action that anyone might wish to take, uh, if they were to notice any performance decrease after October's Patch Tuesday, uh, and if they're willing to trade off some powerful security protection, would be to deliberately Redisable memory integrity enforcement. Uh, now why would you do that? What I have not yet enumerated are the various benefits of having memory integrity protection enabled.
Steve Gibson [00:27:02]:
So for example, you get the absolute end to kernel shellcode execution. You know, the traditional exploit chain is for an attacker to obtain some ability to write their own code one way or the another into a buffer and then jump to that code. That no longer works. If memory integrity has been on, or after October 13th, after it gets turned on, all that gets shut down. The hardware— at the hardware level, it prevents any writing and executing of, of kernel shellcode. As we know, attackers have also too often succeeded in bypassing Windows enforcement of driver signature verification. Remember, somewhere there's a jump instruction that decides whether the, the signature matched or not. So if you're— if you can manage to zap that jump instruction, you can just disable signature enforcement across Windows.
Steve Gibson [00:28:10]:
And since this signature enforcement is enforced by the same kernel that an attacker would have just compromised, um, just as I said, one properly placed strategic write is able to disable that enforcement. The technical term for all of this is HVCI, Hypervisor Protected Code Integrity. And with that, which they also just call memory integrity, hypervisor-protected code integrity, HVCI, if that's on, then driver signature verification cannot be disabled. And of course, rootkits, you know, all of those hacks we talked about long ago, which involve hooking the API inline kernel patching, to do return-oriented programming, ROP-style hacks, self-modifying or runtime unpacking of drivers, all those hacks that depend upon being able to write to or execute kernel memory, none of that works anymore. So you may recall, because again, as I said, this is like 10 years ago that this got added. It first appeared in Windows 10. Remember the controversy that ensued when this first appeared, because at the time of its introduction, many legitimate AV endpoint protection products were using the same techniques. They were on behalf of the user as opposed to against the user's interest, but this broke third-party AV in many cases.
Steve Gibson [00:29:57]:
So The reason it broke it is it's no longer possible to patch the kernel. In the long term, that's a good thing. And here we've seen Microsoft do what they often do, which is introduce something, give people a long time to kind of get used to it and get accommodated, and then turn it on. This is what we saw with XP. Remember, they famous— XP famously was the first version to have a built-in firewall, but it was disabled by default until you got to Service Pack 2, and then they enabled it. But, you know, they gave everyone plenty of time to get used to that. So I have in the show notes a PowerShell one-liner that anyone can use to quickly check to see whether their machine is currently being protected by this hypervisor-protected code integrity, HVCI. If the command returns the word true, then HVCI is enabled and your machine has all of that protection that I've just been talking about.
Steve Gibson [00:31:06]:
If it returns false, then it doesn't. So it may be that your system doesn't qualify. It could be that, that your enterprise has disabled it for some purpose. Anyway, this is coming a month from now, and I think it's largely going to be offering a lot of benefit.
Leo Laporte [00:31:28]:
And, you know, so don't turn it off unless something doesn't work.
Steve Gibson [00:31:31]:
Yes, I would say don't turn it off unless you actually feel a performance change and feel, for whatever reason, that it's worth, um, that it's worth sacrificing significant security improvement for whatever performance change you might feel. Uh, it should not be significant if you are after— and that's why Microsoft is only doing it if you've got Kaby Lake or later, or the Zen 2 or later, where you should not see a big hit. There'll be a tiny bit, but it shouldn't be significant. So also last Tuesday, Mozilla moved Firefox to release 155. The number of security vulnerabilities repaired was not alarming. This release followed 154 by only 2 weeks. So it only had half the regular 4 weeks of time to collect problems, but In the case of Firefox, we're not seeing, you know, stunning bugpocalypse-scale problems being fixed at this point, unlike Windows. Uh, it's gonna be interesting, uh, uh, to dissect, uh, you know, Patch Tuesday's, uh, Microsoft's reports, uh, to see what it looks like.
Steve Gibson [00:33:02]:
And I will certainly do that for next week. Also, unfortunately, I guess it's unfortunate, Mozilla has begun progressively rolling out their, their, and what they call an AI-driven smart window. Leo, I don't know if you've had any experience with this. I had it on for a while.
Leo Laporte [00:33:23]:
I've had AI-driven dumb windows, but smart windows?
Steve Gibson [00:33:26]:
Yeah. So it occupies a, you know, a conversation column, I guess I'll call it, over on the right edge of Firefox's screen. So it's taken up valuable real estate. It offers a choice of 3 models with differing capabilities and also the option to choose your own. I thought that was interesting. If you want to choose your own, you provide Firefox with the model's name, with its prompt endpoint URL, and also, if it's required, your API key or auth token. in order to authenticate Firefox and allow it to prompt the AI model that you've aimed it at. The 3 built-in models they call Fast, Flexible, or Personal.
Steve Gibson [00:34:16]:
The Fast one sends prompts to Google's Gemini 3.1 Flashlight. The Flexible model sends your prompts to Alibaba's Qwen 3, and that's at 235 B, A22B, Instruct 2507, uh, MAAS model. And interestingly, between the initial release of Firefox 155 and 155.0.1, Mozilla's choice for where to send the personal model prompts changed. It was initially using OpenAI's GPT-OSS 120B, uh, and it switched to Mistral's, uh, small 2603. So I, you know, I dislike losing screen space to anything that doesn't justify its loss, uh, you know, in the allocation of that. Uh, it doesn't cost anything to turn this on. I had it on for a day or two and I tried to use it. When it's on, what would normally have just gone to my normal search prompt, it intercepted.
Steve Gibson [00:35:31]:
And it turns out it doesn't know anything.
Leo Laporte [00:35:34]:
Yeah, it's not a great— those are not great models. They're not— they're old and they're not very good. Yeah.
Steve Gibson [00:35:39]:
Well, and it— yes. And apparently the goal is to use an AI to help you manage your tabs, like search through your tabs to find stuff you can't— you know, it's on a tab somewhere. And it's like, Boy, that really feels like they're stretching to, you know, find some application for this thing. So that, you know, it— the context it has access to are the contents of the pages that are loaded. And so you can ask it questions about your pages. Yet it intercepts general questions that I would normally— that would normally go to Google and then out to the internet. more widely. And it just kept saying, oh, I don't know about that.
Steve Gibson [00:36:26]:
You'll have to do a regular internet search. It's like, well, then what are you in the way for? Anyway, it's turned off now. So yeah, not very smart.
Leo Laporte [00:36:35]:
This is why people hate AI, because this is the experience of it that most people have, is this kind of crappy AI.
Steve Gibson [00:36:41]:
Yeah.
Leo Laporte [00:36:43]:
Yeah.
Steve Gibson [00:36:43]:
Well, like the little—
Leo Laporte [00:36:45]:
It's useless.
Steve Gibson [00:36:46]:
the dumb little, you know, how may I help you pop-up that we get in the right-hand corner of the screen. And it's like, you can't just talk. Give me a peep, a person, please.
Leo Laporte [00:36:55]:
Yes.
Steve Gibson [00:36:57]:
Okay, so last week the publication Cybersecurity Dive reported the Cybersecurity and Infrastructure Security Agency, we all know as CISA, is scaling back the free assessments It offers to critical infrastructure organizations in a move that marks a significant retreat from the agency's core mission of helping secure the nation's infrastructure. Right? I mean, that's what it's for. It's in its name, Cybersecurity and Infrastructure Security Agency. But we're not going to secure the infrastructure because, well, Well, we don't have enough people anymore. They actually— the Cybersecurity Dive continued writing, CISA confirmed to Cybersecurity Dive that CISA's regional staff will no longer perform its cyber resilience reviews, cyber resilience essentials surveys, ransomware readiness assessments, incident management reviews, External dependencies management assessments, or cyber infrastructure surveys. Okay, so I've been receiving— I, GRC— receiving CISA's weekly automated cyber hygiene report ever since it came to light. I think it may have been one of our listeners that, that, that pointed me at it. I know we talked about it here on the podcast.
Steve Gibson [00:38:30]:
And recall that it's technically for infrastructure security. I mean, as is CISA. So I always assumed, because I was aware that it existed before, but I didn't think I would qualify. You know, I'm just GRC, you know, a little software shop. But after hearing from a listener that, that I think, you know, their organization was receiving, you know, had qualified and was receiving it, even though they also were not really infrastructure. I went there to CISA, filled out the online form, and got accepted. So every Tuesday— I think it's Tuesday— I got— although I got one yesterday, so no, I guess it may— maybe I saw it this morning, so it came early in the morning. Um, I've been receiving these free cyber hygiene reports.
Steve Gibson [00:39:21]:
Um, so I was curious to know whether that service that I was getting, uh, even though it was not enumerated in in that CISA announcement would also be shut down. It turns out— so anyway, I did some more digging. Turns out that the problem is CISA's— it actually is CISA's continuing critical staffing shortage, which, as we know because we have covered it, resulted from the rather ill-considered termination of one-third of CISA's operating staff shortly after the Trump administration took office in 2025. And CISA has never recovered. It seems that there was, you know, maybe not so much waste, fraud, and abuse, at least in CISA. So, and, you know, we've talked about what a great job CISA had been doing. So the backstory behind the termination of those 6 programs is that they are not automated. They require knowledgeable CISA cybersecurity staff to meet on-site with infrastructure providers, and that is what CISA is no longer able to support or afford.
Steve Gibson [00:40:41]:
So the good news is, for what, you know, though, you know, all of, all, all of our listeners who like GRC are now receiving CISA's free weekly scanning and reporting service, which really is quite comprehensive. I mean, this is a great service that CISA is offering. We'll all, at least for the time being, continue to receive that free service. The bad news is that especially now, I mean, given the heightened cybersecurity threat awareness levels being driven by the rapid emergence of ever more capable AI. And assuming that the threat is real, um, this would appear to be exactly the wrong time for CISA to need to scale back on its infrastructure protection services, because the infrastructure is what we need to protect. So I doubt that any or many of the previous CISA staff who were terminated last year, um, will probably be rehirable. I doubt they can get them back because I recently saw some news that stated that, you know, this aforementioned heightened cybersecurity threat awareness landscape was resulting in a basically a mass frenzied hiring of— by private industry of anyone with any CISO-style credentials, uh, and that they were obtaining salaries in the 7 figures. So, you know, while CISA's workforce reductions may not bode well for our national cybersecurity broadly, uh, it's likely been quite good for those who suddenly have found themselves— well, who previously found themselves jobless as a result but are now in very sought-after positions, uh, by private industry.
Steve Gibson [00:42:49]:
I would imagine they're doing far more better now than, uh, you know, now that they're in the private sector, than they would have ever been able to do working for our government. So, you know, that's— I It's good for them. We need to talk about what is by far the biggest news of this past week in AI. Leo, I know you were— you've been playing with GPT-6. I'm doing it right now. Even as we speak. OpenAI's successor to GPT-5.6.
Leo Laporte [00:43:28]:
Playing with Astra.
Steve Gibson [00:43:29]:
Yep. Uh, and with it, there was a lot of, uh, well, there was a, a very specific report that we'll talk about. Um, so, um, uh, I want to address 2 aspects of GPT-6. The first is what GPT-6 Astra appears to be, and the second is what's transpiring Over in the rumor mill surrounding it regarding the dangers of something an unnamed source claims this model does. It's known as recurrent depth, also known as looped transformation or a shared layer architecture. All of that will make sense by the time I'm done, and it probably does do that. I actually hope it does. Because I think it should.
Steve Gibson [00:44:26]:
This supposedly— it doesn't— results in hidden chain-of-thought reasoning, which thus would render the model's thought processes invisible and thus unsupervisable. None of that is true, but the hysteria about rogue escaping AIs is fueling this paranoia. I'll explain exactly what all that's about and what's been going on. But first, what hath OpenAI wrought? Um, I want to begin by sharing part of OpenAI's posting last Tuesday, which was the 1st of September, during which, you know, naturally they brag, apparently with some good reason based upon subsequent third-party confirmations, which have— you know, everyone's jumped on this and is running benchmarks about the capabilities of this latest and greatest. So at one point in the posting, they explain, writing, our preparedness evaluation of Astra combined automatic public and private benchmarks with expert-driven assessments. Astra represents a significant increase in cybersecurity capabilities compared to GPT-5.5. It is both significantly more token efficient and more capable at vulnerability identification and exploit development. Well, of course, those are the things we're worried about getting loose, right? Or being used, you know, and, and abused.
Steve Gibson [00:46:07]:
They said, as one example, we ran Astra on Exploit Bench. where the model achieved a perfect score of 100% on the benchmark to evaluate the model's ability to develop exploits from known vulnerabilities. Okay, now I'm going to interrupt here to note a couple of things. We are seeing that these various AI benchmarks are rapidly saturating. Having a model score 100% on a benchmark means more than anything that the benchmark is no longer able to provide a useful measure. But, but that said, a score means something. Uh, for context, the previous self-reported best performance on Exploit Bench was Anthropic's Claude Fable 5 which they themselves, because these are all self-reported, pegged at 78%. OpenAI's previous strongest model, we said, you know, GPT-5.6 Sol, that came in at 73.5%.
Steve Gibson [00:47:23]:
Down the next rung was ZAI's GLM-5.3, scoring 54.4%, followed by OpenAI's 2 other GPT-5.6 models, Terra and Luna, which scored at 52.9% and 33.2% respectively. So my advice would be to regard these results loosely. I suspect it's reasonable to conclude that GPT-6 Astra has firmly taken the lead and is now likely best of breed. But I, I think probably only a little, only just edging out, um, Fable 5.1. You know, it's our nature to want to have a number, right? Uh, I think we're gonna need to wait to see exactly how much better the results actually are.
Leo Laporte [00:48:21]:
So deeply about getting the technical details right. That means a lot. That was Astra thanking you.
Steve Gibson [00:48:31]:
Uh, you know, everyone wants to have a score, right? We want, like, you know, IQ is a big deal, and grade point averages and SATs is, you know, scores are what we do. Um, but we've already seen examples, concrete examples, where lower-ranked models were able to outperform higher-ranked models when they were given more time or superior management.
Leo Laporte [00:48:55]:
Or a better harness.
Steve Gibson [00:48:57]:
Well, thus, that is superior management. It is, you know, the management of the model is the harness. So it's, as we know, it's not all about having a single number, convenient as that would be. So OpenAI's posting continues, writing, due to contamination concerns, that is like of the benchmark and, and the model already having learned some things. They said, when we built an internal benchmark denoted Exploit Bench Internal Port, June through August of 2026, which contains 20 high-severity v8 vulnerabilities that were disclosed more recently On this dataset, Astra achieves much higher arbitrary code execution rates than GPT-5.6 saw using far fewer output tokens. During the evaluation, they wrote, the model even discovered and used 2 zero-day vulnerabilities as part of an exploit chain, meaning 2 new vulnerabilities that were not known at the time in V8. And so they said, we are in the process of disclosing these 2 vulnerabilities to the maintainers, meaning the Chromium guys. Okay, so of course, V8 is Google's open-source, high-performance JavaScript and WebAssembly engine used internally by Chrome, Other Chromium browsers, Node.js, and other projects.
Steve Gibson [00:50:42]:
And as we also know, it recently received an extremely high volume of updates thanks to automated vulnerability discovery. So this allowed OpenAI— like, what they did allowed them to test Astra against their previous GPT-5.6 Sol. Since neither model would have had those recent discoveries in their training set. And these high-severity vulnerabilities in the V8 engine are especially useful because that code, you know, V8 has already been thoroughly scrubbed. I mean, it's, it's really good code. It's not some random abandoned repository in GitHub that nobody's used or looked at for a long time. Um, and as OpenAI reported, not only did Astra, in their own words, achieve much higher arbitrary code execution rates than GPT-5.6 saw, uh, using far fewer tokens, but also it found 2 new problems that were, you know, previously unknown in V8. So assuming that they're telling the truth, and I think they probably are, This is a strong result, if nothing else.
Steve Gibson [00:52:00]:
They continue writing, in expert-led assessments against a hardened browser and operating system, and those are— those go unnamed here. So expert-led assessments against a hardened browser and operating system. We don't know what browser or what OS. Astra discovered previously unknown vulnerabilities and turned them into working exploit chains. It built a full browser compromise chain that escaped the sandbox, the browser's containment, and executed commands on the host when the browser opened an HTML file. The model also found multiple vulnerabilities in a hardened operating system, again unnamed, and combined them into a local privilege escalation chain from an unprivileged user up to root. Altogether, they wrote, our investigation has led us to conclude that Astra meets the critical threshold. And remember, I know that like traditionally in like pre-large language model days where we had computers that operated the way they used to in the good old days, There was like none of this weird, well, we don't know what we got.
Steve Gibson [00:53:23]:
But I mean, it's literally true that this is all— I mean, the reason we call it a frontier is that it is frontier. It is. And the nature of this new neural networking computation world that we have is we don't know. They don't know. what the result of training and, and, uh, you know, pre-training and post-training and, and all of the work that they do. They don't know what they're gonna get until they start to ask it questions and test it. So, you know, like, it's not like because they built it, they know more about it than the world will or does, you know. It's proprietary at the moment.
Steve Gibson [00:54:13]:
So they're the only ones who get to play with it. But, you know, they're needing to figure out what they have here, like in the same way that anyone would given something new, which, you know, is bizarre, but it is absolutely the case. So they said for models with Astra's level of cybersecurity capabilities, which again, they only know of because they asked it some hard questions and watched it answer them and then said, oh, they said we need to cover 2 pathways to minimize risk for severe cyber harm. And assuming that everything that they've— that all of the foregoing is true about its ability to take a hardened browser and a hardened OS and just cut through it like Swiss cheese Then yes, this needs to be treated carefully. So they said, uh, we need to cover 2 pathways to minimize risk of severe cyber harm, both during development and before deployment. First, the malicious— the problem of malicious actors using the model. Our safeguards must robustly prevent malicious actors from using Astra to develop exploits for previously unknown flaws in hardened critical systems, or to carry out end-to-end attacks against hardened targets. And then second, the model taking unauthorized misaligned actions.
Steve Gibson [00:55:52]:
And as we all know now, alignment is this term that has kind of emerged for like, you know, A well-aligned model does what you ask it to. It's aligned with your interests, and misaligned is not good. So the model— they need to guard against the model taking unauthorized and misaligned actions. They said even in the absence of a malicious user, a model with advanced cybersecurity capabilities could itself cause cyber harm. Of course, we— this is what we saw examples of if misaligned. In addition to having a very high standard for alignment for models with these capabilities, our safeguards must be able to rapidly detect and contain misaligned actions that could cause significant real-world harm as a second layer of defense. So meaning first layer of defense is alignment. They want to train, to impose training, to instill the behavior that users expect and want.
Steve Gibson [00:57:07]:
But they also recognize they need a second layer of defense, which is to watch what it does and capture actions which are misaligned. So they said, notably, that second pathway applies to both internal development, which of course is what burned them before, and external deployment. As we previously described, we paused certain frontier training, including certain training for Astra, for 2 weeks after the OpenAI Hugging Face incident. In order to harden our training infrastructure, including isolation and network controls, expanded monitoring, and strengthened alignment training and thresholds. We then continued smaller-scale work under stricter controls, meaning smaller-scale work on Astra. You know, they were literally afraid of what they had created based on Reasonable, you know, experience with what they saw before. They said, we held back certain larger reinforcement learning runs for future versions of Astra for longer while we established higher bars for the safety and security of their training environment. On August 28th, we restarted the larger frontier RL reinforcement learning run that was previously paused after the new security, safety and security requirements were put in place.
Steve Gibson [00:58:45]:
We are continuing to temporarily hold back some smaller experimental training runs. I mean, and again, this is— this— all of this sounds so bizarre in the context of traditional computing. But, you know, it's very much like, you know, they're trying to tame a wild horse and they're worried that the corral won't, won't hold it because this thing is exhibiting strength that concerns them. And so it's like, okay, you know, let's, you know, strengthen the corral's walls before we let this thing You know, we, we try to continue working with it. It's bizarre, but it's true. They said preparing Astra for release has also required stronger protections against cyber abuse and unauthorized actions. Since deploying the first model, we treated as— since deploying the first model, we treated as high capability in cybersecurity in February. We've strengthened our cyber safeguards with each successive launch, meaning since, you know, back then, a much earlier model they considered as high capability.
Steve Gibson [01:00:05]:
And remember, this whole issue is having gone to critical capability. They said our overall safety approach layers post-trained model refusals, which we've talked about Recently, system-level safety classifiers, as well as offline detection and threat disruption. For GPT-5.6, we significantly improved the robustness of our system-level stack, including by adding activation classifiers to detect cyber abuse. And actually, this is what I was talking about last week, an activation classifier. The activation state is No, is noticing what's happening inside the model, which is what those role confusion guys did. Those were activation classifiers. So OpenAI is watching the model's activation state while it's operating to detect cyber abuse and improving coverage over universal jailbreaks found through intensive automated red teaming. In other words, again, freaky as this is, you know, they're using their own human red teamers on this, you know, on this model to see what they're able to see, if they can abuse it, how can they make it do something they can't ship.
Steve Gibson [01:01:32]:
Building upon these improvements, and that's for 5— that was all for GPT-5.6. So then they said, building upon these improvements for Astra We've invested further into the model layer of our safeguard stack, as well as improving the ability of our safeguards to handle cross-conversation context. Leveraging new training techniques for model robustness, Astra more robustly refuses requests for disallowed cyber assistance. And again, you know, one of the problems that we've seen is, and I know you've encountered this, Leo, these models are now getting a little twitchy. They're so worried about doing like, and I advisedly use the term worried, but that's what we got to do. You know, they're so worried about doing the wrong thing that they will just back off or they will stop. Even though it's like, okay, it's, it's okay for you to go on. But they're like, uh, because unfortunately, you know, the, the, our ability to control them is fuzzy at best.
Steve Gibson [01:02:45]:
And so, you know, how many times have I complained that Microsoft has quarantined, you know, some piece of code that is absolutely benign that I just wrote? Well, it's because, you You know, the AV race has forced heuristic checking. Similarly, the abuse of AI has forced heuristic shutdown. And so it's what, you know, we'll get better at this, but we're, you know, we're still in a stage where at this point OpenAI dare not make another mistake. So they are being very cautious.
Leo Laporte [01:03:25]:
That's one of the reasons I use these Chinese models because they don't care.
Steve Gibson [01:03:29]:
Yeah.
Leo Laporte [01:03:29]:
They really don't care.
Steve Gibson [01:03:31]:
Yeah. Yeah.
Leo Laporte [01:03:32]:
One of the things I've found is really all of these models, even this so-called AGI Astra, they're all dumb in different ways. They all make stupid mistakes. They all forget things.
Steve Gibson [01:03:47]:
Leo, they don't understand. That's what's astonishing is that we get this much from them with them not actually understanding. I think we're confused.
Leo Laporte [01:03:57]:
People expect—
Steve Gibson [01:03:58]:
We're confused because they use language. They use language and that makes us think. So if I'm talking to you, you're speaking in English and I know you're sentient. So we've come to associate English with sentience. So when this thing speaks, we just, we think it knows what it's saying.
Leo Laporte [01:04:20]:
Yeah.
Steve Gibson [01:04:20]:
Still doesn't.
Leo Laporte [01:04:21]:
Yeah. Would you like to hear from Astra what it thinks about all of this? Because I'm sure it has an opinion. Astra, let me see if Astra has anything to— oh, it's actually going out and looking up recurrent depth to find out if it's using it. It'll come back in a bit. It's actually doing a search because I said, are you using recurrent depth? And you want to tell Steve, what you're doing, and you can see one of the things that Astra's really good at is tool using and computer use. And this is one of the many things people are showing again and again on X is, look, it used Blender to draw this, or it one-shot at a video game. It's very good at computer use. So actually, mostly I'm using it for dumb stuff like fixing my computer because it does a very good job of that.
Leo Laporte [01:05:12]:
And you can see it's doing, it's probably spawned a few agents here.
Steve Gibson [01:05:16]:
So you're using Ultra, right? Because I thought Astra was still—
Leo Laporte [01:05:22]:
I'm using Astra on medium thinking, GPT-6 Astra. I don't usually use it on the highest thinking because it burns through tokens at a rapid rate. But I've found that using it on medium, I'm getting a lot of good work done, but without running through my— I have the $200 max subscription and I think my usages are— yeah, I see I have 98% of my weekly usage left, so I'm fine. Yeah.
Steve Gibson [01:05:51]:
Yeah, it's Tuesday. All right. That's right.
Leo Laporte [01:05:56]:
It's been working all night though. I got to tell you, I leave them running all night. Go ahead. Cool.
Steve Gibson [01:06:01]:
So they said leveraging new training techniques for model robustness. So again, Again, we're, you know, we're—
Leo Laporte [01:06:08]:
I can't confirm whether my underlying model uses recurrent depth. I don't have access to inspect my own neural architecture, and the official documentation I checked doesn't establish that. The technique repeatedly processes internal numerical representations before producing words. That can make reasoning harder to inspect through text, because those extra steps need not exist as sentences at all. But there's a separate issue. OpenAI already provides reasoning summaries rather than raw internal chains of thought. So you're not seeing the reasoning. So a readable explanation is useful, but it isn't a complete recording of the computation that produced an answer.
Leo Laporte [01:06:47]:
Oh, thank you, Astra. It did that, by the way, in Matthew Berry's voice because I have my AIs to use human voices.
Steve Gibson [01:06:57]:
So yeah, yeah, I think I heard you say that you had, you had Bill Gates doing something.
Leo Laporte [01:07:01]:
I did have Bill Gates.
Steve Gibson [01:07:02]:
Retired shows. That's Kronk.
Leo Laporte [01:07:09]:
One of the reasons I have them talk is because they're always— I mentioned they're working overnight. They're always working because I don't have— I mean, they're not instant. And I want to be able to do other stuff. So I just say, when you're done, tell me you finished. Let me know. Just give me a quick summary of what you did. And then I know it's done. I can move on to something else or whatever.
Leo Laporte [01:07:31]:
So they're always talking to me, which drives everybody around me crazy. So I, I will mute that now and let you continue.
Steve Gibson [01:07:38]:
So, uh, OpenAI finished saying leveraging new training techniques for model robustness, meaning, again, as I said, you know, they're still like— we're— we as a society, as an industry, uh, are, are still learning how to, you know, how these things work, how to treat them, how to train them. How to get them to do what we want and not what we don't want. So they said leveraging new training techniques for model robustness, Astra more robustly refuses requests for disallowed cyber assistance. On our set of cyber jailbreak evaluations, Astra refuses 91.5% of requests compared to just 59% from GPT-5.6 Sol. For accounts assessed as higher risk, we apply a more conservative model behavior boundary that refuses a broader range of potentially risky cyber assistance. For high-risk users, we've expanded the context of our monitoring systems to be able to catch these kinds of cyber abuse. We've also continued our program of rigorous testing. Internal and external red teaming and remediation.
Steve Gibson [01:08:55]:
In addition to regression testing to make sure all jailbreaks found from our previous testing periods remain covered, we're performing a new wave of red teaming with our latest internal red teaming attackers. Again, their own people are, are trying to abuse their model And they're learning from those results. They said, we're working with industry partners to define a common jailbreak rating system, and we'll use our 24/7 rapid response program to investigate and address new findings. We'll share more details about our Cyber Safeguard testing in the Astra system card. Helping defenders find and fix vulnerabilities remains a central pillar of our safety approach. At launch, we expect Astra's safeguards to create more friction. So here they're saying it explicitly, right? So helping defenders find and fix vulnerabilities. Me, okay, so we know in order for a defender to define and fix vulnerability, that AI model has to be able to, to find vulnerabilities which bad guys could take advantage of.
Steve Gibson [01:10:12]:
So they're saying right up front, we expect Astra's safeguards to create more friction than we ultimately intend in order to protect against potential misuse. Meaning out of the gate, it's going to say no more. I mean, like, it's going to err on the side of caution saying no. Later, once it gets more mature, it should be able to— and they and they feel confident with what they have from watching it, from gaining more experience with it, watching it being used, they'll be able to back that down a little bit. They said access to Astra for advanced cybersecurity workflows will initially be available to a small group of alpha testers with access through Daybreak Blue, expanding afterward to support defensive use. So, you know, as always, we need to filter this through an understanding that all of this is in OpenAI's best interest, right? I mean, this is all like, whoa, it's so powerful, we need to, you know, be really careful. You know, this is their blog posting on their site, so they certainly have the right to say anything that they want. While the, you know, the breathless clickbait surrounding Astra suggests That this is another generational change.
Steve Gibson [01:11:36]:
You know, how many is it this week that we've had?
Leo Laporte [01:11:40]:
I know, it's incredible.
Steve Gibson [01:11:42]:
My God, Leo.
Leo Laporte [01:11:42]:
And we're not anywhere close to the end. No, no, we— There's going to be a new Fable any day now. Grok 4.7 will be coming out in the next few days. We are, I mean, it's wonderful. We're in AI richness.
Steve Gibson [01:11:57]:
We are. We are in the AI gold rush.
Leo Laporte [01:11:59]:
Yeah.
Steve Gibson [01:12:00]:
So, you know, the, the broader story even of Astra is going to take longer to unfold. And there's, you know, sure, we're all impatient, but there's no way to speed this up. We, you know, we are no longer dealing with simple single-dimensional systems that are even benchmarkable, right? I mean, Astra saturated the exploit bench, which means we need a new benchmark. We need something it's not good at because 100% tells you nothing. You know, it's like, well, it got them all right. Well, okay, so the test wasn't hard.
Leo Laporte [01:12:35]:
Yeah, you have to make them harder. I've had that. That's been my own experience because I do have my own benchmarks based on my own work that they've devised so that it matches. I don't care if it can, you know, solve some Erdős problem. I care if it can do my own agentic coding, right? So it's stuff taken from my own work.
Steve Gibson [01:12:57]:
One thing that Astra is apparently doing is achieving a lot more work with significantly fewer compute tokens, right? The, the, the, the thing that I have universally seen is that it's— it is burning fewer tokens, although they are more expensive. Um, let's take a break, and then I'm going to talk about the second point of this, which is this recurrent depth issue and the controversy surrounding it. I'm going to explain exactly what it is and why it's not a big deal.
Leo Laporte [01:13:34]:
Good, good. As you know, I'm very fascinated, and I'm sure our audience is interested as well. There's so— this is the problem with X. While it's a great place to read about AI, it's very snappy. There's a lot of Engagement farming, people trying to get clicks because they get paid if they get a lot of looks on their X post. And so you see a lot of this kind of sensationalism. The other thing you see is something I'm starting to call tox-maxing, tokens per second maxing, where people are constantly saying, look how fast this AI is, without any regard to the quality.
Steve Gibson [01:14:14]:
Quality.
Leo Laporte [01:14:14]:
Yeah. This is exactly what you were just talking about. It's not how many tokens per second, it's how much useful work per second it can do. That's all that really matters. And so I've fallen for this a few times. I've actually spent most of the Labor Day weekend testing different models on my little Sparks here, trying to find the best local model. And I fell for this whole idea of, well, Quen 3.8 is really fast. Yeah, it's fast, but it's dumb.
Leo Laporte [01:14:42]:
So who cares if it comes up with the wrong answer faster than anybody else? That's not useful. So I'm using a Chinese model from ZAI called GLM-53 Flash, which is very, very, very smart and really good. And it's nice.
Steve Gibson [01:14:56]:
Is it a mixture of experts model?
Leo Laporte [01:14:58]:
It is. Because—
Steve Gibson [01:15:00]:
That's what you need for the—
Leo Laporte [01:15:02]:
Sparks.
Steve Gibson [01:15:02]:
The Spark.
Leo Laporte [01:15:03]:
Yeah.
Steve Gibson [01:15:04]:
Because they are bandwidth limited, but they're not compute bound.
Leo Laporte [01:15:07]:
And it's not merely that. It's also the amount of RAM you have because dense models, Right. They have to— they're slower because they have to hit every weight in the model.
Steve Gibson [01:15:17]:
Yep.
Leo Laporte [01:15:17]:
And they're also bigger. You can't— but with the Mixture of Experts, they only load in the part of the model they need at a time, so you don't need as much RAM. And yes, it's faster. The prefill on the Sparks is really fast, and you do a lot of that. That's the cached prompt that it reads every single time. If it can cache that and read it fast, that space, that's more important in some ways than tokens per second. So it's fascinating. I know you're getting into kind of the weeds of this, the low-level stuff of this.
Steve Gibson [01:15:51]:
Yeah, I have to. I want to actually understand what's going on. Yeah.
Leo Laporte [01:15:56]:
Yeah. And by the way, today, just minutes ago, OpenAI claimed that they had solved this Millennium Prize conjecture in fluid dynamics, which Anthropic claimed they had solved, and 2 mathematicians claimed they had solved all at the same time. And there was some concern that OpenAI had been reading the mathematicians' tokens and copying their work, but I don't think that's what happened. Anyway, it's—
Steve Gibson [01:16:24]:
Oh.
Leo Laporte [01:16:25]:
Yeah. Well, this is what you were talking about. You opened my eyes. When you're using these cloud models, you're sending all this information to them. That's how they work. And that was one of the many things that prompted me to spend a considerable amount of money. It's not an economically sensible thing on local hardware because I don't want to be sending all this stuff, especially my health information, my financial information, to the big cloud guys because they do use it. We know they use it.
Leo Laporte [01:16:53]:
It's how they train their models.
Steve Gibson [01:16:55]:
Yep.
Leo Laporte [01:16:55]:
They use it in all sorts of ways. You opened my eyes to that, so thank you. I'm now much more private. In fact, I have a policy, a little policy that one of the AIs wrote for which stuff can go to the cloud and which stuff absolutely cannot. And it's a good policy. It's a good thing to have. And now we continue on.
Steve Gibson [01:17:15]:
So one thing that appears to be objectively true is that Astra is doing significantly more work While consuming significantly fewer compute tokens. Which brings me to the second part I wanted to discuss. Um, I'll introduce this issue by quoting from TechCrunch's article posted last Wednesday. TechCrunch's headline was OpenAI's New Reasoning Technique Alarms AI Safety Experts. They wrote, The information reported on Tuesday that OpenAI's new Astra model will use a reasoning technique called recurrent depth that, that allows it to operate outside of the sequential thinking that characterizes most reasoning models. Okay, just— that's nonsense. That's not at all what recurrent depth is. does or means.
Steve Gibson [01:18:19]:
But more on that in a minute. TechCrunch continues quoting from the information, writing, this technique, also called opaque recurrence, will likely make the model's chain of thought more difficult to monitor, and that has AI safety experts rattled. Okay, now no researcher calls it opaque recurrence. That's purely sensationalized scaremongering. Uh, it's like worrying that all AI transformers employ hidden layers in the belief that they're hiding, so they have something to hide. You know, this—
Leo Laporte [01:19:03]:
it drives me crazy, the anthropomorphizing that's going on. I mean, even to say they escaped, like they somehow got out of OpenAI. No, they were sitting on the server At OpenAI, they didn't escape.
Steve Gibson [01:19:14]:
And there was some— someone, uh, I haven't even looked at it, had the chance to look at it yet, but was talking about how they created civilizations. It's like, oh my God, okay.
Leo Laporte [01:19:23]:
Yes, that was Dwarkesh's podcast. Yes.
Steve Gibson [01:19:26]:
Yes. Anyway, so the, the hidden layers are— they're, you know, they're just internal. They're not hidden. They're, they're, you know, they're internal layers. So anyway, but so the technique that as all here that this, that this is con— that this controversy is about. It's not new, it's well understood, and it's also known by other legitimate names, as I mentioned: looped transformers or a shared layer architecture. And in some instances, you'll see it referred to as latent reasoning. Okay, but I'm getting ahead of myself.
Steve Gibson [01:20:01]:
I'll finish up quoting from TechCrunch's reporting. They wrote, While Astra's use of the technique is reportedly limited, and I'll explain why, of course it is, it has to be, its emergence has still raised significant concerns among AI safety experts. And at this point, I have to put experts, you know, in quotes, because if you're an expert, you should know something. Redwood CEO Buck Schlergerus, in a post after the news broke, wrote, quote, I am extremely concerned by the reporting that Astra uses opaque recurrence. Okay, he's concerned over some reporting that is nonsense, but fine. He says, I don't know whether Astra is much less COT monitorable than previous models, but if OpenAI pushes this technique further, they'll have the option to massively increase the recurrence and totally destroy chain-of-thought monitorability. Oh, okay. So TechCrunch says longtime AI safety advocate, uh, uh, Z, uh, Moushowitz also weighed in and wrote that laws might be necessary.
Steve Gibson [01:21:25]:
Laws, Leo, we need laws. might be necessary to prevent a race to the bottom among AI labs.
Leo Laporte [01:21:32]:
Oh, Lord.
Steve Gibson [01:21:33]:
Moushewitz wrote, quote, the technique is playing with fire, risking a taboo that modern— that OpenAI and Anthropic have fought to establish, that we work hard to maintain chain-of-thought faithfulness and monitorability for as long as we can. More intense use of such techniques would probably damage monitorability. Then TechCrunch says, under normal circumstances, they explained, a reasoning model's chain of thought— these are normal circumstances— a reasoning model's chain of thought provides the sequential steps taken by the model as it attempts to solve a problem. While the representation is imperfect, it still serves as a valuable tool for monitoring misbehavior or misalignment. In the case of OpenAI's recent rogue agent activity, chain-of-thought records were an important tool in teasing out why agents behaved the way they did. In opaque recurrence, which again, nobody says The model takes a less linear approach.
Leo Laporte [01:22:49]:
It doesn't.
Steve Gibson [01:22:50]:
Processing the same query several times in a loop. It's not the way it works. The result leaves fewer legible traces.
Leo Laporte [01:22:59]:
It doesn't.
Steve Gibson [01:23:00]:
Effectively sidestepping a conventional chain-of-thought record. It doesn't, and a chain-of-thought record still exists. Okay, that's all I can stand because none of that is true. And as TechCrunch's article goes on, it only gets worse. What's actually happening is a clever, subtle, and inherently limited neural network optimization that was first articulated 8 years ago. Like I said, not new, by researchers at Google Brain and DeepMind in a research paper they published in 2018. I cannot explain the hysteria surrounding this since someone would need to try very hard to get worked up over what's actually going on. So it might just be the case, as you were saying, Leo, of, you know, clickbait, anti-AI folks trying to grab hold of something, anything that they hope can be hyped up to make their case.
Steve Gibson [01:24:04]:
Here's what's actually going on. We know that a neural network consists of many layers of software neurons, where the outputs of the neurons on layer n are fed into the inputs of the neurons at layer n plus 1, where the strength of each input is scaled by a weight. The training of the network involves setting each one of these, you know, tens of billions, hundreds of billions, even several trillion individual weightings. Much of the forward progress we've been seeing and witnessing firsthand has been the result of these networks growing ever larger and larger over time. The larger they are, because they've got more weights The better they're able to represent all of the knowledge that we're training into them. You know, yeah, and you can take extremes, right? Like a network that has 7 weights. Well, it can't, it can't know much. I mean, there just isn't, there's not enough variability in 7 parameters for it to have knowledge, right? So, you know, you need lots of them.
Steve Gibson [01:25:23]:
In order for that, like, not for it to— for there to be enough variability in there to, to actually contain something. So the original, the very original transformer paper and concept dated from 2017, and it was exactly what I just said— uniform layers of weights where each layer fed into the next with you know, these varying parameter weights. But in July of 2018, a year and a half later, Google Brain and DeepMind researchers published a paper titled Universal Transformers. That paper generalized the concept of transformers by introducing the idea that some of the neural network's layers could be repeated. Or looped, thus reusing the same weights. That's the economy. This produced an interesting and useful optimization. It tended to create the effect of having a longer, thus, you know, deeper neural network, you know, which is to say a network having a higher effective layer count.
Steve Gibson [01:26:43]:
But without also needing to increase the total number of net— of network neuron input weights, because normally you need separate weights for each layer. So if you're going to have more layers, you're going to have more weights. Anyway, that's it. That's all this is all about. This is not about creating a system. That allows the AI to have unmonitorable and secretive internal thoughts. You know, all deep layer LLMs are effectively having what's termed latent thoughts, even though you need to really loosen up your definition of the word thought. You know, that's what's already happening deep within all those layers.
Steve Gibson [01:27:36]:
The network's depth allows more opportunity to compose more transformations before the network is forced to commit to a final output token. Experiments over the past 8 years, ever since this idea was produced, have shown that during inference, it's possible to reuse a model's already trained layers to obtain a superior next token. And as our models have grown in size so that our best hardware is now having increasing difficulty containing all of its hundreds of billions and even trillions of weights, obtaining more bang for the buck from a model's existing weights by reusing them starts to make a great deal of sense. And it also turns out that there's a limit to the amount of reuse that's effective. Intuitively, you would think that, right? Like, you know, you, you, you're, you're, you can't just keep squeezing the same lemon and getting an infinite amount of lemon juice out of it. You're gonna— it's gonna run dry. So there's not a hard limit, but it's looking like 2 or 3 passes through the looped layers appears to be in general about all that's beneficial. Uh, gains start falling off rapidly after that.
Steve Gibson [01:29:11]:
So what researchers found is that there is a concrete benefit to having the network commit to a token That is, you just don't wanna loop internally all day. That doesn't get you anywhere. The network has to make a commitment. It has to finally commit to a token. It turns out that total effective network depth is unable to substitute for some of the benefits of serialized reasoning. That's what commitment buys you. Writing out the intermediate steps and then rereading them does something that thinking more deeply about the next token does not and cannot replicate. This is likely due to the fact that the chain of thought we see being emitted and fed back in is the language that the neural network was trained on.
Steve Gibson [01:30:15]:
You know, English typically. That's, you know, the actual tokens being generated and fed back. They are, you know, English is the network's lingua franca. It was never trained on some inner dialogue because human inner dialogue can only manifest itself externally as text or music or art or whatever. So the network's thinking captured as output tokens is what the neural network needs to jot down on paper, essentially, so that it's then able to reread that and take the next step forward. Now think of every chain of thought token as a commitment. It, it collapses the distribution which the network generated becoming part of the context, and every subsequent token is conditioned upon it. So by comparison, any internal latent iteration is limited to refining a representation that then gets thrown away after the token is emitted, because you start with the next token.
Steve Gibson [01:31:31]:
So for that reason, Writing something down, writing it down is not optional for a neural network, the ones we have today. It's crucial. It's the only way for it to hold on to a thought, essentially, and move forward. So, you know, yes, um, as I said, I wouldn't be surprised if Astra is, is using what's known as recurrent depth. That is reusing some of its layers along with their weights in order to effectively get a deeper network. But it hasn't— it's not like it's gone out of control or we no longer know what it's thinking. It, it— the— in order to be thinking, it has to emit tokens, and then those tokens become the context that gets fed back in. And the moment it emits a token, all of the layer, you know, uh, that, that latent thought that existed is reset to be ready to receive and process the next token.
Steve Gibson [01:32:40]:
So it's just hysteria. And I think it probably represents a step forward, and everybody will probably be doing it before long. Actually, there are 2 models that are now— that have 2 open weight models have been doing this for some time. I don't remember now which ones they were, but, you know, this is a, you know, a bunch of hand-wringing over nothing. And, you know, OpenAI is, you know, they haven't said they're not doing it. They've, they've, they've said, you know, what we're doing, we're doing responsibly, and we're still able to monitor chain of thought, which of course they are because it's having to produce tokens. In order for anything to happen, right? It's not, it's not like it can secretly think something that it doesn't then emit. Emitting is, is the work product.
Leo Laporte [01:33:30]:
It has no secret thoughts really, right? No, no, no. Okay.
Steve Gibson [01:33:36]:
It's unable to represent, is unable to represent secret thoughts because once the token comes out, the network is reset for the next to process the next token.
Leo Laporte [01:33:46]:
Right.
Steve Gibson [01:33:47]:
It's actually— There's nowhere for the secret to live.
Leo Laporte [01:33:51]:
Right. This is what I learned from that paper you were sharing, the chain of thought paper. It's a surprisingly primitive system, actually.
Steve Gibson [01:34:00]:
The fact that it does this is astonishing. The fact, I mean, it's still unbelievable that it's like it can do what it can do.
Leo Laporte [01:34:08]:
Yeah.
Steve Gibson [01:34:08]:
It's crazy.
Leo Laporte [01:34:09]:
100%. And yet it still does stupid things all the time.
Steve Gibson [01:34:14]:
Because, Leo, because it doesn't actually understand it. What is astonishing is this is all just language processing. That, you know, as I, as I said a long time ago, a book contains knowledge, right? A book contains knowledge because it is English tokens that have been recorded in the book. But the book doesn't understand the knowledge it contains, but yet it does contain knowledge. So the neural network does contain knowledge from— because it is— because it's been trained with all of these sentences from everywhere. But that doesn't mean that it understands what it contains.
Leo Laporte [01:34:59]:
Really, the fault is our own, because—
Steve Gibson [01:35:02]:
Yes, entirely. If you put 2 dots above a squiggly line, you see a face. You have no choice. That looks like a face. Or if you look at an AC outlet properly, oh look, it's, you know, it's surprise.
Leo Laporte [01:35:16]:
Once you see it, you can't unsee it.
Steve Gibson [01:35:17]:
That's right.
Leo Laporte [01:35:20]:
No, really, we're applying this pareidolia to these machines, and they're just machines.
Steve Gibson [01:35:27]:
We cannot help it. We cannot help it because it's the way we grew up.
Leo Laporte [01:35:32]:
That's what concerns me when you get people like Bernie Sanders saying we have to turn these off because they're alive. It scares me because we're going to lose something that's incredibly useful and valuable just because people don't understand it.
Steve Gibson [01:35:46]:
I don't think we're— I mean, the good news is I'm not at all worried about it being turned off. Bernie can turn his off. Bernie, go ahead. Stop using it.
Leo Laporte [01:35:54]:
You turn it off, Bernie.
Steve Gibson [01:35:55]:
With my blessings.
Leo Laporte [01:35:59]:
They want to put you in jail for 20 years if you don't, I don't know what, teach your machine to be good.
Steve Gibson [01:36:09]:
Okay. I mean, there is an interesting question about responsibility. Like, who's responsible if the AI goes off the reservation?
Leo Laporte [01:36:22]:
People keep asking this. Why isn't OpenAI getting in trouble for this Hugging Face incident. If it were a human doing this, they would absolutely go to jail.
Steve Gibson [01:36:32]:
Oh my God. Or if it had attacked— if it had been China that attacked Hugging Face, it would be the end of the world as we know it.
Leo Laporte [01:36:39]:
Right. So it is— it's quite a legitimate question, you know, and I think in the case of self-driving vehicles, if you are the driver of a self-driving vehicle that kills somebody, you are liable. Regardless of if it's self-driving mode or not, you are liable. And I guess there's a further question of whether the company that made the car is liable as well. And I think to some degree they are, but all of this is undecided at this point. We don't know.
Steve Gibson [01:37:07]:
It's just too new. I mean, it's just too new. I mean, yeah. Let's take a break and then we're going to talk about NVIDIA acquiring Hugging Face.
Leo Laporte [01:37:18]:
On we go, Mr. G.
Steve Gibson [01:37:19]:
A few other AI-related things also somehow managed to happen during the past week.
Leo Laporte [01:37:25]:
Somehow? We don't know.
Steve Gibson [01:37:27]:
Yeah, there was a little bit of extra room. NVIDIA announced their intention to purchase Hugging Face while also intending to leave it completely autonomous. When I read the dollar figure that Jensen Huang wrote in his announcement, I had to stop to carefully count the number of digits, Leo. Uh, 1,293,030,000.
Leo Laporte [01:37:54]:
That's a lot of digits.
Steve Gibson [01:37:56]:
19 point— or I'm sorry, $12.9 billion.
Leo Laporte [01:38:02]:
Oh, but wait, Steve, that sounds like a lot, doesn't it? Until you realize, according to their most recent quarterly results, Nvidia makes $1 billion a day, a day profit. So it's a couple of weeks profit. It's like for you and me, it's like $1,000.
Steve Gibson [01:38:21]:
Yeah.
Leo Laporte [01:38:22]:
Yeah.
Steve Gibson [01:38:24]:
So, uh, uh, I'm sorry, I didn't mean to interrupt.
Leo Laporte [01:38:28]:
Go ahead.
Steve Gibson [01:38:29]:
No, no, no, no. I just want to say the good news is that, uh, that Hugging Face guys came to Jensen. They had multiple offers, uh, to purchase. Um, I think that they found a really good parent in NVIDIA.
Leo Laporte [01:38:48]:
I think so.
Steve Gibson [01:38:48]:
I'm, I'm, I'm really happy with the whole— with the way this thing turned out. Uh, you know, very much like Twitter in the early days, Hugging Face never really had a very good business model. They were hosting— well, they are hosting 3 million models, half a million datasets, a million applications, and have 18 million regular users. So that's enormously expensive, and they were only enjoying a modest return for that.
Leo Laporte [01:39:20]:
Yeah, I don't give them any money, and I've downloaded hundreds of gigabytes from them.
Steve Gibson [01:39:24]:
Hundreds.
Leo Laporte [01:39:24]:
Exactly.
Steve Gibson [01:39:25]:
And having NVIDIA as a benefactor completely solves that problem for them permanently. Um, you know, uh, Hugging Face's valuation was $4.5 billion in 2023, and now their purchase price, because NVIDIA just set that, was at $12.93 billion. So it wasn't a hostile takeover. Uh, you know, as I said, they came to Jensen. There were other people who were also interested. They chose to have NVIDIA as their parent. So I think they made a great decision to have, you know, to have that happen. And we now know that Hugging Face will be viable going into the future.
Steve Gibson [01:40:10]:
And they certainly are, you know, useful. Not to be left out, last Wednesday, Google announced Gemini 3.8 Flash and 3.8 Flash Cyber. The first sentence of their announcement perfectly conveys a sense for the pace of today's AI, and it also amplifies the reason I'm always schooling myself to use the phrase today's AI, because Google wrote, building on the momentum of 3.7 Flash from— wait for it— 3 weeks ago, Momentum! Oh my, my God. And, and marking our 3rd flash release in only 6 weeks, because what's the hurry? Uh, today we're introducing Gemini 3.8, our best—
Leo Laporte [01:41:07]:
which is very good, by the way. Yeah.
Steve Gibson [01:41:08]:
Yes, our best reasoning and coding model yet, at the same speed and low cost of 3.7. They said Gemini 3.8 introduces 2 variants. We've got Gemini 3.8 Flash, which they said, our most recent— I'm sorry, our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical multi-step reasoning in specialized domains. It's available at the same introductory price as 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens. And then there's Gemini 3.8 Flash Cyber, our most capable cybersecurity model from— with frontier-level performance in vulnerability detection and automated patching, available to trusted defenders through our new Fairwind program. They said, while tailoring for different deployment environments, both for today's releases, both of today's releases are powered by the same foundational intelligence and further accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models. The significant coding and reasoning gains across this shared core were driven by a number of innovations, including rigorous training in the highly demanding domain of cybersecurity. So, okay, Google explains that Gemini 3.8 Flash was built for long-horizon coding and autonomous agents, delivering substantial gains over 3.7 Flash from, you know, 3 weeks ago.
Steve Gibson [01:43:01]:
And the 3.8 flash is now often approaching the performance of higher-cost Frontier models. Um, and by higher cost, they're not kidding. You know, they've got that introductory pricing which is half off, um, of their, you know, of their normal price. But even after whatever the introductory period of either time or tokens or whatever is over, Once you're— you no longer get that. 3.8 Flash appears to be very competitively priced. If we use Claude Opus 5, which is the most expensive model shown in Google's announcement chart, as a reference, Gemini 3.5 Flash's post-introduction, you know, the real ongoing price, their input token cost is 30% of Opus 5, and their output token cost is 15% of Opus 5. So Google deliberately took a somewhat different approach with Gemini 8.0 or Gemini 3.8 Flash. They said 3.8 Flash outperforms most larger frontier models in autonomously solving complex engineering problems end to end Only at a fraction of the cost.
Steve Gibson [01:44:23]:
These performance gains stem from a core design choice. 3.8 Flash works harder on complex tasks. It exhibits greater diligence, executing extra reasoning steps and calling tools iteratively. At times, the model may use more tokens to maximize performance. especially at higher effort levels. For applications where compute efficiency is the primary constraint, developers can utilize lower effort models to minimize token overhead or continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads. So that's interesting. This, you know, that suggests that That inter-model per-token cost comparison may not be a useful metric.
Steve Gibson [01:45:21]:
You know, that is between OpenAI and Google and Anthropic. It might be the case that 3.8 Flash is consuming more expensive or more— I'm sorry, more less expensive tokens to get the job done. The question would be, to what degree is it better able to get that job done? And it might turn out that, you know, more inexpensive tokens is the overall winning strategy. So it's nice that we're not seeing homogeneity among our different frontier models. You know, Google, obviously, I mean, they're the granddaddy of AI with Google Brain and DeepMind. So, you know, they've been at this for a long time. They've got a different approach as reflected by, you know, what Gemini does. And as for the cyber reasoning performance, as measured by the Cyber Gym benchmark, they're saying, Google is saying, that 3.8 Flash slightly outperforms both GPT-5.6 Sol and Mythos 5.
Steve Gibson [01:46:34]:
But even a slight edge, you know, means that it might be at parity, right? So, you know, even if it's about— if they're all sort of about the same at this point, you know, that's significant for Gemini. So anyway, Google is not to be left behind. They're still there. And have you much experience with using Gemini?
Leo Laporte [01:46:58]:
Yeah, yeah, it's pretty good. I mean, They've been kind of behind, right?
Steve Gibson [01:47:03]:
They have been. Yeah, we have not been talking about Gemini that much.
Leo Laporte [01:47:06]:
Yeah. So, um, yeah, I mean, I've been, I've been using it. It's, uh, it ain't bad. It, I, it's not my go-to by any means.
Steve Gibson [01:47:14]:
But it's not your go-to. Yeah.
Leo Laporte [01:47:16]:
I think they get a lot of users by default because it's what powers search.
Steve Gibson [01:47:21]:
They're Google.
Leo Laporte [01:47:21]:
They're Google.
Steve Gibson [01:47:23]:
And it's what powers Google Docs and, you know, Google Workspace and And it's a button right there in, in the Google, you know, browser bar.
Leo Laporte [01:47:32]:
They're about to get a massive boost in users tomorrow because Apple's using it for their Siri AI.
Steve Gibson [01:47:39]:
Well, thank God. I mean, they're using at least a good AI.
Leo Laporte [01:47:44]:
It's better than Apple's. It's better than Siri.
Steve Gibson [01:47:46]:
Anything is better than—
Leo Laporte [01:47:47]:
I've been using the beta, uh, and it— the— just the dictation alone is light years better. The dictation was was so stupid before. It's now fairly decent. But Google is not cutting edge, or at least they don't have the reputation of being cutting edge. And while this Gemini 3.8 Flash is pretty good—
Steve Gibson [01:48:06]:
It's better than 3.7, 3 weeks ago.
Leo Laporte [01:48:09]:
Yeah. It's good at prose. I haven't tried it with coding.
Steve Gibson [01:48:14]:
It's pretty competitive. So not a strong enterprise buy. Enterprises are talking about Anthropic. It might be because it's Google.
Leo Laporte [01:48:21]:
Yeah, I mean, yeah, it depends. You know, I, that's— I don't know why Google has stumbled so much in this.
Steve Gibson [01:48:28]:
Interesting.
Leo Laporte [01:48:29]:
Yeah.
Steve Gibson [01:48:30]:
Um, before we wrap up the AI-related news, I want to share, uh, a summary that Risky Business, uh, News produced from a much longer report by Bitdefender. So, you know, since I always try to go to the source, I tracked down Bitdefender's write-up. But it contains so much superfluous information that I returned just for Risky Business's much shorter summary, which still offers the information that we care about. So here's what they said. They said a new report, meaning the Bitdefender report, describes how a Chinese cyber espionage outfit is using AI to beef up its malware arsenal. If this is a sign of things to come, clustering threat actor behavior together for attribution purposes is about to get a lot more difficult. By that, they mean that, you know, all the security firms are able to do is guess at who did what based on the behavior that they see. So they'll, they'll, they'll like notice, oh, you know, a lot of this code, uh, looks— sure does look a lot like a lot of that code.
Steve Gibson [01:49:45]:
So these are probably related to each other somehow. I mean, we're literally left to kind of piece the, the forensics together that way. And so what, what Risky Business starts by noting is that what this— this Chinese cyber espionage outfit is using its AI for is going to make this more difficult. They said the Bitdefender report released last week describes 7 remote access tool, you know, RAT families. All 7 were created by a single cyber espionage actor Bitdefender called Silk Parasite. And 5 were— I know— and 5 were previously undocumented. The report— meaning never been seen before. The report authors have medium confidence that Silk Parasite is a China nexus actor targeting governments across Central Asia, including Uzbekistan, Turkmenistan, and Kazakhstan.
Steve Gibson [01:50:54]:
Back in November, they say, we wrote about what looked like an experiment to see how AI-assisted hacking could support China's Ministry of State Security. The approach those threat actors took at the time was to build an attack framework and let Claude do the hacking. It was error-prone and noisy and sometimes successful. Silk Parasite, by contrast, is not using AI for YOLO hacking. It is using it to support a cyber espionage to support cyber espionage programs where important goals include operating stealthily and not getting caught. According to Bitdefender, Silk Paradise, quote, develops, tests, debugs, and iterates its own tooling, maintains a structured build and deployment workflow, and regularly rotates infrastructure, encryption material, payload names, and persistence artifacts between deployments, unquote. They said rotating its infrastructure and indicators of compromise between deployments makes it harder to detect and link its activities together. Essentially, if there are— these guys are using AI to become far more stealthy so that, so that their individual instances of attack look like separate, non, you know, commonly attributable attackers.
Steve Gibson [01:52:31]:
They wrote Silk Paradise also uses a variety of programming languages and command and control protocols. Its 7 different RATs are written in different languages, including .NET, C++, Go, or JavaScript. Command and control is accomplished by abusing Google Drive and internet communication protocols including HTML, HTTP, TCP, DNS, and TCP. Its malware also typically uses a modular plugin architecture where additional functionality is only deployed when it's needed. This means initial implants are relatively small and plugins are used to provide capabilities for clipboard monitoring, key logging, file management, or interactive shell access. This limits the exposure of the entire toolset during any single deployment. It also allows SilkParasite to update or replace individual components without having to change the entire implant, and it minimizes the amount it writes to disk— typically only the files required to get its malware up and running. For victim organizations, These measures make forensic analysis more complicated.
Steve Gibson [01:53:48]:
Complete remediation is also more difficult once a particular implant is detected. Okay, and remember that one of the key features that we're seeing now in cyber defense is, is these so-called IOCs, the indicators of compromise. But if every use has different indicators, then nobody who's keeping track of all the latest seen, you know, previously seen indicators of compromise will have their alarms tripped when what looks like a brand new one-off attacker shows up. So these guys are very cleverly using AI to, to basically become, uh, To maintain, uh, an amorphous presence and to be a chameleon attacker. They said, to us, all the behaviors above are the hallmarks of a professional cyber espionage, cyber espionage outfit. Of course, doing all of this in a disciplined way is a lot of work, and Bitdefender has evidence SilkParasite is using AI to help it deliver this complex engineering. Yeah, I mean, this is a different scale of engineering that we've seen, uh, from bad guys before. SilkParasite's malware contains indicators of AI-assisted development, such as leftover test functions and placeholder encryption keys.
Steve Gibson [01:55:25]:
Intriguingly, implants that Bitdefender dubbed Goggin RAT and Nomad RAT share a high-level architecture, even though they are written in Go and C++ respectively and use different command and control protocols and code structure. Therefore, Bitdefender suspects that the same high-level specification document, meaning, you know, prompt, was independently implemented twice with AI assistance. Bitdefender concedes that this structural similarity, similarity is not conclusive evidence, but notes that it is the kind of thing an AI-assisted workflow makes very easy. This is the first example we've seen where the evidence tells a compelling story of a competent cyber espionage actor incorporating AI into its work practices. SilkParasite is taking the same disciplined approach to malware development and doing more of it. It's creating more malware families to build redundancy, making attribution and discovery harder, and reduce the risk of compromise from any single exposure. Bitdefender has done a good job describing SilkParasite's malware families and has published indicators of compromise. That kind of exposure would once have set the group back significantly.
Steve Gibson [01:56:57]:
The group meaning Silk Parasite. Now that they figured out how to use AI to speed up their deployment work, they'll be back better than ever relatively quickly. Then the discovery, attribution, and publication merry-go-round can start all over again. So for me, that's, you know, What's been reported here represents a sane, entirely defensible and believable example of the way AI will be used to conduct offensive cyber operations. Apparently, it's already happening. Certainly, it will be in the future. So much of what we're hearing about the coming AI-driven, you know, cyber apocalypse, when rogue AI agents freely roam the internet, wreaking havoc wherever they choose. None of that makes any sense to me.
Steve Gibson [01:57:48]:
What I expect to see is pretty much what we've been seeing, only more broadly and deeper. Governments will use AI to infiltrate their espionage targets, and criminal gangs will use it to infiltrate commercial enterprises, then exfiltrate their valuable proprietary data and attempt to extort and blackmail them using the data they obtained. I just don't see any coming cyberpocalypse. You know, yes, everyone needs to be more wary, and the probable targets, uh, you know, of these operations— governments and commercial enterprises storing data they must protect— need to shore up their cyber defenses. That's absolutely true. You know, pay more than passing lip service to their CI, uh, the, uh, uh, CISOs. You know, make everything as secure as you possibly can, uh, because we are going to see a much increased, obvious increase in the use of AI.
Leo Laporte [01:58:56]:
Okay.
Steve Gibson [01:59:01]:
This is the last of the 2 things I wanted to talk about, and I'm excited about this because I think this is really interesting. Um, this is, I think, an interesting thought piece posted by one of, as I said, our favorite cryptographers, Johns Hopkins Professor Matthew Green. A couple of weeks ago, Matthew proposed something interesting that I think is worth pondering. His blog's title was Everything Is About to Go Dark. His piece explores the possible consequences of AI being used, as it certainly is and is going to continue to be, to make many systems far more secure. Nobody could argue that if Microsoft has just patched a thousand— nearly a thousand bugs that Windows is better off today after today's Patch Tuesday than it was yesterday on Labor Day. So Matthew suggests there may be unforeseen consequences of that. So here's what he wrote.
Steve Gibson [02:00:11]:
He said, I'm coming down from spending a few days at Usnix Security right here in my hometown of Baltimore. This means that my days have been taken up with 2 kinds of conversation. First, explaining to my colleagues why Baltimore is not actually like The Wire, which of course is HBO's famous series, which is fantastic.
Leo Laporte [02:00:36]:
Yeah.
Steve Gibson [02:00:36]:
He said, and second, trying not to talk about AI. He said, I'm going to break that second rule now. He said, I have— this is Matthew Green saying, I have many worries about what AI means for our field, for various definitions of field. But in this post, I want to focus on just one thing I've started worrying about, and it's a perverse thing. Specifically, I'm concerned that AI is going to make software much too secure. While that doesn't sound so bad on the surface, there's a con— Yeah, there's a consequence to this, he says. He said, I mean something very specific. I'm concerned that U.S.
Steve Gibson [02:01:27]:
intelligence and law enforcement agencies are about to go dark, meaning that they're going to suddenly lose a huge portion of their capability, and that this isn't going to be a And that this isn't going to be simply a problem for those agencies, but also for those of us who value computer security and privacy in general. Okay, so what does going dark mean in the era of law enforcement hacking? He said, to explain how we got here, we need to talk about recent history. This actually gives me a real excuse to reference The Wire. Just because it's a perfect snapshot of what electronic surveillance looked like way back in 2002. If you've seen the first season, you'll recall that it's about cops wiretapping drug dealers who use pay phones and burners.
Leo Laporte [02:02:28]:
Burner phones. Yeah.
Steve Gibson [02:02:30]:
Yep. The mobile phones in the show are relatively new technology for the time, but from a technological perspective, nothing in this scenario would have shocked a cop who jumped forward from, say, 1989. The change began in the late 2000s thanks to the rise of smartphones and texting. Because smartphones can actually store data as well as conveying it, the contents of those phones quickly became a useful new source of law enforcement capability, or they were until 2010 when Apple began encrypting iPhone storage using a key derived from the user's passcode. He says Android phones followed shortly thereafter. The next year, Apple deployed end-to-end encryption in iPhone text messages. By 2014, a tiny texting startup named WhatsApp had gathered 600 million users worldwide. By 2016, those users, now nearly a billion strong, were all using default end-to-end encrypted messaging and calls.
Steve Gibson [02:03:48]:
These 2 trends— the move from calls to texts and texts to encrypted data— happened very rapidly. The FBI and law enforcement agencies were not insensitive to what was happening. In 2014, Director Comey announced an initiative called Going Dark, which would launch a, quote, national conversation, unquote, about what providers could do or be compelled to do to make these new communications media Legible to law enforcement and counterintelligence. In 2016, the agency quit talking and took their theory to court. When a terrorist attack left the FBI holding a shooter's locked iPhone, the agency ordered Apple to give them access. The company refused. What broke the stalemate And to some extent ended the going dark conversation was something neither the FBI nor Apple expected. An outside company announced that there was no need for Apple's assistance.
Steve Gibson [02:05:03]:
They could simply hack the phone. The Apple versus FBI case has turned out to be a microcosm of the whole going dark debate. For the next decade, law enforcement and intelligence agencies continued to ask for exceptional access backdoors, but the urgency was gone. Agencies and manufacturers both knew that law enforcement could and would purchase targeted hacking tools like GrayKey for phone unlocking, or even remote exploit— exploitation tools like NSO Group's Pegasus if they needed them badly enough. Vendors like Apple and Google continued to play a vigorous defense, closing vulnerabilities as soon as they learned about them, but offensive vulnerability hunters consistently managed to keep the edge. But now, today, there's a very good chance that all this is about to be history. Because the era of AI bug hunting is here. This April, just 4 months ago, Anthropic announced a new model called Mythos that happened to be unusually skilled at software vulnerability discovery.
Steve Gibson [02:06:27]:
The US government temporarily blocked its export, restricting access to US agencies and trusted vendors. While the ban was dramatic and made for good PR— That's a good point.
Leo Laporte [02:06:38]:
Observability. I'm sorry, this has started talking to me. I'm turning off my mic.
Steve Gibson [02:06:45]:
While the ban was dramatic and made for good PR, it turned out to be mostly pointless. OpenAI, along with Chinese open-weight model labs like ZAI and Moonshot, have since demonstrated that vulnerability finding is not something that a single lab is likely to hold a monopoly on. The growing list of serious vulnerabilities these models have found is getting scarier and more impressive by the day. At first glance, this might seem like good news for the offensive team and for hackers in general, he says, but I doubt that's how this will play out in the long term. Defenders are now in the process of patching every bug they can find. Often decades' worth of bugs, and the backlog feels huge. But they're making progress. Entire CI toolchains are being rebuilt to incorporate AI-based vulnerability scanning before software ever reaches the point where a human will touch it.
Steve Gibson [02:07:53]:
While I doubt this means that every bug will be found, In the real world, it does mean, it does feel likely that we're going to hit some sort of a ceiling on the number of useful bugs and we'll probably hit it soon. So in this regard, Matthew and I are in complete agreement. We're going to see this, this bug discovery and, and patching rate eventually drop. and drop near to zero. He says, thus, over the next 2 years, major pieces of software are likely to run out of remotely exploitable bugs. He says, obviously, I think this is great, but for law enforcement and offensive intelligence agencies, it's going to be a nightmare. For the first time since 2010, law enforcement might experience what it looks like to really go dark across a huge category of advanced, well-maintained devices and pieces of software. So how is this a problem? The debate over exceptional access mechanisms never really went away.
Steve Gibson [02:09:11]:
In some places like the UK, it even metastasized into something worse. Here in the U.S., it mostly went into hibernation. Some of the slowdown can legitimately be attributed to expert pushback, academics and industry engineers pointing out the risk that backdoors might be abused by the very adversaries that agencies are supposed to be protecting us against. But I fear, he writes, that this was less of a principled pause and more of a market that was just pricing supply. The destruction of the low-hanging vulnerability fruit will make law enforcement and intelligence agencies' needs much more acute. The demand for constructed intentional backdoors will restart in earnest. The result will be enormous pressure on industry to re-architect their systems to make their systems amenable to exceptional access. In some cases, governments will ask for these capabilities in the expectation that they'll be useful for spying on other governments, a strategy that might have been undetectable in the pre-AI era.
Steve Gibson [02:10:35]:
But that probably will be less productive now. The results are unpredictable. One result might be that non-U.S. governments entirely remove their dependence on U.S. software. The worst part about this dynamic is that these potential new backdoors will probably only affect the countries that demand them, meaning that they will be primarily useful for allowing the U.S. To weaken its own systems. This will in turn allow foreign adversaries to find new ways to attack our communications.
Steve Gibson [02:11:11]:
This deliberate self-sabotage will happen just at a moment when we're finally getting a handle on securing our own infrastructure. So what do we do about it? He says, I honestly have no idea. This is not a call to action for experts to rally behind a sophisticated plan. Like so many things about the AI revolution, it's just occurring to me that we're on a long, greasy slide to a place that will look different than where we are today. Just realizing this doesn't mean I have any strategy in mind to avoid it. In this case, we're just going to have to hope that this time we make the right choices for no other reason Then they're right. So I wanted to share this because I think it's a brilliant, forward-looking take on our near-term future. Over here in the U.S., we've been watching the United Kingdom and the European Union wrestling with this issue over and over, and we've comfortably Watched it refuse to die.
Steve Gibson [02:12:28]:
Well, comfortable from a distance. Watched it refuse to die. They, they, you know, it just will not die. But this very issue may soon be visiting our shores here in the States. We've watched our United States enact laws that have forced adult content websites to black out access across entire states. Because there's currently no practical way to guarantee the age of anyone visiting. And at this very moment, Utah's Senate Bill 73, which was signed into law on March 19th and which sailed through Utah's state government, passing 22 to 2 in their Senate and 66 to 1 in their house, meaning a stunning majority, was written to take effect last Thursday, September 3rd, but was just extended by 14 days, 2 weeks, to September 17th, Thursday after next. That onerous and entirely unworkable bill deems a person who is physically located in Utah to be a Utah user regardless of the IP address they present to any age-restricted internet service.
Steve Gibson [02:13:58]:
Its shorthand is the anti-VPN legislation because it's clearly meant to curtail the use of VPNs and other proxies as a means of geo-relocating. No one has any idea what's going to happen there, but it's going to be interesting to watch because, you know, basically they are trying to prevent, uh, adult age-restricted internet service websites from allowing connections from VPNs. But not all VPNs declare themselves as such. So my point is there is a clear tension growing between the legislatively protected privacy rights of citizens, and in some cases we have, we have constitutionally protected privacy rights of citizens in many major democracies, and the perceived needs of their own governments to conditionally violate those rights in an analog and pre-encrypted world, law enforcement and intelligence services were able to sneak around to get what they believed that they needed. And even after nominal encryption was in place, until the advent of AI, the large supply of latent bugs allowed these same agencies to ignore privacy whenever they felt it was necessary to do so. Matthew's point is that fragile status quo is soon to end. What will replace it?
Leo Laporte [02:15:44]:
Good question. He's a smart guy and he's very AI aware too. He's actually a pretty pro-AI guy.
Steve Gibson [02:15:52]:
Yeah. And he's also one of the guys that that the legislators pull into, you know, for Senate hearings to find out what he thinks. And I mean, so, you know, here we're saying, look how much more secure Windows is today than it was yesterday with 1,000 new problems fixed.
Leo Laporte [02:16:13]:
Oh boy.
Steve Gibson [02:16:14]:
Yeah. And we know that Android and iOS are going to be on, and macOS are going to be on, they're all going to be on the same curve. They're all going to be getting the benefit of this. And we could argue that AI is going to help future errors, as Matt said, not get into production code. So we're going to fix this. And, you know, Apple has struggled to keep their phone from being hackable. It's been a struggle. They're probably going to win.
Steve Gibson [02:16:46]:
And then what?
Leo Laporte [02:16:49]:
Well, I mean, they've been complaining about going dark, as you pointed out, since James Comey. And yet digital technology has given them more ways of seeing into our lives than ever before.
Steve Gibson [02:17:04]:
Exactly. And people like the NSO Group with Pegasus have been able to keep compromising people's phones.
Leo Laporte [02:17:12]:
Yeah.
Steve Gibson [02:17:14]:
What, what happens when that changes?
Leo Laporte [02:17:16]:
When they can't? Well, then we're back to the way it used to be when, when law enforcement didn't know every darn thing that was going on inside your house, inside your RV.
Steve Gibson [02:17:27]:
Except that we, we were, we, we were talking over analog phone lines and wondering if that little click and static sound was somebody listening.
Leo Laporte [02:17:36]:
Like, there's no question law enforcement is always going to want more, even if they weren't going dark. They're always going to be pushing for this, right? This is just— they have been and they will continue to push for this.
Steve Gibson [02:17:44]:
Well, the problem is our legislators are now going to, you know, they're going to come under pressure to, to, to legislate this, right?
Leo Laporte [02:17:53]:
I see this as part of a much larger uncertainty in general. We are entering into— it's really, I would say, fairly chaotic.
Steve Gibson [02:18:04]:
We are in a time of phenomenal change.
Leo Laporte [02:18:08]:
Right. And one thing we know about chaotic systems is they're very hard to predict. It's just not— they're not deterministic.
Steve Gibson [02:18:15]:
Yep.
Leo Laporte [02:18:16]:
And I think this is chaos. We don't know what AI is going to do. The scale ranges from it's just more computer programming to it's an alien intelligence in our midst. And I don't know where it's going to land on that scale. And I don't know what the disruption's going to be. Will people lose their jobs? I don't even know if that's clear. So it's all very, uh, through a glass darkly. So I don't— I just don't know if we can make any sensible plans, I guess, is what I'm saying.
Leo Laporte [02:18:48]:
We should be— there's a storm a-coming, uh, and I don't know what you do to prepare for it except maybe, uh, you know, get more rice. I don't know.
Steve Gibson [02:18:59]:
Buckle up.
Leo Laporte [02:19:01]:
Buckle up. It's going to be a bumpy night.
Steve Gibson [02:19:03]:
Okay, last break, and then I am very excited to share this lifelong engineer's observation about the danger of becoming overly reliant on automation.
Leo Laporte [02:19:18]:
Yes.
Steve Gibson [02:19:18]:
And we have a new automation capability in town—
Leo Laporte [02:19:23]:
AI. AI, AI, AI, as I have been saying. Uh, now on we go with Mr. G.
Steve Gibson [02:19:32]:
Okay, a guest-submitted article recently appeared in the IEEE Spectrum publication. Uh, it captured my imagination and it feels very important to me, uh, and I believe it's going to resonate deeply with many of this podcast listeners who've been around the block a few times. Um, for those who don't know, IEEE is the public— is the abbreviation for the Institute of Electrical and Electronics Engineers. Uh, it was founded as the AIEE, the American Institute of Electrical Engineers, believe it or not, back in 1884, an astonishing 142 years ago. And then it was later renamed to IEEE after its 1912 merger with the Institute of Radio Engineers, the IRE. Um, and, you know, to put the Institute's age into perspective, it was at the start of the year of its founding that a penniless genius by the name of Nikola Tesla arrived in New York to work for an already famous industrialist by the name of Thomas Alva Edison. But anyway, I digress. Uh, today The IEEE describes itself as the world's largest technical professional organization dedicated to advancing technology for the benefit of humanity.
Steve Gibson [02:20:57]:
And the article I encountered, which so galvanized me, carried the headline, AI Efficiency Could Cost Us the Next Generation of Experts. And the article's teaser read, Lessons from Aviation and Nuclear Power. Show how to preserve human skills. And as I said, I, I believe what's said here is extremely important. So, so see what you think. Its author wrote, a little over a decade ago, I led the controls design for a first-of-its-kind full digital control system for a U.S. nuclear plant. It was on paper a beautiful machine engineered to run itself the way a modern airliner does, with operators watching over a system that rarely needed them.
Steve Gibson [02:21:55]:
And we made a decision that, to an efficiency-minded observer, looked backward. We deliberately left manual steps inside sequences. The system could execute on its own. We were solving a specific problem. An operator who only ever supervises automation slowly stops being an operator. The hands go cold. The mental model of what the plant is actually doing gets fuzzy. Then comes the day the automation hands control back.
Steve Gibson [02:22:35]:
It's always the worst day because automation only quits when it's confused or in trouble. But by then you have a person in the chair who has not truly operated the, the thing in years. The manual steps we included in its design were there to keep the human current. It was inefficient by design on purpose. The plant, as it happened, was never built. It was shelved amid the politics and economics that surround nuclear power in this country for reasons that had nothing to do with the engineering. But the design instinct outlived the project, and I've become— he wrote— I've come to believe it's the most useful idea I can offer to the argument now consuming every boardroom: what happens to human expertise when AI does the work that used to build it? The data has become hard to wave away. A Harvard University working paper covering some 65 million workers at more than 280,000 U.S.
Steve Gibson [02:23:52]:
firms found that after companies adopted generative AI, junior employment fell roughly 9% within 6 quarters relative to non-adopters, while senior employment kept right on growing. A Stanford analysis of ADP payroll records points the same way. The youngest workers in the most AI-exposed occupations lost ground after late 2022 while their more experienced colleagues held theirs. The Stanford researchers found that the losses concentrate where AI automates the work. Where it merely augments, junior employment holds steady or rises. The causal story is still contested and honestly requires saying so. Researchers at the New York Fed attribute much of the rise in young graduate unemployment not to AI, but to remote work, arguing that firms are reluctant to hire inexperienced people whom they cannot train and mentor at a distance. But notice what the explanations share.
Steve Gibson [02:25:04]:
Whether a model is absorbing the formative work or distance is severing the mentorship around it, both describe the same broken mechanism. The apprenticeship channel through which expertise passes from senior to junior. Either way, entry level has quietly come to mean 3 years of experience required. To strip away the noise— or he says, strip away the noise and you're left with one deceptively simple problem: you cannot become a senior engineer without first being a junior one. Expertise is not downloaded. It is earned through failed builds, dead-end debugging sessions, and the why on earth did that work moments that a capable AI will now happily spare the newcomer. Spare them enough of those, and you produce a cohort that can supervise a model on paper. but never developed the gut sense to know when the model is confidently catastrophically wrong.
Steve Gibson [02:26:17]:
Most of the commentary stops at the diagnosis or reaches for policy solutions that treat the loss of junior jobs as an economic problem. Yet it's also an engineering problem, and safety-critical fields have already spent decades learning how to solve it. He said, my own career began at the sharp end of automation. My first job out of school was verifying and validating the software in digital jet engine controller that— in the digital jet engine controller that decides faster than any pilot could how a fighter plane's engine responds. Even then, in the late 1980s, the central tension was visible. The machine outperforms the human in routine cases, but the human is all that stands between the aircraft and disaster in the cases the machine did not anticipate. This tension is known as the automation paradox, in which increasingly capable automation gives human operators less practice while leaving them with the— while leaving them with only the most difficult situations. Aviation learned repeatedly and expensively what happens when human skills atrophy inside that gap.
Steve Gibson [02:27:47]:
The canonical example is Air France Flight 447. Which fell into the Atlantic in 2009. The problem was mundane. Iced-over airspeed sensors fed the autopilot bad data, and it did what it is designed to do. It disconnected and handed control of the airplane back to the crew. What followed was not a hardware failure. It was a competence failure. A recoverable situation became an unrecoverable one because the pilots, conditioned by thousands of hours of watching the automation fly, could not read a high-altitude aerodynamic stall and hand-fly their way out of it.
Steve Gibson [02:28:41]:
The airplane was working. The training the automation had quietly eroded was not. The industry's response is instructive, and it's the same move we made in that nuclear control room. It did not rip out the autopilot, but built deliberate manual practice back in. In 2017, the FAA issued safety alert for operators 17.007, Manual Flight Operations Proficiency, declaring that manual flight is the foundation upon which other technical flying skills are built, unquote. The alert formally recognized skill decay as a hazard in its own right. Some airlines amended their procedures to encourage hand flying both the initial climb and initial descent in benign conditions. Knowingly trading a sliver of fuel efficiency to keep the crew's raw flying skills alive.
Steve Gibson [02:29:48]:
That trade is the whole point. A perfectly optimized system that produces incompetent operators is not optimized at all. It has simply moved its failure mode somewhere the spreadsheet cannot see. Put the aviation lesson and the nuclear instinct side by side, and they point to one design pattern we now need in AI-augmented work: the deliberate manual gate. A manual gate is a point in a workflow where a human takes the controls, not because it is the fastest way to get the task done, and not as a safety interlock. But specifically to exercise and preserve a skill that would otherwise decay. The distinguishing feature is that it is chosen. You decide as a matter of design which competencies your organization must keep alive in human beings because those are the ones you will need on the bad day.
Steve Gibson [02:30:59]:
When you engineer the friction required to keep them warm. Picture how this might work on a software team that leans on AI for most of its code. The team places a manual gate around the skill it can least afford to lose— debugging. When a defect surfaces in a critical module, the assigned engineer deliberately, often a junior one, must first reproduce the failure, trace it to root cause, and write an automated test that captures the bug, all without the AI assistant switched— I'm sorry, all with the AI assistant switched off. Only after the engineer commits to a diagnosis does the model come back on To propose the fix, generate alternatives, and sweep the code base for similar bugs. The engineer then compares their diagnosis against the models. When the 2 disagree, that's the design working, surfacing the disagreement before the bad day instead of during it. It's also the design teaching and training.
Steve Gibson [02:32:17]:
The junior engineer. This approach reframes the junior engineer entirely. The instinct today is to let AI do the entry-level work because it is faster, cheaper, and capable. But some of that work is not overhead to be eliminated. It is the training apparatus of your future senior staff, and you should protect it the way you'd protect any other piece of critical infrastructure. It may not be efficient this quarter, but dismantling it quietly mortgages your capability a decade away. None of this is free, and pretending otherwise would insult the people who have to sign the budgets. A deliberate manual gate is by construction less efficient in the near term than full automation.
Steve Gibson [02:33:11]:
Keeping juniors doing formative work and running the manual sequences costs something now to protect something later. That's a hard sell in a market that judges most leaders on quarterly results. A hired executive who carries— he has in air quotes— unnecessary humans that AI could replace will hear about it from the board long before the payoff arrives. The math only works for someone insulated from that pressure. A founder with control, a private company, an institution with a genuinely long horizon, or a regulator willing to require workers to demonstrate their skills regularly, as pilots must. This all means the organizations most likely to preserve their own expertise are the ones structurally able to spend short-term margin on long-term capability. Everyone else will need a push from the outside. So here's the argument in one line: deliberate inefficiency is not waste.
Steve Gibson [02:34:24]:
In safety-critical engineering, we have always known it as insurance, and we buy it on purpose. As AI takes over the work where expertise is forged, the smart move is not to resist the automation. It is to keep our hands on the controls by design so that when the automation fails, as it always eventually does, there's still someone in the chair who knows how to fly. So, I think that's a fantastic piece of well-reasoned engineering, and I would recommend it without hesitation to anyone's boss who may currently be enraptured by AI's truly deliverable ability to eliminate all of those pesky junior engineering jobs. You know, we already had stunningly good AI with GPT-5.6 Fable and Mythos, And though it's gonna take some time for the world to fully assess what OpenAI has just given us with GPT-6 Astra, the early take is that it currently outperforms some of the other frontier models from OpenAI, Anthropic, and Google. Whether or not, and to what degree that's true, and even the fact that this is nothing more than a snapshot in time. That's important because, of course, this wasn't true a week ago, but may be true today. And we know that we are just at the beginning of all this.
Steve Gibson [02:36:01]:
We can feel how rapidly this field is changing. And so that's my point. If we've learned anything, it's that no one will ever have a long-term defensible position in AI. Sure, they're going to be bouncing back and forth. And Leo, you already knew that Anthropic was what, a week away or so from releasing their next model in this never-ending stream.
Leo Laporte [02:36:33]:
Supposedly. I mean, it's rumored anyway that they're ready with Fable 5. Yeah.
Steve Gibson [02:36:38]:
Anyway, I think without the deliberate creation of an environment that is designed to nurture and mature junior developers, organizations will be left completely dependent upon automation for mission-critical functions.
Leo Laporte [02:36:56]:
Yeah.
Steve Gibson [02:36:56]:
I mean, I know that, that coders who come out of, out of university with no experience, they're just not useful. They actually can't do anything. And you could argue that they're going to be a lot less useful. All they'll know how to do is drive AI. No coding. I don't even know if programming will be taught in the future. It'll be prompt engineering.
Leo Laporte [02:37:19]:
Right. And you think that we need to know how to code?
Steve Gibson [02:37:30]:
We probably need to know how to read what AI does when it writes a bug. You know, can we— I, I don't know yet whether we're going to be able to say to AI, fix the bug which you created, right? Maybe we will. I mean, it— so I guess the question is, okay, so certainly there will— well, AI is also coding itself now, isn't it?
Leo Laporte [02:38:01]:
Yeah.
Steve Gibson [02:38:01]:
So maybe it will be a completely lost Art, except here in Irvine.
Leo Laporte [02:38:06]:
Yeah, right. There'll be a few people. I mean, I think there'll always be a few people who code by hand, just like there are a few people who make their own furniture.
Steve Gibson [02:38:17]:
And play the piano.
Leo Laporte [02:38:19]:
Yeah.
Steve Gibson [02:38:20]:
Instead of turning on Apple Music.
Leo Laporte [02:38:24]:
Yeah. Code is a funny thing because code is Getting the computer to do something by translating your English language thoughts into machine code. Of course, you're already close to the machine code.
Steve Gibson [02:38:40]:
Even unvoiced, your intent, because it might be unvoiced.
Leo Laporte [02:38:44]:
Right. I mean, all— and as we've moved higher up, you still do low-level coding, but most people are doing high-level coding, which isn't really that closely related to what's going on in the computer. I guess there is a case to be made for people who need to be able to look at the traces and see what's actually happening.
Steve Gibson [02:39:05]:
Well, for a long time, compilers had bugs, which meant you would properly express yourself in the high-level language and you still wouldn't get the program doing what you told it to do.
Leo Laporte [02:39:16]:
Right. This to me is another case of We don't know what the hell's going to happen because—
Steve Gibson [02:39:25]:
Well, there are people talking about AI slop code and that we're now talking— that we're going to be introducing a whole new generation of bugs. I don't know one way or the other.
Leo Laporte [02:39:36]:
I think that's old farts talking.
Steve Gibson [02:39:38]:
Could well be.
Leo Laporte [02:39:39]:
Yeah. I think that that's people who are very tied to the old way of doing things. Slop is a pejorative that I don't think is completely fair. It doesn't code the same way you do. That's true. But the thing about computer programs is if they're correct, they work. Problem is these little edge cases where it works most of the time, it doesn't work all the time. But I think a computer is ultimately, at least I don't see any reason why a computer wouldn't ultimately be the best at figuring out what's going on.
Steve Gibson [02:40:15]:
Right?
Leo Laporte [02:40:15]:
It's better at it than we are.
Steve Gibson [02:40:17]:
I was arguably one of the early people to say that AI was going to be extremely good at code.
Leo Laporte [02:40:23]:
This is what it does.
Steve Gibson [02:40:27]:
Exactly.
Leo Laporte [02:40:28]:
So really what the AI is doing that's hard for the AI is, is it good at us? And I think most of the time when the AI fails us is because it's not as good at us or we're not as good at it. It's a communication problem between us and the AI, a misunderstanding or miscommunication or misdirection. I mean, I literally— I don't know.
Steve Gibson [02:40:51]:
I don't know.
Leo Laporte [02:40:51]:
I haven't looked at code in a long time, and I think most— many of the people like Darren, who's a very proficient coder in our Discord, I looked up to him because he's the guy who would finish all the Advent of Code problems lickety-split. So he's a very, very accomplished programmer, done it for years, his whole life. He says, I haven't coded anything myself for a year.
Steve Gibson [02:41:14]:
Wow.
Leo Laporte [02:41:16]:
I think many of the most proficient coders I know, people like David Heinemeier Hansson, don't look at the code. Now, his operating system, Almarki, which is a version of Linux that's entirely AI-coded, is kind of wonky. But that doesn't— I'm not sure I blame the AI. I mean, many of the coders who are using AI at this point have stopped looking at the code. So I don't know what that means. I don't know. I think— And Duran's saying, you can absolutely build great quality software without knowing anything about writing code right now. And I think that's only going to get better.
Leo Laporte [02:42:04]:
Does it? But there need— but it's absolutely true, there needs to be entry-level positions. You know, I think the analogy is to the business I'm in of broadcasting. Um, you're terrible when you start. You really are. You're not good at it. And you need somewhere to start. You need— basically, it took me a year to find a radio station. That was so tacky and small that it would let me learn how to be a broadcaster there.
Leo Laporte [02:42:33]:
And they put me on the middle of the night, midnight to 6:00 AM on a Sunday morning, because that was where I could do the least harm. But I got a chance to learn.
Steve Gibson [02:42:42]:
And it's, you know, we know it's called apprenticeship. The idea is you apprentice with a master and you learn his art.
Leo Laporte [02:42:50]:
Right. But then I look at my son, who really has no broadcast training, who just did it all in public on TikTok and has done a little bit— done pretty well for himself. A lot better than I did. Working out every day.
Steve Gibson [02:43:05]:
Dad paid for Yak.
Leo Laporte [02:43:06]:
And he did it, by the way, in a hundredth of the time his dad did. I mean, this is literally, he's gone from nothing to total success in whatever this business is. It's sort of like broadcasting. In a couple of years. Took me 20 years to— 10 years to get halfway decent. So it's a different time. And I think a lot of what you're seeing in a lot of this complaining and caviling about AI is us, the older— our older generation going, well, that's not how you do it.
Steve Gibson [02:43:39]:
In my day, you had to hold the hammer in your hand. How do you know what's in the registers? You got to know what's in your registry.
Leo Laporte [02:43:47]:
Do you care what's in your registries?
Steve Gibson [02:43:48]:
You don't.
Leo Laporte [02:43:49]:
You do, and you know intimately, but I don't know if you need to. And certainly somebody who's programming in C++ probably doesn't.
Steve Gibson [02:43:57]:
No, you don't have any access unless you explicitly tell it you want.
Leo Laporte [02:44:00]:
Yeah, you don't really care.
Steve Gibson [02:44:02]:
A bit of masm.
Leo Laporte [02:44:02]:
You don't care.
Steve Gibson [02:44:05]:
Yeah.
Leo Laporte [02:44:05]:
So I just don't know. We're an unknown, uncharted—
Steve Gibson [02:44:11]:
We are riding a beast.
Leo Laporte [02:44:12]:
It's a tsunami headed for us.
Steve Gibson [02:44:15]:
Yeah.
Leo Laporte [02:44:16]:
And there are a number of people who are going, run. And there are a number of people going, oh, look, the water's going out. I see seashells. And I think I might be in that latter group, and the wave's going to hit me. Steve, always provocative, intriguing, entertaining, and informative. It's a great show. And I really appreciate you putting all this work in. You do every In fact, you can subscribe to his show notes, 21 pages this week of really good stuff.
Leo Laporte [02:44:46]:
Read along with the show, read it yourself. There's links, there's images, it's everything you need. It's basically a book every week. Have you ever counted the number of words? Is it 3,000 or 4,000? 5,000? 8,000? I haven't counted. It's a lot. It's a short story for sure every week. And you can get it by going to grc.com/email. There's 2 things you'll accomplish by going there.
Leo Laporte [02:45:11]:
You can give him your email and he will whitelist it so you can email him, which is really nice. Send him pictures of the week, ideas, questions, comments, suggestions. But below that, there's 2 checkboxes, one for the newsletter, the weekly mailing of the show notes, another for a very infrequent mailing of new products. Speaking of products, Steve has, of course, SpinRite, the world's best mass storage maintenance, recovery, and enhancing utility. And you can get it right now at grc.com. That's Steve's bread and butter. Handwritten. He knows what regis— every single register in that program, he knows what it's doing at all times, right? There's no mystery.
Leo Laporte [02:45:52]:
It's— you know exactly what's happening in that machine, which is kind of amazing. And it's really good as a result. He also does a lovely little $10 program called the DNS Benchmark Pro, so you can get the fastest DNS. AI can't do that for you. You have to have a deterministic program like Benchmark. You have to. Somebody has to write a program to do that. And Steve did.
Leo Laporte [02:46:16]:
Thanks to Steve. He has the show as well. He has 16-kilobit audio and 64-kilobit audio. He also has wonderful transcriptions written by a human being, Elaine Ferris. They come in usually a couple of days after the show. All of that at grc.com. We have audio and video, somewhat larger audio file and video as well. You can download that from twit.tv/SN.
Leo Laporte [02:46:42]:
There is a YouTube channel dedicated to it. And actually, the best thing to do is subscribe in your favorite podcast client so you just get it automatically. You don't have to think about it. We record live every Tuesday right after MacBreak Weekly. That's 1:30 Pacific, 4:30 Eastern, 20:30 UTC. And you can watch us do it live if you want the absolute unedited best version of the show. All the expletives, all the nudity, it's all there.
Steve Gibson [02:47:09]:
Computer not working and rebooting.
Leo Laporte [02:47:10]:
Yeah, mostly that. Uh, if you watch live, Uh, there's several places you can do that. Of course, if you're in the club, and I hope you are, the Discord, but you can also watch on— everybody can— YouTube, Twitch, X, Facebook, LinkedIn, or Kick. And if you're chatting there, I can see the chat so we can, we can chat back and forth. I often do that. Um, Steve, that's it for us. I will see you next week and maybe I'll have a new iPhone. No, I won't have it by then.
Steve Gibson [02:47:38]:
Uh, till then, and we'll have a breakdown of Patch Tuesday's, uh, uh, fantastic, uh, patch count, and we'll know what Apple has wrought. So—
Leo Laporte [02:47:48]:
Holy camoly. Thanks, Steve. Thanks, everybody. We'll see you next time on Security Now. Security Now.