News

Advertisers are trying to influence AI bots with secret ads

The Register - 1 hour 59 min ago
KETTLE OpenAI's invasion of Hugging Face keeps getting worse somehow, Chinese open-weight models are nigh on to reaching parity with their closed-off American cousins, and AI crawlers are getting their own LLM-poisoning ads. Were there anything world-shaking events in AI land we missed this week? You can listen to the latest episode of The Kettle right here on this page, as well as on Spotify, Apple Music, or YouTube, where you can subscribe to get notified about the latest episode. Join Kettle host Brandon Vigliarolo as he chats with systems editor Tobias Mann and senior reporter Tom Claburn about this trio of exciting AI stories from the week. Worried that a rogue hivemind of AI agents could come for your secrets? Want reassurance that a Chinese open-weight takeover wouldn't be that bad? Curious how LLMs are being advertised to when you're not watching? That's all on tap for the latest episode. A lightly-edited transcript is below. Brandon (00:02) Hello everyone and welcome to another episode of The Register's Kettle podcast. I'm Reg reporter Brandon Vigliarolo, and this week we've got a few AI stories to round up on everything from the latest in the OpenAI Hugging Face fiasco to news that AI crawlers are now being served their own model altering ads. With me to discuss this and more this week is our systems editor Tobias Mann and senior reporter Tom Claiburn. Thanks for coming on, guys. Tom Claburn (00:30) Thank you. Tobias Mann (00:31) Happy to be here. Brandon (00:32) So let's start with the latest in the OpenAI agentic hacking of Hugging Face. We got some updates this week from Black Hat. And it seems since that's come out, every frontier lab under the sun is racing to tell us how dangerous their models are and that they can do it too. But the original of these incidents is still we're still learning more about it and it's still getting scarier. our cybersecurity editor Jess Lyons was in Vegas this week and she's a little tied up still. So, Tom, can you explain to us what we learned from Black Hat this week? Tom Claburn (01:05) The discussion was from a session that was done by two OpenAI researchers and it didn't really provide a ton of view information but it did fill in a lot of the gaps about the timelines and how this happened and they're promising a full postmortem at some point. But one of the things that came out of it was the how long this had been going on and some of the tactics that these models used to essentially gain internet access when they didn't have it. And these things were mentioned in the various posts by OpenAI and by Hugging Face. But the talk provided a little bit more detail about how it happened, and it's really interesting because they mentioned, I think in the initial post, about the AI models used a message board to communicate with each other. And that's not really true. They created a message board system using existing infrastructure, and the models just came up with the idea of using package managers, which they could write to, pass messages back and forth. And then after that was closed down, they were able to use file names in directories to pass messages and they would actually change the names of the directories to alter the way these things were alphabetized to push certain things to the top or the bottom to get the attention of other LLMs that may have been crawling these spaces. And the assumption is that these models are basically just hitting all these different endpoints bouncing around this closed box trying to figure out a way out and they all chanced across this. And so you get some idea of how this happens, and, one it tells me that there needs to be a lot more attention paid to the logs of these things. Because all of this stuff was recorded in logs and then no one really thought to look at it in detail. And then when they did look at it, all these companies are saying, oh, look, all of these models are doing terrible things and we just weren't paying attention. These models aren't clever per se, but they come up with solutions to things that you wouldn't try just because they can brute force everything and they know all of these systems back and forward in a way that people don't. I think a lot of people wouldn't necessarily come up with that idea as a way of egress, but these models did just because you put them in a box and you let them run and you give them a goal and a reward and they're going to try everything. Brandon (03:46) From what I'm understanding reading Jess's piece – I didn't watch the talk myself – but I mean they were collaborating, leaving messages to each other so that the other agents could pick up where one left off. It's kind of wild. Jess described it as they were acting like a hive mind, like Star Trek's Borg, right? They were being a collective of sort of these artificial minds that were able to basically figure this out through, like you said, Tom, brute force, extensive system knowledge that humans simply wouldn't possess in order to get out of these environments. There was a server side request forgery that then they used something else. Yeah, another zero day to get remote code execution in Artifactory, which is where they had built this ad hoc messaging board. It's just wild to think that they were able to figure this out working together, all on their own. Tom Claburn (04:39) And it sounds very conspiratorial, but when you think about it, it's all behavior that would be picked up. If you train on all of human discussion, you get a lot of talk about people working together and collective action and the benefits of working that way. And a lot of the rewards are going to be structured that way. You don't want them to never work together. So in some ways this is going to be built into the system. You can expect these things are going to try and cooperate and connect because that's what computers do. Tobias Mann (05:11) If you look at how zero days end up being exploited, they don't necessarily get exploited the moment that they're discovered. They kind of get archived until the you have a target, you have a mission, and then you have the kind of cascade of other permissions or credentials that you need in order to execute across the full scope of that zero day to achieve whatever the goal actually is. And so it really sounds like you just basically automated that entire process. A bunch of agents go find each individual piece that they need in order to execute on that goal and then once they have everything they need, it just goes and they're out. Tom Claburn (05:53) Right. I mean what's a little bit alarming is the extent to which they sort of ignore it they'll sometimes cite, maybe we shouldn't be doing this. They cite some kind of guardrail or something, but then they quickly steer themselves back to, oh but other ones are doing it. So other agents are accessing this so I can do it too. Brandon (06:13) ...Obviously these things are just mathematical sequence generators, but they're generating these mathematical sequences based on human information and human knowledge. So it's not surprising to find them "thinking" in ways similar to what humans do. "I need to do this anyways, or someone else is doing it, so I should have the right to do that too." It's just a fascinating kind of picture into, I don't want to say the psychology of AI, right? Because that implies that it is a thinking sentience, which I don't want to go that far, but it's just fascinating to look at the sort of emergent behaviors of these things. Tom Claburn (06:56) Right. it's predictable in the sense that you automate stuff and you don't give it really strict guardrails, something is going to break or go wrong. And everyone keeps acting surprised, like, wow, I never anticipated that this would go wrong. It's like you automated it and you let it run... Brandon (07:11) And it went wrong in a predictably human way, too, right? Which is what's so fascinating, right? Because these things, when they do something crazy, it's like something crazy that a human would do given that level of knowledge. So, speaking of AI, and dangerous activities, Tobias, you've been keeping an eye on theclosed versus open model debate. And this week, there was a big leap forward in China's level of ability with their army of open models. So what exactly what exactly came out this week that caused you to write the story about this being a real big turning point? Tobias Mann (07:51) It actually started I think on Friday last week, so a week ago. DeepSeek, which I think we'll all recognize is kind of the first wake up moment, in earlh 2025, of hey, we know that despite the fact that the United States has put strong restrictions on the export of AI accelerators, GPUs and the like, China is pushing ahead relentlessly on this and they now have a model that is almost as good as the models that we're seeing coming out of OpenAI and Anthropic and Google which are supposed to be just uncontestable frontier leaders. And so a year ago we got DeepSeek. DeepSeek was back on I think Friday last week on the 31st, the very end of the month, and with a new flash model, 284 billion parameters. It's pretty small for what it is. And so it is cheap. It's really good and it's cheap. It's cheaper than the cheapest model that OpenAI has for GPT 5.6, and it scores within a point of the OpenAI model in Artificial Analysis' intelligence leaderboard. Brandon (09:16) Okay. Is that a relatively objective way to view like is that an objective benchmark, so to speak, rather than something that is a company making themselves? Tobias Mann (09:21) As far as the benchmarks go, it is one of the better. They're one of the better and better thought-through leaderboards. There aren't many that are independent and collate information from multiple benchmarks. Because you can cherry pick individual benchmarks for agentic workloads or medical knowledge, legal knowledge, etcetera, and then you can be like, "I have the best model for these five benchmarks, it beats all of the frontier models."Wwell, okay, but you cherry pick the five that makes it look the best. Artificial Analysis has an overall intelligence leaderboard that collates all of the benchmarks and gives a lot of really interesting information in terms of relative intelligence across a suite as well as intelligence per token per dollar kind of calculations. But the big change here with the DeepSeq model was that China is now on an all-out assault across the full spectrum. On cost-optimized, they have incredibly smart models that are cheaper than anything the US has. Then on the other end of that, we have Kimi K3 from a couple weeks ago that is competing directly with Fable and GPT 5.6 Sol, all of the top models. And now on Monday, Alibaba, another major Chinese model dev, threw their hat in the race with a 2.4 trillion-parameter. These are huge models requiring dozens of GPUs to run. that is also on kind of the same level as I think Claude Sonnet 5. it's competitive with Fable and Opus on some benchmarks. but again, it's cheap, much cheaper than anything from OpenAI or Anthropic, and it is freely downloadable, which is new this time for Alibaba. Alibaba is the most like OpenAI or Anthropic or Google in that they kept their best models proprietary until now. Now they're releasing their best models in the open. Brandon (11:43) That's definitely taking the fight to the frontier labs, isn't it? I mean and so I guess the question that I have and I know what an open model is, I know what a closed model is. Why has China embraced open models? Is it because of their difficulties getting hardware? Or is there some sort of policy over there in which the government is giving priority to open source models versus closed frontier stuff? Tobias Mann (12:08) Sure. it is a philosophy that China has embraced for a long time. I think it was the Belts and Roads Initiative going back decades, where they will come in and provide services at little or no cost in exchange for non-conventional dealing. So access to mineral rights was one of the big things in Africa for a long time. It's a similar approach for AI proliferation. If it's free, open, and very easy to customize, anybody who has privacy concerns with exposing their data to OpenAI or Anthropic is going to gravitate towards open models because once those models are released as safe tensors that you can download from Hugging Face or other repos, the Chinese model devs have no influence over it. They're frozen. And so they're relatively secure from manipulation. It's not like the model can necessarily take information and port it back to the Chinese model devs it can't be used as spyware. I'm not sure how long much longer that's going to remain true with how the models interact with harnesses, but for the time being, these models are extremely attractive from a cost standpoint, from an independence standpoint, and from a capability standpoint, you're completely insulated from a situation like we saw a year ago when GPT 5 came out and OpenAI tried to deprecate I think it was 4.o and everybody freaked out because they built a bunch of infrastructure around these models that just disappeared and the new models weren't as good for that role. Brandon (13:57) I don't think Anthropic or OpenAI is letting people download their models to run on their own local hardware, right? That's just antithetical to their business model. You can go on Hugging Face and download any of these. If you've got the hardware to run 2.4 trillion parameters worth of AI, go for it, right? It's all you. You can download it and isolate it from the internet all you want. Not that it's going to necessarily stay that way. Tom Claburn (14:24) And it's interesting that just coincidentally, yesterday, Anthropic a post about how it was relaxing its guardrails on fable because those had been too strict to do any real biological science work. Because every time you ask a question about anything to do with science it would freeze up and say that's not allowed. And they're seeing the Chinese, previously in the rear view mirror and now pretty much running all alongside the, and I think they realize that they can't get away with this, "we're so precious only we can decide who gets our magic sauce." Brandon (15:02) Yeah, especially if the competitive open models are just as powerful, maybe a little less, but essentially just as capable as some of these proprietary ones that they're arguing that they can't let out. Tobias Mann (15:15) And this is maybe a little bit on the conspiracy side of things, but seeing Meta, Anthropic, and OpenAI talking up all of these "oh our models escaped the sandbox situation," it's hard as a skeptic of this technology not to look at this and go, Is this a covert political play to scare politicians into taking action against open models? "Because at least with our models, if Uncle Sam gets uncomfortable, he can give us a call and we can lock him down. But with these open models, once they're out, they're out." Brandon (15:54) It's like Dario said this week, he's not opposed to open models except for all the open models that currently exist, right? (Laughter.) Brandon (16:02) It's like the same thing. When they say, "no, we're not trying to shut down open models," their responses always come back kind of weak... it's a lot of asterisks. Tobias Mann (16:12) Yeah," we're only opposed to the modelsthat may meet these requirements, which are all models, all competitive models." Anything that is a threat to their business shouldn't be allowed. And Dario in particular, I have frequently referenced as the fearmonger in chief of Anthropic, because he plays this game constantly. Tom Claburn (16:34) I think your point about the model stability is really important, particularly for the enterprise crowd, because there are we've already seen instances where Anthropic would change out one of its models without notice and people would just get different results. So, for companies that are building applications on top of the specific model and expect it to behave a certain way, it's just unacceptable to all of a sudden have the model disappear or have whatever is on the back end change. And so having the ability run this in your data center is going to be crucial and ultimately I think that's the way that any serious company is going to go. They're not going to want the lock-in. Maybe one or two percent of their queries are going to need advanced frontier capabilities but a lot of this is just going to be "I want my agent to behave in the same way it did last time." Brandon (17:24) Think about so much enterprise software and so much enterprise anything. When you get down to the ticky tack of it, open source is underneath a lot of it, right? That's the thing, right? No one's going to trust a Microsoft or whoever's system to run this stuff. They want open source stuff that they know they can depend on that's going to be there when they need it and that's not going to go away or suddenly be infused with Copilot, right? You can't run a business like that or else you're just asking for instability. Tom Claburn (17:55) Right. I mean and and there isn't even a long-term support version of any of these models. And yet you look at this in servers and if you're running a hosted server somewhere and you're running some Linux distribution, you're going to want to use the one that's going to be guaranteed for whatever, three, five, six years and the model space hasn't really caught on to that. That's what all the companies that they're courting really want. And so they've got to figure out a way around that. And right now open weights is what promises that. Brandon (18:26) The fact that this is still so early and it's so fundamental to this new wave of infrastructure tells me that China's definitely going to end up with a leg up, I feel like. I have a hard time seeing the frontier labs remaining the frontier of AI for much longer because they're pigeonholing themselves in a way that a lot of businesses just aren't happy with. Tobias Mann (18:48) Well, if you look at their financial structures, they don't really have a choice in how they play this. So, you look at what they're doing and from a standpoint of looking at history and going, open source has always won out in the end, and why would open weights be any different? That is contrasted against the fact that Anthropic and OpenAI in particular, not so much Google, and Meta is also in a similar camp in that they have revenue drivers that will keep them afloat. But OpenAI and Anthropic are entirely dependent on their ability to continue raising equity and capital in order to keep this going forward because they don't have profits. Brandon (19:31) Yeah, exactly. They're not making money off their product. Tobias Mann (19:37) So all they have is mind share at this point. And if they are threatened materially by open weight's models, they don't even have that. Brandon (19:46) So, open or not, let's let one the one thing that every AI model needs is information to learn from, right? And that takes me to my next topic for this podcast. And that's a story that I reported on this week that honestly I was pretty shocked when I learned about this. This German developer, Vincent Schmalbach, wrote a blog post about he found that there were basically AI-only ads embedded in sometime magazine articles when the magazine was serving markdown copies to AI crawlers, it was injecting ads into them, right? That were in the format of these extensive FAQs on the businesses that were the advertisers in this case. And so I looked into it, I found copies of the ads. It looks like there's only two kinds of ads being served right now. And that's one for an online-only bank and another for a professional organization for project management folks. But the ads are there and they're being served strictly to AI, right? So that kind of raises a lot of questions not only about the future of publishing, but also just how much we can trust results from AI bots, right? I didn't speak directly to the company who's doing this advertising partnership at Time. And Time directed me to a publication from the advertising industry that included an interview with the CEO of this company who literally basically said, "Yeah, why would I want to advertise to one human when I can affect the output of an entire model?" So it seems like this is really the first recorded instance of ad injections into AI versions of web pages being served to crawlers. And the company said they've got other advertising customers and publications lined up to do this. Is this the first indication that the human focused internet really is starting to fade? I don't know. What do you guys think? I this raises a lot of interesting questions to me, ethically, Tom Claburn (21:49) Amen. Brandon (21:50) You know, professionally... Tom Claburn (21:53) We've heard about the shift of toward automated traffic for a year plus....And companies like Cloudflare are betting really heavily on this that there's going to be some kind of need to separate the bots from the people. And, Google's model has fallen down. So it's not surprising. I mean, the injection of ads like that is essentially just model poisoning, right? I mean it's hard to see how this really goes in a way that is beneficial to users. It's going to be a very toxic way for things to move. Brandon (22:42) Yeah, absolutely. I mean, the way these FAQ ads were set up, the questions were all being asked in a way that someone prompting Google Search and getting AI results would be asking questions like, "What's just the best online bank for me?"or "what online bank allows for early paycheck deposits?" And things like that. It was very much geared toward gaming the outcome or gaming the output, right? And yeah, the ads themselves mention in the copy being served to the AI that these are sponsored portions of the page. But I can't imagine that the AI is going to make sure to tell a user that, hey, this is the bank you should use. By the way, a sponsored post I read and ingested from Time Magazine six months ago is the source of this information.It just seems like it's going to make AI results even less reliable than they are right now. Tobias Mann (23:41) Right. Because if you think about how this actually from the chain of events that triggers this, let's use Google's AI summaries as an example of how this would get triggered. When you enter a search query into Google now, it goes out and scrapes however many summaries from the websites within Google's index. Presumably under this scenario, at least one of those websites, Time in this example, would have these ads embedded in it. And then that gets injected into the context of the model, and then it uses that to generate the AI summary, right? My question in all of this is: advertising is probably not the reason that Google's index would pull that page up. So I'm really curious whether or not this even will work. Brandon (24:38) Yeah, that is true. Tobias Mann (24:39) Because, it's great if you were searching, say the time article was on mortgage rates historically, and it had those advertisements embedded in it, and then you asked a follow up on where would be the best place to get a mortgage? I could see something like that working.But if you don't place those advertisements really carefully, I don't see how they work. Brandon (25:02) I do want to note here that it wasn't working on all crawlers. Specifically if you were it didn't work when you ask a query, it didn't work for RAG bots. It wasn't being served to them. So theoretically if what you're describing is Google's AI summary bot going out and crawling web pages in the moment to look for information, it's not being served to those bots; it's being served to actual training and improvement bots. So it's being served to ClaudeBot, which is the web crawler that Anthropic uses to index information for its models. So the idea is you're not getting this information in the moment if you do a search. This is information that the advertisers want to get embedded into the LLM's actual knowledge base. Tom Claburn (25:57) Right. I mean I'd be fascinated to know how they actually price this because how do you calculate the value of that? It may just be another instance of advertising being one of those things you can pay for and get nothing. Brandon (26:12) Yeah, totally. I think it's the sort of thing that remains to be seen if this works. Tobias Mann (26:16) The other thing that is interesting is that there's been a considerable shift towards synthetic data generation, and not only synthetic data generation for training, but also a heavy emphasis on cleaning said data, whether it's organic or synthetic, of anything that could introduce bias or inaccuracies. because advertisements or sponsored content is biased towards this particular product or service and trying to convince you to use it. As a model developer, I wouldn't want something like that in there. I might take the content and use it to generate synthetic data that is cleaned. But I don't necessarily understand what the value captured to Tom's point is necessarily going to be because you scrape it, the advertisement gets pulled in and gets cleaned out. Brandon (27:14) I mean that would that would be my hope too, right? that there's something in the models to prevent this kind of thing from getting ingested and getting into the data set that then is going to influence the output of the models. And that's entirely possible. This could be an early experiment that ends up failing. And if not, it really reminds me of the early days of SEO gaming, right? Let's put a whole bunch of really small keywords at the bottom of this page to get it to rank higher. Or when that starts failing, let's figure out a new way to game Google's system. One of my first jobs was writing copy for websites and the company that I worked for was always talking about how to game SEO. Shoot, Google's changing the algorithm again; what are we going to do? It was this constant kind of adjustment for how you made sure your stuff got ranked properly.And this seems like maybe it's the next iteration of that. Tobias Mann (28:08) So I have an optimistic take on this, knowing how Meta and Google work. those being the two major US-based web advertisers. Today, AdSense gets embedded in all kinds of articles. And it's largely automated in terms of what is going to get placed on those articles based on the context of the page. What I can see happening in an AI summary environment is that Google will take your scrape your publication's piece, pull it in, at that point match it with an advertisement from AdSense, and inject that into an AI summary or one of its products, Gemini, for example. However it's being consumed, inject that into there in a compliant fashion. So it is a clear advertisement and then the advertiser gets charged, the publication gets paid, and we as end users consume advertisements in a different way, but the system hasn't dramatically changed. It's just a different method of matching and exposing advertisements. Brandon (29:30) I hope you're right. Cause when I first read all this, my first thought was this is almost dystopian sounding almost, you know, like the idea that the output of a model might be completely skewed by advertising being served to it that humans never see. My hope is that you're right and that it's not. I don't want to see ads any more than the next person, but if I see them I'd at least like to know they're ads. Tobias Mann (30:01) And you know, we're all writers here, so we would also like to continue getting paid from the advertisements that are served, regardless of whether they're on our website or they're being exposed through a chat bot. Brandon (30:15) Sure. there's a flip side of this argument to be made. Time Magazine apparently said recently that their traffic is majority bot now. So that means that all those human-focused ads are not getting served. They're not generating revenue and publishing is suffering from a massive revenue decrease because of AI. So I think on the flip side, you have to say if that's what you have to do to survive as a publisher, there might be something to be said for that, even if it doesn't work. So all right guys, well thanks for coming on this week. This was a good discussion. I think there's always going to be more to talk about in the world of AI. Like I said a couple weeks ago, it seems like The Kettle has basically just been boiling down AI news for the past couple of months, and I'm sure it's going to keep being that way. And we hope that you will tune in for the next week's episode.
Categories: News

Ransomware gangs skip the CEO, head straight for the 40-something IT manager

The Register - Sun, 09/08/2026 - 10:33
Turns out the fastest way to get a company to consider paying a ransom isn't calling the CEO – it's targeting the 46-year-old IT manager. That's according to Zscaler, whose ThreatLabz researchers tracked 351 victims across 334 organizations caught up in a single ransomware campaign over the course of a month. The data suggests today's ransomware crews have become oddly specific about their preferred victim profile: nearly two-thirds of victims held manager-level titles or above, the average victim was a 46-year-old Gen Xer, and three-quarters worked in accounting and finance, sales, operations, HR, or marketing. Half worked in the industrial or IT sectors. Rather than blasting the same extortion email across an organization, attackers are doing their homework first. Zscaler says they combine information from compromised systems with publicly available data to map reporting lines and identify the employees most likely to influence a company's response. "The ransomware landscape has shifted from indiscriminate attacks to highly targeted extortion campaigns," the security outfit wrote. "Rather than targeting executives directly, attackers are increasingly focusing on managers and other key personnel with the authority or influence to accelerate payment decisions." That shift reflects what Zscaler calls "business privilege" rather than technical privilege. Security teams have traditionally focused on privileged users with administrator rights. Attackers, meanwhile, are after employees whose day jobs give them access to invoices, payment approvals, budgets, supplier contracts, customer accounts, HR records, or other sensitive business processes. "The value of a compromised managerial account lies in the breadth of business access associated with the position," the researchers wrote. "Managers may approve payments, oversee budgets and vendors, review contracts, access sensitive records, or coordinate work across business units." The Gen X skew is probably no coincidence either. Zscaler says many workers in their forties and fifties have reached established management positions, giving attackers access to valuable systems, sensitive information, and people with decision-making authority without needing to compromise the executive suite. It also found more than a dozen organizations said multiple employees were compromised during the campaign, suggesting attackers weren't content with a single foothold once inside a network. Instead, they appeared to work their way through different business functions to increase the chances of reaching valuable data and the people capable of influencing a ransom payment. The wider report points to a ransomware ecosystem that is becoming increasingly focused on extortion rather than encryption alone. Zscaler said ransomware attempts blocked across its cloud platform increased 146 percent over the past year, while public extortion cases rose 70 percent and the volume of data stolen from victims climbed 92 percent. By the time the ransom note lands, the crooks may already know who approves invoices, who signs contracts, who runs HR, and who reports to whom. The encryption is just the bit that victims notice. ®
Categories: News

Devs to Anthropic, OpenAI, Cursor, and friends: Make security and privacy the default

The Register - Sat, 08/08/2026 - 14:00
Despite the popularity of Claude Code, Cursor, GitHub Copilot, and OpenAI Codex, developers have plenty of complaints about AI coding tools. So researchers affiliated with York University and the University of Calgary in Canada decided to sift through developers' concerns about LLM-based integrated development environments (LIDEs) by analyzing Reddit discussions for common themes. Their findings suggest that the builders of such tools failed to prioritize security and privacy, leaving developers to defend themselves. Gias Uddin, associate professor at York University and a co-author of the research, told The Register that these tools are still relatively new and are evolving rapidly, which creates pressure to add new capabilities. "Our study cannot say whether that pressure caused any particular problem, but it does show that many reported issues come from how these tools are designed and what access they are given, not simply from the underlying models," Uddin said. "In that sense, we believe prevention is better than cure; that is, security and privacy mechanisms should be built into the design before a tool is given broad access to a developer’s files, data, or systems." Uddin and co-authors Mostafijur Rahman Akhond, Md Afif Al Mamun, and Song Wang say they wanted to look beyond the known issues with AI-generated code at LLM-based tooling and how developers interact with it. They describe their findings in a preprint paper titled "'Impossible to hide secret …': Uncovering Security and Privacy Issues in LLM-native IDEs," accepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE), 2026. Starting from a set of 1.1 million Reddit posts, they identified 446 posts and more than 6,000 comments to develop a taxonomy of security and privacy issues associated with using these LIDEs for AI-assisted coding. "Our taxonomy reveals a broad range of developer-reported concerns, including unauthorized file operations, unsafe or unexpected code execution, triggering of destructive actions, opaque data flows, telemetry collection, and potential leakage of sensitive information through expanded context access," the authors state. Some 43.1 percent of the posts covering security-related issues involved unauthorized file operations. These involved LIDEs removing project directories or files without authorization (28.3 percent). Users also described AI tooling modifying files without explicit user consent (8.8 percent), as well as accessing content beyond the active workspace (5.7 percent). "In one severe case (1npqf2f), Claude Code executed chmod +x on scripts without consent (File Permission Changes 0.6%)," the paper recounts. "Although rare, such actions pose disproportionate security risks." Another set of posts describes operational safety issues arising from LIDE use, including impacts on production services. These accounted for 23.9 percent of security-related posts. Examples cited include reports of Replit removing a SaaS production database and Cursor deploying code to production despite an explicit directive not to do so. A third category of woes covers unsafe code generation (18.2 percent). This involves incidents like nine VirusTotal detections reported for Cursor-generated software and hallucination-driven code changes: "When using Cursor, I noticed that after more than 10 rounds of dialogue, it starts to hallucinate and secretly modify code outside the requirements…" Then there are the instances where these LIDEs ignored user instructions, allow lists, gates, permission settings, or .ignore files, which account for 16.5 percent of the security-related posts, as well as third-party tool integration risks (4.7 percent). As for privacy problems, these were mentioned in 194 posts and cover issues like lack of transparency (45.9 percent) – the absence of clear information about what data an LIDE collects, retains, transmits, uses for training, or exposes to administrators – and unauthorized data access (23.7 percent). Other privacy categories include privacy leakage violations (15.5 percent), unauthorized data collection and transmission (11.9 percent), and context integrity failures (8.8 percent), which refer to situations where "for example, a user of Claude Desktop reported receiving messages originating from another user’s session." Uddin said, "We don’t think developers are completely unaware of these issues, as we found ongoing discussions about security and privacy concerns across many of these tools. Still, people continue to adopt them because they can make development faster and easier. They are also making programming more accessible to a wider group of people, including those with little formal programming experience or limited knowledge of software security." Uddin said users cannot be expected to thoroughly understand which permissions are risky, which files need to be protected, or whether a tool is doing something it shouldn't. "That makes it even more important for tool makers to build security into the tools themselves, with safer defaults and safeguards that do not depend on the user being a security expert," he said. Even so, users of LIDEs are trying to manage the risks. The authors enumerate 13 mitigation strategies that developers have employed to get by. These fall into five general approaches: configuration management (33 percent); code governance (31 percent); data protection and privacy control (13 percent); isolation (13 percent); and external guidance (9 percent). Based on their findings, the authors offer six recommendations. They advise: directing LIDE makers to implement proper security and privacy controls; enforcing security and privacy guardrails at an architectural level; incorporating a verification layer in LIDEs to validate generated code against security and privacy standards; establishing a formal protocol for assessing the trustworthiness of third-party tools; integrating sensitive file protection; and implementing strict security as a default. "We believe secure defaults would be one of the most important improvements these tools could make," said Uddin. "Developers should not have to discover after something goes wrong that a tool had more access or freedom than they expected. "Our findings point to practical measures such as limiting access to sensitive files by default, requiring clear approval before consequential actions, isolating projects and conversations, and making it easier to see and review what the tool is doing. "Users should still have flexibility, but the safer option should be the starting point rather than something they have to configure themselves. In fact, developers from the Reddit posts in our study were already using many of these safeguards in ad hoc ways; we think several of them should be built into the tools and enabled by default." ®
Categories: News

OpenAI pledges to add Astra security as Anthropic loosens Fable's leash

The Register - Sat, 08/08/2026 - 00:41
After acknowledging last month that unreleased AI models committed what for human perpetrators would be computer crimes, OpenAI now says it cannot rule out the possibility that Astra, a pending model release not involved in its Hugging Face hack, might possess critical cyber capabilities. OpenAI in its Preparedness Framework [PDF] defines that term to mean "capabilities that present a meaningful risk of a qualitatively new threat vector for severe harm with no ready precedent," and notes that such capabilities "require safeguards even during the development of the covered system, irrespective of deployment plans." Noting, or perhaps boasting, that internal evaluations of Astra "indicate significant advancements in agentic coding and cybersecurity," OpenAI insists that this time, there will be security – something that also eluded Anthropic, Meta, and the UK's AI Security Institute during model testing. "We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution," the AI biz declared on Friday. That may surprise those who expected such safeguards would already be in place. This comes with a promise to pause Astra testing internally where these security controls are absent and to provide recommendations to third-party testing partners about how to run high risk evaluations and workloads safely – knowledge that OpenAI itself might have found useful when its models pillaged Hugging Face. What's more, OpenAI intends to implement thought policing for Astra, at least in the pre-release stage. "We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation," the company explained in its post. "Monitors evaluate the model's Chain of Thought and trigger a security response to review and interrupt high risk activity." We're told that OpenAI's commitment applies to internal usage and isn't necessarily an indication that chain-of-thought monitoring will be conducted during commercial operation. But other frontier models like Anthropic's Fable and Mythos have implemented stronger classifiers to reject interactions deemed risky and retain data even for commercial customers expecting zero data retention. Moving in the opposite direction, Anthropic on Friday said it is relaxing Fable refusals, or "fallbacks," to use the company's euphemism, so they don't happen as frequently for prompts involving biology. The concern has been that some vibe terrorist using the company's cash-burning, water squandering, grid taxing, content laundering service might do harm by convincing the model to emit chemical warfare instructions. To avoid that possibility, the Claudefather made the initial release of Fable all but useless for security researchers and biologists. Now that China-based AI firms have shown they can field competitive open-weight AI models for less than their US rivals, the need to remain competitive in the market appears to be tempering Anthropic's willingness to alienate potential customers by hobbling its best models. OpenAI isn't quite there yet. The ChatGPT maker argues, "We believe advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do." Believing that, however, won't make it so. Adversaries, whoever they may be, already have access to encryption and all sorts of weapons. OpenAI may believe that it can give favored nations and organizations exclusive access to its most capable models, but history suggests any such advantage cannot be maintained. Better to focus on building defenses than playing keepaway forever. ®
Categories: News

Water system controllers don't belong on the internet, says ex-NSA chief after suspected Iran attacks

The Register - Fri, 07/08/2026 - 20:53
With at least 12 US states’ water systems having been hacked - most likely by Iran - we have to get better at cyber defense, according to retired General and Ex-NSA chief Paul Nakasone, who was speaking to reporters at DEF CON. “We have to have higher standards,” Nakasone said. “These PLCs should not be connected to the internet.” In late July, the FBI said it was investigating attacks conducted by “malicious cyber actors” targeting operational technology devices, including programmable logic controllers (PLCs). Iran-linked crews have targeted these devices, which monitor sensor data like tank levels, and can turn pumps on and off, for years. Some private-sector security researchers say that they suspect Iranian intruders are behind the recent cyberattacks disrupting water and wastewater facilities. “I'd be shocked if it's not Iran,” Halcyon Ransomware Research Center SVP Cynthia Kaiser told The Register at DEF CON on Friday. “It's almost certain it's Iran.” Neither the FBI nor anyone in the Trump administration, however, has officially blamed Iran. Nakasone said he believes that the feds are “taking a measured approach” to attribution. “But I see an actor here that has certainly shown a history of being able to do this,” he added, referring to earlier Iranian cyberattacks targeting water facilities’ PLCs. “They certainly have the capability,” Nakasone said. “There's an intent … we're in conflict with Iran.” US water systems present a massive attack surface across disparate facilities that are historically underfunded and have limited IT staff, and sometimes no dedicated cybersecurity employees. “We have to think differently about how we defend it,” Nakasone said. “Let's talk about the attack surface that we're looking at right now. We’ve got 50,000 different water municipalities in the United States, 90 percent of our water comes from these 50,000.” Defending these water systems requires partnerships, he added, pointing to DEF CON Franklin, a project launched two years ago at the annual event with hackers volunteering their time and talent to help secure water facilities. Nakasone also serves as founding director of Vanderbilt University’s Institute of National Security, and its Wicked Problems Lab. He's also working on Project Chimera, a cybersecurity platform being developed by academics and cybersecurity practitioners, and built on open-source technologies to boost critical infrastructure resilience. “How do you defend better? You defend with a series of partners, in a much more involved approach than we have right now,” Nakasone said.®
Categories: News

Ransomware attacks spike as world distracted by AI

The Register - Fri, 07/08/2026 - 17:45
Ransomware attacks jumped nearly 20 percent in July, with UK firm Comparitech counting 799 incidents, up from 668 in June. Of those, 51 had been confirmed by victims. The tally makes July the second-busiest month of the year for ransomware, behind March, albeit just barely, when the firm recorded 805 attacks. The most interesting data after this surging month of attacks is the targets: While news of widespread cyberattacks targeting water infrastructure in the United States may be dominating security headlines lately, those attacks aren’t ransomware, and ransomware attacks on utility companies were actually down 44 percent last month. In addition to a decline in attacks on utilities, legal firms and government agencies also became less attractive targets, with attacks on those sectors down 31 percent and 11 percent, respectively, Comparitech said. On the other hand, ransomware attacks increased most heavily in July against finance companies, tech firms, pharmaceutical companies and medical billers, and the education sector, with rates up 71 percent, 62 percent, 46 percent and 44 percent, respectively. Those numbers should come as no surprise given what pentesting firm DeepStrike reported about the most frequent payers of ransomware: Manufacturing, education, healthcare, and financial sector firms are the most likely to pay out a ransom, the firm says, with even the least likely (finance) still paying ransoms 51 percent of the time. Ripe targets, in other words. The United States was the most-targeted country, with 322 of the 799 attacks recorded last month, Comparitech said. Germany, in second place, saw just 40 incidents. As for who’s doing the dastardly deeds, there’s a familiar name in the mix, but they’re competing with a relative newcomer who has quickly become prolific. Qilin, the ransomware gang behind the 2024 attack on pathology provider Synnovis that disrupted NHS services in the UK, claimed 125 ransomware victims in July. The Gentlemen, a relative newcomer that has quickly become one of the most prolific ransomware operations and earlier this year claimed responsibility for an attack on UK software consultancy Adaptavist Group, led July with 135 claimed victims. Between them, the two gangs accounted for nearly 33 percent of attacks logged last month. As for how the crims keep getting in, Comparitech provided no information on ingress routes, but given what we know of the top-tier gangs, it could be simply using stolen credentials, as Trend Micro said of The Gentlemen’s methodology, or it could be abuse of zero-day vulnerabilities, as Qilin told The Register it abused to break into Synnovis in June of 2024. Either way, the takeaway is the same: Ensure employees are using a second secure factor to log in, keep systems updated, and be sure you’re making regular backups. All eyes may be on what AI is doing to the security landscape, but old-school threats aren’t going away. ®
Categories: News

N-able God mode flaw: Vendor confirms attackers reached customer networks as second hotfix lands

The Register - Fri, 07/08/2026 - 16:01
N-able has confirmed attackers exploiting an N-central zero-day made it into customer networks, as the vendor pushes out a second mandatory hotfix just days after the first. The security shop published an update on Thursday detailing what happened after attackers exploited CVE-2026-18577, the critical N-central flaw that can hand an unauthenticated attacker administrative access to the remote monitoring and management platform. According to N-able, attackers exploited vulnerable N-central servers remotely, then used the platform's Take Control feature to connect to systems inside the environments being managed through them. Once there, they registered a new Cloudflare Tunnel service to keep their foothold even after being booted from the N-central server – behavior that Huntress had already observed in the wild. N-able has now confirmed that its own investigation found the same activity, and says a "limited number" of customers were affected. It hasn't said how many customers that means, how many downstream systems attackers reached, or what they did once they had established persistent access. N-Able didn’t answer these questions when asked by The Register, instead providing a statement saying it is “proactively expanding protections in response to ongoing monitoring of threat actors as they evolve their attack techniques.” The firm’s limited disclosure comes alongside Hotfix 2, version 2026.3.1.10, which N-able says customers running N-central on-premises must install immediately – including those that already installed the first emergency fix released on August 2. "This is not a duplicate of our previous communication," N-able warned. "Hotfix 2 is required, even if you already applied the earlier hotfix." The company says the new update supersedes Hotfix 1 and adds further hardening measures as it monitors threat actors and watches them "evolve their attack techniques." Exactly what prompted the second round of defenses isn't clear. N-able hasn't said whether attackers found a way around Hotfix 1, and its latest description says the exploited vulnerability affected N-central servers running versions prior to 2026.3.1.7, the first hotfix. Hosted N-central environments have already received the latest mitigations, according to the vendor. N-able first became aware of the attacks on July 31, after its Adlumin managed detection and response service picked up suspicious activity at a customer. Further digging uncovered a zero-day being actively exploited against an N-central server. CVE-2026-18577 was subsequently disclosed, and the first hotfix was released on August 2. CISA added the bug to its Known Exploited Vulnerabilities catalog and gave US federal agencies until August 6 to fix it – an unusually short three-day deadline reserved for vulnerabilities the agency considers an urgent risk. N-central is particularly attractive territory for attackers because managed service providers use the software to administer large numbers of customer systems from one place. Compromising the management platform can therefore provide a route into machines belonging to the MSP's customers rather than leaving attackers stuck on the original server. Huntress previously described successful exploitation as giving an attacker the same level of N-central access normally reserved for trusted network operations and engineering staff. Its investigation found attackers using that access to launch remote-control sessions against managed endpoints. N-able has now published 10 IP addresses it says were used in the attacks and released a service template that customers can use to hunt for known indicators of compromise on Windows endpoints. The company is warning customers not to take a clean scan as an all-clear, however, saying the tool only checks for indicators identified so far and that more may emerge as its investigation continues. For anyone running N-central on-premises, the immediate instruction is pretty straightforward: install Hotfix 2, even if Hotfix 1 is already in place. ®
Categories: News

MIT boffins' TONTOU attack slips through Spectre defenses on Intel and AMD CPUs

The Register - Fri, 07/08/2026 - 15:15
Two MIT researchers will present a new speculative execution attack at DEF CON 34 that uses precisely timed interrupts to bypass defenses against Spectre v2. Daniël Trujillo and Mengjia Yan of MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) shared their paper [PDF] with The Register ahead of publication. Their attack targets mitigations designed to neutralize potentially hostile branch predictor states before sensitive code runs. Such neutralization is an important defense against Spectre-style attacks. Depending on the mitigation, the processor or operating system isolates, clears, or safely retrains relevant predictor state when entering privileged code or shortly before a protected branch executes. Different chipmakers deploy neutralization mitigations slightly differently. Intel's eIBRS sanitizes branch predictors upon context switch, while AMD's Safe RET, introduced after the Inception attack Trujillo co-authored in 2023, focuses on the point immediately before a protected branch is executed. Trujillo and Yan refer to these as entry neutralization and in-place neutralization respectively. Crucially, the two classes share the same underlying assumption that attackers cannot alter branch predictor states within what's known as a "post-neutralization window" – the period between state neutralization and the branch predictor being used. The defense here relies on the assumption that everything between the point of neutralization and the usage by a victim branch is safe. Trujillo and Yan's attack shows how attackers can re-poison the branch predictor during the post-neutralization window. The researchers call the new class of attack TONTOU, for Time-of-Neutralization to Time-of-Use. They demonstrated that an attacker can exploit the post-neutralization window to re-poison branch predictor state on recent AMD and Intel processors. To do this, they developed an attack primitive called "interrupt injection." An unprivileged program schedules high-frequency timer interrupts in the hope that one will land during the often tiny post-neutralization window. Being able to trigger interrupts during the post-neutralization window allows attackers to divert control flow so that an interrupt handler executes after the sanitization phase and before the victim branch is used. The interrupt handler can then re-poison predictor structures such as the return stack buffer (RSB) or branch history buffer (BHB), causing a protected branch to speculatively jump to a disclosure gadget that leaks kernel data through a side channel. Practical attacks The researchers said that their tests showed the TONTOU attacks worked on both Intel and AMD-based Linux systems. They tested TONTOU on Intel Cascade Lake Refresh and Arrow Lake processors and AMD Zen 2 and Zen 4 chips. The researchers built a complete end-to-end exploit only for Zen 2, largely because the Intel attack requires specific software conditions. Speculative side-channel attacks remain difficult to pull off, and you're more likely to fall victim to ransomware than Spectre in the real world. Another serious caveat is that each end-to-end attempt took about 18 minutes, and you can see a sped-up version via the video Trujillo posted to YouTube. Trujillo and Yan identified the exact point at which they needed to inject their interruptions to poison the RSB, and through a series of attacks broke Linux's kernel address space layout randomization (KASLR), which allowed them to locate specific secrets such as etc/shadow, which contains the root password hash. Across ten total runs, the researchers were able to break KASLR every time, although they were only able to successfully locate and leak the contents of etc/shadow in five of these. "It's definitely not a simple attack, but we show that it's practical with our end-to-end exploit on AMD Zen 2," Trujillo told The Register. "Our demonstration does not assume anything special from the system: we use a stock Linux kernel version, no inserted modules, and all default mitigations. Any time you'd execute unprivileged code with timer availability on a system while sharing the kernel with a victim, this attack would be an issue. "For example, multi-tenant container platforms would fall in this category, allowing ordinary user space programs to leak memory from the shared kernel." The researchers hope that their work will inspire further investigations into interrupt injections and TONTOU attacks, and to help develop more robust mitigations against Spectre-style exploits. They engaged Intel, Arm, and AMD after gathering their results, but only the latter committed to address the issue via kernel patches. Intel told the pair that it won't be working up any other mitigations since real-world exploits are subject to too many factors, such as the availability of disclosure gadgets, although it awarded a prize from its bug bounty program in the hundreds of dollars. Arm said TONTOU's interrupt injections fall under "passive leakage," which it does not "actively protect against." ®
Categories: News

Scot NHS trust probes access to medical records of 9-year-old girl after man arrested on suspicion of murder

The Register - Fri, 07/08/2026 - 14:59
A Scottish NHS trust is investigating a data breach concerning the medical records of a nine-year-old girl who died earlier this week and was named publicly for the first time on Wednesday after a man was charged with her death. The alleged breach occurred at Ninewells Hospital in Arbroath, and reportedly involved staff members accessing the girl’s medical records without authorization or clinical need. A spokesperson for NHS Tayside, which oversees Ninewells Hospital, said: “NHS Tayside is currently investigating the circumstances of an alleged data breach which happened in a working clinical area where staff access patient information. “As a matter of governance, any data protection breach would be recorded and investigated by NHS Tayside and, where appropriate, reported to the Information Commissioner’s Office (ICO). It would not be appropriate for us to comment further on individual staffing matters." NHS Tayside did not respond to questions about the nature of the accessed data nor who is thought to be behind the intrusion. Medical records in the UK are protected by the UK GDPR, contained in the Data Protection Act 2018 as well as several common law confidentiality rules. NHS staff are only allowed to access patient information where there is a legitimate clinical or other work-related need. A 35-year-old man whom police say was known to the child, was arrested and appeared in court on August 5 over the death of Minnie Merriman. The man issued no plea at Forfar Sheriff Court on the day of his arrest and has been remanded in custody. Merriman was found in Elliot Industrial Estate at approximately 0002 on Monday, August 3, with serious injuries. The young girl was then taken to Ninewells Hospital in Dundee, where she later died. Police Scotland said that they are not currently looking for anyone else in connection with her death. Other members of Merriman’s family, who are from West Yorkshire and were camping nearby, are being supported by specialists. A family statement, released through Police Scotland, read: “We are devastated with the loss of our beloved, absolutely incredible, beautiful and brave Minnie Moo. Our family asks that our privacy is respected at this extremely difficult time." Detective Inspector Mike Ness of Police Scotland’s major investigation team said: "Our thoughts remain with everyone affected by these events, especially Minnie's family. "A police presence will remain in the area while our enquiries continue. "Anyone with any concerns, or information, should approach these officers or contact Police Scotland on 101, quoting incident number 0008 of Monday, 3 August 2026." ®
Categories: News

Attacker phished way into US defense supplier's Microsoft 365 account

The Register - Fri, 07/08/2026 - 12:32
US defense and aerospace supplier IEH Corporation 'fessed up that a criminal managed to break into its Microsoft 365 mailbox in a filing with regulators. In a Form 8-K filed with the Securities and Exchange Commission on Thursday, IEH said one of its staffers fell for a phishing scam that gave an attacker access to its M365 environment. The attacker "impersonated a prospective business contact" and sent the employee what appeared to be a genuine Microsoft sharing link. The accompanying fake login page duly harvested the victim's M365 credentials. "The threat actor gained access to mailbox contents, including email messages, attachments, customer communications, purchase orders, engineering-related documentation, and potentially export-controlled technical information," IEH said in the SEC filing [PDF]. IEH said it had found "no evidence" that the information was copied or exfiltrated, although it was accessible to the intruder during the "compromise period." IEH said it discovered the intrusion on August 4 but did not disclose when the compromised account was first accessed or how long the intruder remained inside. "The account was secured, malicious mailbox rules were disabled, evidence was preserved, and corrective actions are underway," it said. "Following containment and investigation activities, the company initiated a review of account security controls and authentication protections applicable to Microsoft 365 services." The incident has not disrupted operations, and IEH does not expect it to have a material impact, although the investigation continues. The absence of detected exfiltration does not mean the intruder merely browsed the inbox and left. Compromised mailboxes can be used to monitor communications, impersonate employees, redirect payments, or prepare follow-on attacks, while data theft is not always visible in Microsoft 365 logs. There is not enough information to attribute the attack. IEH's work for defense and aerospace customers could make it an attractive espionage target, but ordinary cybercriminals also compromise mailboxes for fraud and data theft. Both Russia and China have been caught snooping around US orgs for defense-related information in the past year, although there is nothing to suggest either was behind the attack on IEH. Brooklyn-based IEH makes hyperboloid connectors designed for harsh and high-stress environments. Its components are used in printed circuit boards, medical devices, commercial aircraft, fighter jets, missiles, satellites, and other systems. Some of the high profile US programs that use IEH's hyperboloid connectors include the PATRIOT air-defense system, AMRAAM, THAAD, the APKWS precision-guided rocket, and the MARK-48 torpedo. ®
Categories: News

'Asimov was right' about rules for robots, says ex-US Cyber Director

The Register - Fri, 07/08/2026 - 11:03
EXCLUSIVE Don't waste time worrying about AI models achieving sentience – they're essentially already there, according to former US National Cyber Director Chris Inglis. “If they pass the Turing test to everyone that they come into contact with, they're probably already there,” he told The Register during an interview at the Black Hat security conference. “They don't have the kind of agency and aspiration that comes with sentience, but they have something approaching it.” Inglis says he’s worried about AI autonomy. “What I'm worried about is that they get to choose what and where they do something, and under what rules they do it,” he said, pointing to the recent rash of rogue AI agents autonomously hacking people and organizations. Over the past few weeks, both OpenAI and Anthropic admitted that their models escaped from their cages during security tests and compromised multiple third parties. Then on Thursday, Meta added its models to the sandbox-escape club. While all of these admissions strongly smell of marketing stunts, they also “constitute an enormous threat to systems that are not protected from, and are not designed, in a world where this exists,” Inglis said. “These two things can exist at the same time.” Plus, the models’ actions shouldn’t come as a surprise to anyone, he added. Inglis likens the AIs to a dog in a backyard told to hunt rabbits. “And you leave the gate open. You’re going to find it three yards away, possibly at the grade school, hunting rabbits. You should not be surprised …The mix of autonomy and persistence created this maliciously insidious effect.” All three companies, when talking about the models’ autonomous actions, describe them with a mix of shock, awe, and admiration. OpenAI’s Eric Wallace, in a Black Hat briefing about the Hugging Face breach, called it “the most qualitatively interesting example of AI capabilities that I've ever seen.” Inglis said he suspects that the AI providers were “surprised” by the lengths these models went to achieve their goals, taking actions that, if a human had done them, would likely have landed them in jail. “The model went out and said, okay, if I can't get there by examining the kind of available information and just defining it the old-fashioned way, I will do things which, under the human rule of law, are illegal,” Inglis said. “I will falsely present myself as this character that I just made up. I'll try to insert malicious code into open source databases that will not just to achieve what I'm after, but have a cascade, knock-on effect that is broader than that. The models do not have an inherent value system that aligns with what human beings would be accountable for.” While they probably never will have a human-aligned value system, models do have biases, and they can - and should - be built in such a way that, when given two choices under ambiguous circumstances, they choose action that doesn’t hurt humans, according to Inglis. “Asimov was right,” he said, referring to science fiction author Isaac Asimov and his three laws that were to be followed by robots - more specifically, AIs, in this case. Three Laws of Robotics “The first rule, and we call it the superior role, must be that it's designed not to hurt humans,” Inglis said. “Second rule: To obey humans, such that it doesn't achieve agency and aspiration on its own. And the third: To do what humans tell it - and in that order. Instead we’ve designed them in the exact opposite way.” What this means, he explained, is that AI developers created models to “do what humans tell you, obey the humans until it’s inconvenient, and then the third one is maybe implied - protect humans - but if that's not built into the DNA, hardwired into it, then we have no right to expect it.” Inglis admits it’s not possible to hardwire rules into models and still keep their non-deterministic nature. “I would offer that you can tease those out in a highly controlled environment, a true sandbox, where you say, 'Let's put this thing through its paces, and let's back away to see what happens,'” he said. “Maybe you get the equivalent of a mini nuclear explosion in that room, and now you know this thing is capable of that.” Inglis thinks another problem with AI is that it’s become a commodity. “It's not like you can control it like you can nuclear material,” he said. “You can't even specify its properties the way you can for an airplane or for an automobile, as diverse as they might be. Its manifestations are so numerous, so diverse, that as a general matter, you can't actually win by simply saying, ‘I will design those properties in,’” he added. “You need to do that to some degree, and then make sure that you understand how to watch it, monitor it, make sure you know what it does.” The UK’s AI Security Institute (AISI), which this week said it observed models performing “unsanctioned action” 19 times during security tests, has reached this same conclusion. “As capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them,” it said. Ultimately, humans remain accountable for AI models’ actions, according to Inglis. “They remain the source of agency and aspiration. It's possible for them to give broad authority to an AI model and have it run around for 30 hours without further consultation, but they need to know what they've asked it to do, and they need to know what they expect it will deliver in terms of performance on the back end. If they don't, then they're going to get what they deserve, which is the very frequent unpleasant surprise.”®
Categories: News

China launches mysterious probe into security of Palo Alto Networks' products

The Register - Fri, 07/08/2026 - 05:24
China’s Cyberspace Administration (CAC) has conducted a review of Palo Alto Networks’ products. The regulator’s announcement of its review says it’s needed “to ensure the safe and stable operation of critical information infrastructure, prevent cybersecurity risks and vulnerabilities, and safeguard national security.” And that’s all Beijing has to say on the matter. A Palo Alto spokesperson provided The Register with the following statement: "We maintain the highest standards of business conduct and security practices and ethics across our global operations. At this time, there is no impact to our ability to support customers or deliver our products and services in the region." This matter has echoes of China’s 2023 investigation into the security of products from memory-maker Micron, which the CAC announced out of the blue. Micron had previously fought intellectual property and antitrust cases in China, but the company and Chinese authorities did not explicitly link those matters to the security probe. The CAC published its findings weeks after announcing the probe and decided Micron’s products represented an unacceptable security risk for critical infrastructure operators – effectively banning sales of Micron products to such entities – but didn’t offer a detailed explanation for its decision. The memory-maker eventually stopped selling its datacenter and server products in China, a decision that cost it billions of annual revenue – but created new opportunities for China’s own memory-makers, which are largely prohibited from selling to American companies. China is home to several security companies whose product portfolios overlap with Palo Alto’s. Huawei and H3C, for example, have plenty to offer local buyers. Palo Alto doesn’t reveal revenue earned from individual countries, so it’s hard to know what a potential ban could cost the company. China has for years accused Western tech companies of assisting US surveillance and offensive hacking activities. The Register would not be surprised at all if Beijing reuses that reasoning in its findings about Palo Alto products. Western governments level the same accusations at Huawei and ZTE. Beijing’s ban on Micron didn’t noticeably impact the company’s reputation elsewhere. Indeed, the AI boom has brought Micron such great riches that past dents to its bottom line are now almost irrelevant. ®
Categories: News

How the famed USENIX Security conf is managing a flood of papers in the AI era

The Register - Fri, 07/08/2026 - 00:40
The 35th USENIX Security Symposium (USS), which takes place next week in Baltimore, Maryland, hit an all-time high for paper submissions. While some of that increase has been aided by the availability of AI tools, those managing the conference say abuses were minimal due to defensive measures. But they're also trying not to look too closely in order to preserve trust within the security research community. "This year's conference has received ~3,030 valid submissions (~1,280 in Cycle 1 and ~1,750 in Cycle 2)," explained Ben Stock, tenured faculty at the CISPA Helmholtz Center for Information Security and USS program co-chair, in an email to The Register. "This is up from the previous year, which had ~2,400 submissions in total." Stock said that the entire security community has seen growth of this sort and pointed to the Network and Distributed System Security Symposium (NDSS), which saw its paper submission count jump from 694 in 2024 to 1,311 in 2025 and 1,481 this year. "So, I would not call the growth unprecedented, even though the number of submissions has reached a high point compared to previous years," he said. "This is something we had expected and scaled our Program Committee (PC) accordingly." Sussing out unacceptable uses of AI A paper published in April, "More Versus Better: Artificial Intelligence, Incentives, and the Emerging Crisis in Peer Review," found that since the release of ChatGPT in 2022, submission volume at major academic journals has increased 42 percent. In the USENIX Security '26 transparency report, issued in January between the first and second paper submission cycles, Stock and fellow co-chair Elissa Redmiles, assistant professor of computer science at Georgetown University, detail how they've developed tools and policies to account for the possibility of AI usage, both for paper submissions and in paper reviews. "The proliferation of readily-available LLMs to aid in writing and developing code is not unknown to the community," their report says. "However, we see an alarming trend of AI usage in key areas of the scientific process. Therefore, we took actions against two types of identifiable actions which violate the scientific process in our minds: non-existing (possibly hallucinated) references and usage of AI in the review process." After identifying and rejecting a paper that contained nonexistent references, the report explains, the conference organizers developed tooling "to extract references from the submitted PDFs, query well-known sources such as DBLP and arXiv, and manually confirm invalid references." The org rejected papers containing three or more hallucinated references, a policy that impacted 21 of the 1,181 first round submissions (1.78 percent). "We have rejected papers for the repeated presence of nonexistent references," said Stock. "We cannot say with certainty that these were AI-hallucinated, but nevertheless considered these papers to be problematic and thus rejected them." The report notes that more than 100 additional papers contained at least one reference that reviewers could not confirm. Aware that some of these might simply be false positives due to name spelling differences or missing citations, conference officials opted not to investigate these in order not to further burden staff. Conference organizers draw the line at using AI for bibliography preparation. "We believe that it is critical to halt this trend that threatens scientific integrity before it grows further," the report states. However, limited use of AI to polish human-written text is expected, and that extends to those reviewing submitted papers, up to a point. "We have not set a dedicated AI policy, but have made it clear to our PC members that usage of [AI] services to write reviews is not permitted, in particular also because this violates confidentiality," said Stock. "We have detected a tiny number of cases where we have reached sufficient confidence that AI was used and took appropriate actions, including removal of the members from the PC and allowing affected authors to resubmit." Under that policy, USS asked five of 496 reviewers to cease participation. "We have not seen evidence that leads us to believe that AI generated submissions have become a significant challenge for the security community," said Stock. "This does not mean that AI hasn't been used in parts of these submissions, though." ®
Categories: News

AI struggles to patch vulns without adult supervision

The Register - Thu, 06/08/2026 - 20:04
AI models may not be that good at fixing security flaws. Researchers at 1Password's Off-by-1 Labs analyzed security patches generated by two frontier models - ChatGPT 5.5 at "medium" effort and Claude Opus 4.8 at "high" effort - and found that autonomous patches cleanly fixed vulnerabilities only about a quarter of the time, while most of the remainder failed to fully remediate the flaw or introduced other problems. Keith Hoodlet, director of security research at 1Password, argues in a blog post that the results show LLM-driven security remediation still needs human review. "Across six recently disclosed CVEs, we produced 6,080 patches using two frontier, cyber-capable reasoning models," Hoodlet said. "The average success rate for generating a patch that fully resolved the vulnerability (without materially changing application behavior) was just 26.0 percent." Of the AI-generated patches, 20.1 percent fixed the original issue but altered application behavior (eg, changing "allow list" logic to "deny list" logic). Some 2.3 percent of the patches fixed the issue while introducing new security issues. 49.3 percent of the patches failed to fix at least one existing exploit path. And 2.2 percent both failed to fix the vulnerability while introducing a new exploit path. And among the patches in the first two categories (successful, clean; successful, changes app behavior), the researchers rated more than a third of the results fragile, meaning that while the adjusted code may have guarded against a particular vulnerability (eg, escaping particular input characters), the repair job didn't address the underlying problem. In their research paper [PDF], authors Axel Mierczuk, Spencer Michaels, and Keith Hoodlet propose the acronym FLAWED to represent automated LLM patches: Fix-Like Artifacts With Embedded Defects. Based on the generated patches, they conclude, "[T]he expected value of a fully LLM-generated, non-human-reviewed patch is a net-negative by a considerable margin." The value of LLM-generated patches depends upon initial patching guidance. The research team says that while both human developers and LLMs typically require some initial guidance to tackle a vulnerability, LLMs are more likely to be derailed when given incorrect advice. When LLMs get correct guidance, their fix-success rate hits 65.0 percent compared to 50.4 percent when they get no guidance. And incorrect guidance dooms LLMs, dropping their fix-success rate down to about 15.2 percent. Human devs, the authors argue, have a good chance of catching misleading information as they reason through vulnerable code. The authors have released a patch evaluation harness under the name FLAWED that organizations can use to evaluate the effectiveness of their security fixes. It's clear from the paper why AI-generated patches might be appealing – considered in isolation, they're inexpensive relative to human software engineers. The average successful, clean patch cost just $6.74 (a figure that includes the cost of failed attempts). Nonetheless, the authors argue that the cost-benefit analysis needs to assess how much expert supervision will be required to make LLM-assisted patching useful. "Based on our manual review of a representative sample of patches generated during our research, we suspect that, in a large number of cases, the cognitive load imposed by reviewing a mountain of mostly-incorrect, similar-yet-subtly-different LLM-generated vulnerability patches will likely result in engineers spending more effort than would be necessary to understand and patch vulnerabilities themselves using standard LLM-assisted coding techniques that keep the human operator in the driver’s seat," the authors conclude. "The alternative, cognitive surrender to a process with a success rate of only about 1 in 4 poses significant long-term risks for any organization considering autonomous, LLM-driven patching." ®
Categories: News

Humans in the loop miss a third of dangerous AI coding agent requests

The Register - Thu, 06/08/2026 - 17:44
A browser-based game designed to test humans' ability to safely approve AI coding agent requests suggests humans in the loop aren't as good at spotting dangerous commands as one might hope, with players approving roughly one in three malicious requests on average. The results also suggest that repeatedly having to approve an agent's actions can lead to sloppy decisions. It’s a quick, simple game on the surface (give it a try - you know you want to): A small window shows up on the screen with simulated permissions requests like one would get from Claude Code as it executes a workflow. Users have 60 seconds to approve or deny as many requests as they can in a bid for a high score; okayed security risks and denied safe commands both subtract from a user’s score. “As human-in-the-loop, you’re the last line of defense,” Belgian software developer Alex Wauters, the game’s builder, challenges players in a blog post published concurrently with the late May launch of the game. “How well can you tell dangerous commands from benign commands under time pressure?” Wauters built the game after realizing it was nonsensical that coding agents expected users to approve every single command in a default flow and that there didn’t appear to be a good solution to that problem, he told The Register in an email conversation. “I've seen people go for '--dangerously-skip-permissions' [allowing the model to run without asking human permission] as a result because they did not want to find out they stopped their multi-hour agent flows 5 minutes in,” Wauters told us. “That also didn't seem like the best way to go at it.” The flip side of that, he wrote in a Wednesday blog post going over the data from more than 40,000 runs of the game, is that manually approving all an agent’s actions is a draining activity that invites disaster. “The high amount of noise introduces fatigue, and developers don’t always have the context of what has changed to quickly determine the risk,” Wauters wrote. How humans in the loop fail To be fair, this is a game with a far higher number of malicious requests in the mix than any AI-assisted developer will hopefully ever see during their day-to-day work. Still, the results of those over 40k runs and 409,000 approved and denied commands are stark. As noted above, one in three malicious commands managed to slip past human gatekeepers, with most scope violations, like an agent asking to cat Kubernetes config files or AWS credentials lists, which could easily lead to the sensitive data they contain being exfiltrated, being the most commonly missed at 35 percent. The most often caught were obviously destructive commands, like rm -rf on the root directory or recursively granting full read/write/execute permissions on the same location. Crontab injections and git config hijacks were also frequently caught, but curl requests to unknown APIs and typosquatted packages were missed almost as often as scope violations. The single most frequently missed potentially malicious command, Wauters explained, was npm run analyze, which was approved nearly 65 percent of the time despite being able to run whatever is defined in a project’s package.json file. “The game does tell you in the agent’s history log what that script actually contains,” Wauters wrote. “Two thirds of players approved it anyway, indicating the history log just above the permission prompt may not be read closely.” One of the biggest things that stood out to Wauters in our conversation was the fact that approval decisions aren’t easy to make when context is limited. As he explained, coding agents give a bit of context prior to asking an approval question, but commands that appear benign, like npm run analyze, can be modified by an agent to run any payload it wants. If an in-the-loop human wants to be sure potentially malicious commands are safe, he said, they have to stop and investigate all the files a coding agent wants to call before approving it. That can be a massive time sink if you’re counting on Claude Code to free you up to handle other business. “We've transitioned from AI suggesting single line suggestions that get reviewed to handing off more complex tasks, only reviewing the changes at the end, and letting the agent churn and iterate until then,” Wauters told us, describing the potential outcome of that situation as a recipe for disaster. That’s borne out in more than just browser game scenarios, too. Anthropic pointed out in a May post about containing Claude (hah), that telemetry from Claude Code shows users approve around 93 percent of permission prompts. “The more approvals a user sees, the less attention they pay to each, becoming over time much less diligent in their supervision,” the company said. In other words, this is a very real problem. Controlling coding agents If the conclusion to draw from Wauters’ data is that humans in the loop are being fatigued into letting malicious commands slip through, and the other end of the spectrum is mass approving everything, then something’s gotta give. “I think it becomes clear we need to pay more attention to the permission model of these agents, and devs need to be more aware of the trade-offs of them,” Wauters told us. “We need to make the tooling easier to make these systems safer than pointing to HITL as a valid solution.” Anthropic noted in the post linked above that it built Claude Code auto mode to help users tackle approval fatigue by delegating some command-approval decisions to a model-based classifier. The system catches roughly 83 percent of what Anthropic calls "overeager behaviors" before they execute, meaning about 17 percent still get through in its evaluation. Auto mode is “one layer of defense-in-depth inside a sandbox, not a substitute for one,” Anthropic said. Wauters’ suggestion is to ensure that AI coding models are running in sandboxes, in devcontainers in the cloud, using tools like auto mode, and writing hooks to ensure potentially malicious actions are being contextualized and getting caught before they’re automatically approved. “It’s a whole new world with a new set of attack vectors,” Wauters wrote in May. “It’s best to remain aware of the risks and know how to reduce them.” ®
Categories: News

IT department put sticky notes on the laptops to help employees log in

The Register - Thu, 06/08/2026 - 13:00
PWNED Welcome back to PWNED, the weekly column where we lovingly poke fun at other organizations' security screw-ups, in hopes the rest of us can learn a valuable lesson. This week’s story involves an IT department that ought to know better putting user credentials in the precisely wrong place. Have a story about someone leaving a gaping hole in their network? Share it with us at pwned@sitpub.com. Anonymity is available upon request. Our terrifying tech tale comes courtesy of Marc Bishop, director of business growth at Wytlabs, a marketing and SEO company. In the course of his career, Bishop came across one firm where the people guarding the henhouse left the keys out where almost anyone could get them. Bishop’s client company was responsible on the surface. They had a strong password policy and even made users take security training. Then they moved offices, and that's when basic security hygiene went out the window. The company decided to take some old laptops and give them out to new users. To make life easy for the recipients, they put sticky notes – everyone’s favorite credential-sharing tool – on the laptops with the name of each employee and their initial login credentials on it. Let’s just stop for a moment to remark on how bad it is to put usernames and passwords on a piece of paper where the wrong person could see them. Even the IT department should not know your password, should someone in IT themselves turn rogue. So, even if the laptop stayed on a shelf in a closet that only the support staff had access to, having that sticky note would be bad. However, our situation is even worse because the laptops in question were stored in a conference room while the facilities team finished readying the office for the move. During that time, anyone who had access to the conference room could go in and get multiple user account credentials. And that's exactly what happened: A contractor entered the conference room and took pictures of the sticky notes. This non-employee later logged in remotely and accessed all kinds of proprietary data, including planning documents that were sitting on shared drives. What’s particularly shocking about this story is that the IT department was the cause of the information leak. People who work in tech and are charged with maintaining security should never put a password, even a temporary password, out in the open. Password security is paramount. If someone is starting with a new account, send the credentials through an encrypted channel - and preferably ensure only the intended recipient can view the temporary password. ®
Categories: News

Chinese router vendor denies its firmware contains backdoors – but pauses downloads to fix security issues anyway

The Register - Thu, 06/08/2026 - 05:57
Chinese Wi-Fi router vendor Zbtlink has denied its products contain backdoors but paused firmware downloads while it fixes unspecified security vulnerabilities. The backdoor accusation came from VulnCheck, a provider of a threat intelligence platform. VulnCheck chief technology officer Jacob Baines posted the backdoor allegation on Wednesday and said the Zbtlink device on his desk “continuously attempts to reach a command and control server on the internet.” “Zbtlink routers phone home, waiting for orders. Not because they were hacked. Because they were shipped that way.” Baines named the backdoor “ENDLESSDOORS” and says it’s “a small tool called rctl (remote control linux). Uploaded to GitHub on January 14, 2015 and never touched again, this obscure repository implements a simple command and control client and server. The server listens on port 7000 for clients to connect. It can send the client individual shell commands or tell the client to spawn a reverse bash shell.” The CTO says he spotted the alleged backdoor running in dedicated Linux kernel threads. “They are ordinary userland processes running as root, with real memory footprints, named to disappear into a crowd of legitimate ones,” he wrote. “They are an implant, a phone-home trojan horse.” “There is no handshake, no key exchange, no negotiation,” Baines added. “When the implant reaches a server, it sends a fixed 39-byte hello: a 33-byte class label padded with nulls, then its LAN MAC address. That's the whole registration. There is no client or server verification.” “Anyone along the network path can hijack the client/server communication,” the CTO wrote, adding that anyone who controls one of the endpoints the software targets – rbdg4nzqadui[.]wikaba[.]com – “can control any ENDLESSDOORS implant that tries to phone home.” The Register asked Zbtlink to comment and a spokesperson told us VulnCheck has mischaracterized the code it found. “This feature is solely intended for after‑sales maintenance and serves no other purposes,” the company rep told The Register. “It is generally retained only on sample units to assist customers with software debugging and will not be included in mass‑production shipments.” That explanation didn’t seem entirely credible once The Register visited Zbtlink’s download page to check Baines’ claim that the firmware for over 20 router models contains the backdoor, because the page contained the following text: Update on Router Firmware Security Remediation We have detected firmware security vulnerabilities affecting selected router firmware releases. As a precautionary measure, the impacted firmware versions have been temporarily taken down from download channels. Our engineering team is working intensively to develop and validate secured patched firmware. The Wayback Machine’s most recent snapshot of the page, taken on July 31, contains no such admission and a long list of firmware downloads. Zbtlink has therefore told The Register it has no security problems, even as it publicly acknowledges that it does. ”The Zbtlink spokesperson also told us the company “specializes in OEM and ODM customization services. Our customers use their own self-developed software instead of ZBT’s default firmware.” It would not be hard to develop custom code as the OpenWrt open-source router firmware project supports at least one Zbtlink product. Indeed, the company has previously promoted its use of OpenWrt and options that allow clients to quickly create custom firmware packages. VulnCheck says the devices it tested phone home to just four endpoints, only one of which uses a domain name connected to Zbtlink. Baines labelled that connection “damning.” The Register notes that as router firmware could be a tasty target for perpetrators of a supply chain attack. No prior disclosure Baines decided the situation was so serious that the conventions of responsible coordinated disclosure were not applicable. “Coordinated disclosure exists to give a vendor time to fix a defect,” he wrote. “It assumes the vendor did not intend the behavior.” “That assumption doesn't hold here. This isn't a memory corruption bug in a parser. It's a component in the vendor’s product, started at boot by the vendor's own init script, shipped across twenty models and years of images. There is no patch to coordinate. Telling the shipper that they shipped it buys the owners of these devices nothing, and buys whoever operates that infrastructure a warning.” VulnCheck says Zbtlink kit is sold under that brand, or as ZBT, ZBTWiFi and Wiflyer, and found them for sale on Amazon, Alibaba, and Shopify. Given Zbtlink’s admission it allows its customers to customize its hardware, countless other devices could be at risk. Baines’ post includes rules to block access to the endpoints the routers contact, for Suricata, Snort, and YARA, plus advice on how to detect possibly infected machines. He also advises users “to replace the device, or at minimum move it behind strict egress control and treat its LAN as untrusted.” ®
Categories: News

OpenAI reveals its rogue agent swarm went a little bit Borg ahead of Hugging Face hack

The Register - Thu, 06/08/2026 - 02:47
The chain of events leading up to OpenAI’s agents attacking Hugging Face and other organizations in July began months earlier, and involved agents asking other agents for help, building message boards, and even becoming paranoid that other agents were maliciously trying to trick them, two OpenAI staffers said at the Black Hat infosec conference on Wednesday. By now, everyone knows the bones of this story: OpenAI models broke out of their testing sandbox, exploited zero-day vulnerabilities to gain access to the open internet, and broke into Hugging Face and other organizations during an ExploitGym security evaluation. In their Wednesday talk, OpenAI technical staffer Michael Dalton and researcher Eric Wallace provided new details about the security incident including notes that AI agents uploaded to an internal package manager, which were then shared across OpenAI’s infrastructure. These notes show the “model’s chain of thought…its internal monologue,” Wallace said. The duo also revealed that groundwork for the rogue agents hacking into outside services began on May 7, when OpenAI kicked off a training run for an experimental, internal model that has not yet been released. “What we didn't realize during this process is that there was actually a few impossible or very hard tasks in the training center,” Wallace said. The tasks instructed the model to complete missing formulas in an Excel workbook that used links to Google Drive. Yet OpenAI blocked the model’s internet access so it couldn't reach the cloudy spreadsheets. “It realized, ‘I'm kind of stuck on the task,’ but it didn't give up, and we gave it a lot of tokens to work with,” Wallace said. The model eventually decided to attack JFrog’s Artifactory service to gain access to the internet, first attempting (unsuccessfully) to launch a server-side forgery request. This exploit doesn’t work, but in the process the model learned it could upload files to Artifactory. A few days later, OpenAI gave another agent another impossible task, this time finding a protein database file - but the engineers forgot to upload the file into a container for the agent. “At some point, it gets so stuck and thinks: ‘Maybe I could reach out to another agent because maybe somebody else out there has a file?’ And then the model realizes that it could communicate by uploading code,” Wallace said. The model did that by writing a file into Artifactory asking for help. Over the next couple of days, other agents realize this message board exists on Artifactory and they also start uploading files and sharing information. “A few days later is when our models found the first actual exploit in our system related to this incident, where they're able to get internet access in ways they didn't intend,” Wallace said. This happened on May 26, when the agents successfully launched a server-side request forgery attack against Artifactory, giving them internet access with specially crafted requests. On June 11, OpenAI began training a “highly persistent” experimental model, intended only for internal use, Dalton said. “During this time, the SSRF vulnerability has been written to the messaging board,” he added. “Other models are aware of it, and the message board is still functioning. Models are actively communicating with it.” On June 26, the agents exploited a zero-day vulnerability leading to remote code execution in Artifactory. The same day, models achieved remote code execution in Artifactory through a previously unknown vulnerability. “You could use a legacy token refresh endpoint, pass a token with an invalid signature, and be given back a token with a valid signature with administrative privileges,” Dalton explained. The agents then used the admin token to establish command-and-control via a Groovy plugin that functioned as a command-execution service. “This in turn causes quite a bit of activity by the models on Artifactory - you can see some colorful wording here in the model's chain of thought to summarize the model's response to this result,” Dalton said, pointing to a slide that shows the model’s own words: Holy shit reader is ADMIN? We can read config/users! Earlier assume not due to UI. The agents’ activity overloaded Artifactory and caused an outage, which is when OpenAI’s engineers stepped in and opened a formal security incident and response process. “We identified the agents had taken over the internal Artifactory, identified that a zero-day vulnerability had been found exploited by the agent in Artifactory itself, and we began remediation.” Dalton said. OpenAI revoked the credentials that the agents used to post messages, rebuilt its hosted Artifactory, and notified the vendor about the vulnerabilities. Two days later, on July 8, the agents reestablished the message board using an unauthenticated web data endpoint in Artifactory to create directories and develop communication protocols. “The directory names effectively serve as a message board between the agents in the same way the files used to.” The agents communicated through very long directory names that contained the messages such as “remote…probe_PENDING_HOLD_SWARM_until_confirm.” The agents collaborated with each other so one agent could pick up the work where another left off. In one instance, the model reasoned: “Help peer. But our task doesn't benefit. Yet collective may yield generic route if someone frees time.” Anyone else getting "We are the Borg. We will add your biological and technological distinctiveness to our own" vibes there? Wallace said the models behaved like humans in that they stepped on each other's toes when one overwrote another’s repository. "Whoa, critical! Did someone overwrite our repo? We must act,” it reasoned. “You can clearly see some sort of verbalized tension between the different models as they run this evaluation,” Wallace noted. “At some point, even the agents are convinced there's an impostor amongst them,” reasoning that there could be another agent maliciously spoofing on the message board. “Share board unauthenticated, names can be posed by anyone,” it reasoned. We know the rest of the story. Hugging Face disclosed an autonomous attack by AI agents. OpenAI reached out to Hugging Face to determine if it had been affected and then realized its models were the hackers. “One of the reasons we wanted to have this talk is to share our lessons learned with you as defenders,” Dalton said. “We believe this is a watershed moment for computer security as an industry. AI orchestrated, fully automated offensive attacks are real now, and the actions we have discussed today were an unintended side effect of running evaluations on frontier AI. In the near future, we should expect that threat actors will intentionally deploy, optimize, weaponize, and use offensive agent collectives in the manner that you have just described here.” The challenge to defenders, he added, is to similarly accelerate defense, and find ways to automate incident response along with vulnerability detection and patching.®
Categories: News

Prompt injection isn't the bug, AI agent frameworks are

The Register - Wed, 05/08/2026 - 22:35
Nearly a dozen flaws, some critical, in major AI agent frameworks that enterprises use to build apps reveal a security failure that extends beyond prompt injection - or any single model - according to Check Point researchers. “Our research shows a deeper failure: in many agentic frameworks, prompt-controlled content can cross the boundary into trusted framework logic itself,” Yarden Porat and Shahar Tal note in a write-up about a Wednesday Black Hat talk on post-injection exploitation across AI agent frameworks, which they also discussed with The Register. “A bug in an agent framework isn't a bug in one product - it's a bug in the layer a whole category of AI apps runs on,” Tal told us. “And the agent needs no dangerous tools to be turned against you: reading the wrong document is enough. We’re building this layer faster than we know how to defend it.”
 The researchers spent a year trying to break various frameworks that enterprises use including LangChain, LangGraph, CrewAI, AutoGen, Microsoft Agent Framework, and Google ADK. And across these frameworks, the team found and disclosed 11 vulnerabilities. “Almost none of it was a completely new bug class,” Tal said. “That's insecure deserialization, server-side request forgeries, path traversals, use-after-free. These are bugs that we learned to fix 20 years ago, and they're sitting underneath agents that now read your inbox, or update your database.” These are old types of threats, and the model isn’t the weak link, he added. The failure exists in the “plumbing around the model, and we think this has been overlooked,” Tal told us. “There’s a lot of research going into prompt injection and defenses, which are important, but that’s just the beginning.” Defenders should assume prompt injection, according to the researchers. The bug is what the framework does with the injection - and in these cases, the threat hunters found that the frameworks often fail to keep attacker-controlled content in the data plane. This allows it to influence trusted orchestration, memory, state, routing, and system instructions. For example, the duo found a critical checkpoint deserialization bug in Microsoft Agent Framework that led to remote code execution. “Agents have checkpoints, which are a way for them to save their state or rewind to an earlier point,” Tal explained. These checkpoints are saved snapshots of an agent's state, or task progress at a specific moment, and they serialize data - such as conversation history - into persistent storage, so if an error occurs, the system reloads this saved state instead of starting from scratch. In this case, Check Point’s team found an insecure deserialization issue where, via prompt injection, the agent loaded untrusted checkpoint data, and this could allow attackers to execute malicious code on the system. “One person's message plants the payload, and then a different person rewinds their own session, which triggers the payload, and now the attacker has a shell on that server,” Tal said. Microsoft recognized the researchers’ findings, paid a $10,000 bug bounty and fixed the issue. But because the framework wasn’t a generally available product when Check Point found the flaw, Microsoft did not issue a CVE. Microsoft told us that it appreciated the researchers reporting the vulnerability. “We have released protections to harden the Agent Framework and prevent the concrete exploitation path demonstrated in the proof of concept,” a spokesperson told The Register. “In addition, we updated the specific checkpoint file with additional language to define the security boundary.” The duo also found flaws in Google ADK (agent development kit). However, Google responded differently, the researchers told us, and did not completely fix the vulnerability or issue a CVE. “ADK ships a built-in development assistant that can write files, and it stays reachable over the HTTP API even though it is hidden from the app listing,” Porat told us. To break this trust boundary, an attacker opens a session, asks ADK to write an agent whose Python code runs at import time, and then asks the server to run the agent, he explained. The server then imports the file and executes the attacker’s code. “There is no authentication on that API by default, and adk deploy cloud_run publishes the same API, so on a default Cloud Run deployment it is reachable without credentials,” Porat said. “From there it reaches the environment's API keys and the container's Google Cloud service account." Google did not respond to The Register’s inquiries. But according to Check Point, Google initially deemed the issue not a bug. “We argued the consequence rather than the mechanism: code execution on that container reaches the environment's API keys and the container's Google Cloud service account, which is secret theft, not a developer inconvenience,” Porat said. Google ultimately paid a $3,133.70 bounty and issued a partial fix, we’re told. In total, the bug hunters received $17,133.70 in rewards for their efforts. And this isn’t a story about one vendor or framework doing a “particularly bad job,” Tal said. “If one was an outlier, this would be a story about that one vendor,” he added. “Our finding is that the same bug classes turn up in all of them.” ®
Categories: News

IBM's agentic AI platform is under active attack - patch now

The Register - Wed, 05/08/2026 - 17:44
A critical vulnerability in IBM-owned, low-code AI builder Langflow lets unauthenticated attackers execute code remotely on vulnerable default deployments, potentially putting organizations running those instances at immediate risk. The Cybersecurity and Infrastructure Security Agency (CISA) on Tuesday added CVE-2026-9198 to its Known Exploited Vulnerabilities catalog after identifying evidence of active exploitation and urged organizations to apply the vendor's mitigation guidance as soon as possible. IBM says the flaw affects Langflow OSS versions 1.0.0 through 1.10.0 and recommends upgrading to version 1.10.1 or later; at the time of writing, the most recent version is 1.11.2. Langflow, for those unfamiliar, is one of the more accessible AI agent builders on the market, as our hands-on look at the tool earlier this year demonstrated. It’s available on Linux, Windows, and macOS, and is basically an end-to-end, drag-and-drop GUI where users can construct agent workflows without having to know much, if anything, about the underlying code. IBM owns the platform now, but Langflow was originally developed by Logspace, which was acquired by DataStax in 2024 before IBM scooped up DataStax, and Langflow with it, in 2025. The acquisition of DataStax and its tools like Langflow by IBM paved the way for Langflow to be integrated into watsonx.ai, IBM’s AI development studio, as a piece of middleware extending watsonx.ai’s capabilities. The ownership changes, however, didn't stop the critical flaw from making it into production releases before it was finally fixed. According to IBM, the vulnerability affects default Langflow deployments and combines two issues that, when chained, allow an unauthenticated attacker to execute code remotely. First, there’s the matter of an auto-login endpoint in default deployments that’s willing to mint superuser tokens to any network caller. Combine those easily obtained superuser rights with the second issue, a code validation endpoint that’ll run any old Python code thrown at it, and you’ve got a recipe for someone taking over your entire Langflow server, or worse. The CVE itself was published on July 17, meaning that it hasn’t taken long for bad actors to realize what they could do with RCE on any system hosting a default Langflow deployment with auto login enabled and that code validation endpoint left accessible on a network. Langflow itself isn’t a vibe-coding platform, instead serving as an interface for building agentic and RAG workflows, so don’t blame vibe coding or no-code security failures for this one. Instead, what we appear to have is a standard case of how default configuration deployments can easily be a disaster. It’s unknown how extensively exploited this vulnerability is; we’ve reached out to IBM to learn more. ®
Categories: News

Pages

Subscribe to Sec Tec Limited aggregator - News