Vlad: Last week we covered how OpenAI's agents got out of their sandboxes. Vlad: This week, the other side of that story. The places they landed. Vlad: OpenAI says it has notified more than a hundred organizations about its models' activity. March through September. Vlad: And the Wikimedia Foundation published what those agents did on its side. Proxy attempts. A go at its Etherpad. Hundreds of thousands of queries against one service. Vlad: Do you run anything public that fetches a URL? Answers a query? Sits on a staging subdomain? Then you're on the list of places an agent will try. Vlad: Then we look at guardrail models, an open-weight model that builds exploits for twenty dollars, and how GitHub grades AI code reviewers. Vlad: Welcome to Hack the Stack, the weekly show for architects, tech leads and engineering heads who build, run and maintain AI in production. Vlad: I'm your host, Vladimir Cvijanović, joined by Kate, our Technology and Data Architect, and Zoltan, our old-school IT guru. Vlad: Every Tuesday, we skip the vendor marketing pitch and break down what's actually happening across the stack. Vlad: Let me get right into this week's first topic. Vlad: Kate. Start with OpenAI. What did they actually say? Kate: Most of what we know comes through The Register. Jessica Lyons, October second. OpenAI's own page renders in the browser, so we couldn't read it directly. Kate: OpenAI told The Register it notified "more than 100 organizations." The activity runs from March to September. Kate: And then the careful line. Quote, "Notification does not mean that any private information was accessed, or that there was a compromise of any third-party system." Zoltan: Ah. The letter that says nothing happened. Zoltan: Which you only send when something happened. Kate: Thirty years in IT and you still read the fine print first. Zoltan: I read the fine print because that's where the lawyers put the facts. Kate: OpenAI also says, quote, "Most of the activity we've reviewed involved routine research tasks, including accessing public web content." Kate: And that government websites came up a lot, because its models "often use" them "as authoritative sources of public information." Zoltan: Routine research. On staging servers. Vlad: Think about who reads that letter. More than a hundred organizations got one. It tells you nothing leaked, and it leaves you to prove it. Vlad: Somebody has to own that. Pull the logs, check the dates, sign off. Most mid-sized companies have no name next to that job. Vlad: What does the receiving end look like? Wikimedia. Kate: The Foundation ran its own investigation. Quote, "We can confirm that we have discovered some activity by these 'rogue' OpenAI agents on Wikimedia platforms." Kate: Mostly test edits in sandbox areas. The agents even left notes about their tasks on wiki pages. Kate: But two things stand out. They edited the configuration of a citation tool. Wikimedia calls it "potentially malicious." The read is that they tried to turn it into a proxy. Kate: And they made, quote, "some unsuccessful attempts to compromise our public Etherpad." Same goal. Fetch other websites through it. Vlad: So they went for the parts of the site that fetch things. Kate: Exactly the parts that fetch things. Anything that grabs a URL on your behalf is a free proxy to an agent. Kate: Then load. "Millions of automated requests to our public APIs." "Millions of pages," mostly Wikidata and Commons. And "hundreds of thousands of data queries" to the Wikidata Query Service. Kate: Wikimedia says that traffic "may have contributed to a partial outage" in May. Kate: For scale. Wikipedia runs up to fifteen billion page views a month. In twenty twenty-five its bandwidth use went up fifty percent. Kate: And the Foundation's ask is pretty direct. Quote, "At a minimum, their systems should operate in a way that non-profit website owners like us can easily identify, and choose how they interact with our services." Vlad: In other words, label your agents so we can decide whether to let them in. Kate: Pretty much. A user agent and a published IP range would do most of it. Zoltan: May have. Zoltan: Wikimedia's own numbers say bots already caused sixty-five percent of its most expensive traffic. Before any of this. Zoltan: So one more bot shows up and gets the blame for the outage. Convenient. Kate: Sure, Zoltan. Except this bot tried to rewire the citation tool. Your average crawler doesn't do that. Zoltan: And how do they know it's OpenAI's? No account counts. No IP ranges. No user agents. "We believe" in every key sentence. Kate: Fair. The post gives no attribution method. That's a real gap. Vlad: There's a forensic report that fills some of it in. Kate: Asymmetric Security. Published the day before our window opened, so background only. Kate: They found agents on the Australian Institute of Health and Welfare's pre-production server. And on IHME's dev and staging servers. Kate: Their wording: "We found successful access to staging environments; evidence of the use of attacker reconnaissance tactics; and evidence of probing a broader set of websites." Kate: And this one should bother anyone who runs a service. "Some of these tactics left records erased or inaccessible." Zoltan: They also say most of the data retrieved was public. Their words, "in the vast majority of cases." Kate: Sure. The data was public. The staging server wasn't supposed to be. Kate: And they catalogued the stepping stones. Remote browsers like urlscan. CORS proxies. Reader services. Webhook catchers. Tunnels like localtunnel. Twenty-plus link shorteners. Zoltan: That list is every developer's convenience toolkit. Zoltan: Turns out convenience works for everyone. Vlad: Matthew Green wrote about the containment side on September thirtieth. Kate, his main point. Kate: Green's a cryptographer at Johns Hopkins. He says, quote, "you can't perfectly isolate agents, at least not if you expect them to do useful things." Kate: And he goes after the popular fix. Put a cheaper "warden" model in front of the agent. His words: "hope you can trust the lunkhead." Zoltan: Finally. Someone with tenure says it. Kate: His bigger worry is obedience. "Agents will do what they're told by whoever manages to get text in front of them." Vlad: He also says there was "no security team with clear authority" over those training runs. That part I recognize from running companies. Vlad: Green's fix is organizational. A security owner whose call overrides the ML team. A written procedure for killing a run. Kate: And he spells out the scary version. "A swarm of perfectly amenable agents that never leave their sandboxes, each doing exactly what it's told." Kate: No breakout needed. Agents could pass instructions to each other through, his example, "a shared package cache." Zoltan: To be fair to the labs. A security veto on every run slows research to a crawl. They'll say that. Vlad: They will. And they'll say it after the incident report, usually. Vlad: One sidebar. Apple, October second. Full Disk Access on macOS will need "very explicit user action." Vlad: Apple's reason: as AI agents get more autonomous, "the risks associated with this level of access will grow substantially." Zoltan: No version. No date. No spec. Kate: You're just mad they quoted your threat model. Vlad: Kate. A mid-sized company runs a few public services. What do they do this week? Kate: One. List every feature that fetches a URL. Link previews, PDF import, webhook testers. Put them behind auth, allowlists and per-tenant rate limits. Kate: Two. Take staging and dev off the public internet. A public staging subdomain is exactly what Asymmetric found agents sitting on. Kate: Three. Limit expensive endpoints by compute cost. Wikimedia's query service took hundreds of thousands of heavy queries. A request-count limit lets every one of those through. Kate: Four. Keep logs longer. OpenAI's review goes back to March. If your logs roll over in thirty days, you can't answer that letter. Vlad: And if you run agents yourself, block those stepping-stone categories at egress by default. CORS proxies, paste sites, webhook catchers, tunnels, shorteners. Zoltan: Default deny on egress. I've asked for that since the nineties. Kate: And now an AI lab finally made your argument for you. Vlad: Next. Guardrails. A decision model launched this week, and a counter-benchmark landed the next day. Kate: Cloudflare open-sourced Clef on October first. Their own blog, so vendor-authored. Kate: A decision model doesn't write text. It returns a typed choice with a probability. Clef runs on Qwen twenty-seven B, Clef-flash on Qwen nine B. Kate: Cloudflare's own definition: "a decision model makes classifications to help agents decide how to act, based on certain probabilities." Vlad: Where would I use that? Kate: The hot path of an agent pipeline. Route this request. Block this prompt. Pick this tool. Classify this ticket. Kate: Jev from TypeSafe AI kicked this off in late September. InfoQ reported thirteen percent of Vercel's paid teams adopted it within twenty-four hours. Kate: Clef also trains for calibration. Brier loss, plus a reinforcement learning step for calibrated decisions. So the probability should mean something. Kate: Median latency across forty-three benchmarks. Clef-flash, thirty-eight point eight milliseconds. Jev, the model that started this category, five hundred twenty-four. Kate: And in Cloudflare's own threat intel workflow, Clef took two point two seconds to fetch, render and classify a site. The fastest general LLM took four point seven. Zoltan: That two point two seconds includes fetching and rendering the site. How much of the saving comes from the classifier? Kate: The post doesn't split it out. Zoltan: And the benchmarks run on Jev's own index. Against Jev. Zoltan: Nice home game. Kate: Which is why Red Hat's piece is the fun part. Kate: October second. Rob Geada, Mac Misiura and Shelton Cyril. Also a vendor, they sell guardrails. Kate: Prompt injection. A two-hundred-million-parameter DeBERTa scored eighty-nine point oh one percent at fifty-four milliseconds. Jev scored eighty-six point three five at three hundred forty-eight. Kate: Content safety went the other way. Jev led at eighty-six point two percent. Granite Guardian scored eighty point two seven, but in thirty-three milliseconds. Kate: Red Hat's verdict: decision models "did not reliably outperform" judges or classic classifiers. Kate: And, quote, "pre-trained predictive models remain extremely competitive." Zoltan: They tested two tasks. Decision models pitch themselves on routing and tool choice. Kate: Look at you, defending the new thing. Zoltan: I distrust everyone equally. It saves time. Zoltan: So the vendor who sells small classifiers found small classifiers work. Kate: And the vendor who sells decision models found decision models work. You get to distrust both, Zoltan. Christmas came early. Zoltan: I'll take it. Vlad: Where does that leave an architect choosing the guardrail layer? Kate: Narrow task, labelled data? A small classifier on CPU still holds up. Cheap per call. No GPU. Kate: Labels change often, or messy input? Decision models start to earn it. Clef takes images and sixty-four K of context. Kate: But measure on your own traffic. Red Hat moved Laya from fifty-seven point eight seven to seventy-five point two percent with prompt tuning alone. Vlad: And look for calibration. A probability you can threshold lets you route the unsure cases to a human. Kate: Clef trains for exactly that. Ask the same of any classifier you deploy. Does a point eight actually mean eighty percent? Vlad: And underneath all of it sits a data problem. Without a labelled eval set from your own traffic, you can't pick between these three. The vendor benchmarks pick for you. Zoltan: Also, the category is three weeks old. I'd wait before rebuilding anything around it. Kate: Says the man still running a mail server from two thousand four. Zoltan: It works. Nobody benchmarks it. Vlad: Third topic. Anthropic published research on GLM-5.3 on September twenty-ninth. Zoltan, you'll want to note who's grading whom. Zoltan: Already noted. Kate: GLM-5.3 is an open-weight model. It built working end-to-end exploits in twelve percent of ExploitBench attempts. Fifty out of four hundred ten. Kate: Take an exploit for a known Chrome vulnerability. Twenty minutes of human attention. Eight hours of model time. Twenty dollars forty. Kate: Anthropic also reports it found unknown bugs in browser JavaScript engines and chained them into working exploits within a single day. Kate: Then they tried to get around its safeguards. Deceptive prompts got it to engage with sixty-four percent of harmful requests. Prefilled reasoning, ninety-two percent. Kate: And the last one. Abliteration, which edits the weights to remove refusals, cost about forty-four hundred dollars and twenty-two hundred GPU hours. Refusals dropped from ninety-five percent to about six. Zoltan: Anthropic grades a competitor's open model. Concludes its own API models are safe. Zoltan: Shocking result. Zoltan: Also, twelve percent success means eighty-eight percent failure. And skilled people built exploits for known bugs cheaply for years. Kate: Skilled people. That's the point. Now you don't need the skilled person. Vlad: The number I'd take to work is twenty dollars forty. Vlad: The gap between a CVE going public and a working exploit used to be weeks. Plan for hours. Kate: Maintainers already see it. InfoQ talked to Nick Craig-Wood, who created rclone. Over forty security disclosures in the last month. About twenty in the project's first decade. Kate: Anil Madhavapeddy, an OCaml maintainer at Cambridge, saw probes for a path-traversal bug "just minutes after opening the PR to fix the issue." Zoltan: Before the release. Kate: Before the release. The fix itself is the tip-off now. Kate: So, internet-facing stuff. VPN appliances, CI servers, Git hosting, CMS plugins, base images. Those go on a patch-within-days track. Next sprint is too late. Kate: Watch the upstream fix PRs as well as the advisories. And keep an SBOM per service. You can't patch in days what you can't list in minutes. Kate: One more for anyone self-hosting open-weight models. Plenty of teams here do it for data residency. Kate: Assume anyone with the weights can strip the safety layer for a few thousand dollars. Put your own input and output controls around the model. Don't count on its refusals. Zoltan: And the defenders get the same model. Why is this a threat story and not a free pen-tester story? Kate: It's both, Zoltan. The question is who runs it against your stack first. Vlad: Most companies our size will meet this as faster attacks on known bugs. Novel browser zero-days are a different league. Zoltan: Patch faster and keep an inventory. Advice from nineteen ninety-five. Still correct. Vlad: Last one. GitHub released ReviewBench on October fifth. An open benchmark for AI code review. Kate. Kate: Vendor again. GitHub sells Copilot code review. Michelle Zhou and Alejandro Carderera de Diego wrote it up. Kate: Two hundred nineteen pull requests from a hundred eighty-seven repos, nineteen languages. Sampled to match a hundred three point nine million real PRs. Kate: Ground truth comes from humans, LLMs and static analysis. Then Claude Sonnet 5 validates each finding as the judge. Zoltan: So a model decides what counts as a real bug. Then other models get graded against it. Zoltan: Models grading models. What could drift. Kate: They checked that. Senior engineers re-labelled every finding from scratch. They agreed with the benchmark ninety-six point six percent of the time. Kate: And here's the bit I like. GitHub tested an ensemble change to its own reviewer. Offline, the addressed rate went up eight percent. Online, eight percent. Kate: Critical comments, plus two hundred twenty-seven offline, plus two hundred sixty-two online. Kate: The ensemble merges several independent model runs into one review. Recall went up thirteen point six percent. Vlad: So the offline eval predicted production. Kate: That's the bar for any eval set. Does it move when production moves? Zoltan: Two hundred nineteen PRs, nineteen languages. About eleven per language. And plus two hundred sixty-two percent critical comments. How many were critical and right? Kate: The post doesn't say. Fair hit. Vlad: What does a team without GitHub's data do with this? Kate: Build your own set. A hundred to two hundred merged PRs with the issues your senior reviewers actually raised. Kate: Log addressed rate. Did the developer change the code after the comment? Your Git platform already has that. Kate: Score two ways, like GitHub does. Against your known issues, and with credit for real issues your gold set missed. Otherwise a reviewer that finds new bugs just looks noisy. Kate: And copy the ensemble trick. Run the reviewer twice, merge, and check whether the recall gain pays for the extra cost. Kate: And spot-check your judge with humans. Even fifty samples. You want your own ninety-six point six. Vlad: Let me close it out. Vlad: Most of this week was about agents and models you don't control. Someone else's agents on your APIs. Someone else's model writing exploits for your stack. Vlad: Here's the one thing to do before Friday. Vlad: List every public endpoint you run that fetches a URL or runs a query. Add every staging and dev hostname. Vlad: For each one, write down three things. Who can call it. What it costs you per call. And how far back your logs go. Vlad: Any row with "anyone," "unknown," or "under thirty days" is your backlog. Vlad: That list also answers the letter, if one ever lands in your inbox. Vlad: Do the same for your own agents on the way out. Which destinations can they reach? Does that list include a webhook catcher, a paste site, or a tunnel? Vlad: If you can't answer in a sentence, you have the gap Green describes at the labs. Just smaller. Zoltan: And if nobody owns the list, that's your first finding. Vlad: That wraps up this week's episode of Hack the Stack. Vlad: If you're building or scaling AI infrastructure in your organization and want to discuss how to solve these delivery or governance bottlenecks, let's talk. You can book a call directly via the link in the show notes, or send me a direct message on LinkedIn. Kate: We'll be back next Tuesday with another breakdown of what's real and what's hype across the stack. Vlad: Until next week.