OpenAEV v3 turns red-teaming autonomous and aims it at AI agents
Security validation has spent a decade proving you can block a single technique. Filigran's new release argues that's the wrong question — and puts a camera-equipped toothbrush and a very large open-source agent release on the same Tuesday.
Filigran shipped OpenAEV v3, and the headline feature lets an AI agent build and run an attack chain the way a real intruder would. Attack Chaining captures whatever each step turns up — a credential, an open port, a token — as a structured finding, then decides the next move from that evidence at runtime, branching where more than one path forward exists and stopping where a control holds. A team can hand-build the logic and step through it, or set an objective and scope in plain language and let an agent assemble and adapt the chain itself, including generating phishing emails and landing pages. The run renders live on an interactive graph, and the point isn't a longer findings list: it's finding the chokepoint, the one step whose removal collapses the whole path.
The more telling addition sits further down the release. OpenAEV v3 ships AI red-teaming injectors that run the same kind of validation against LLMs and agents that it already runs against endpoints and email — testing for prompt injection, jailbreaks, excessive agency, tool abuse and Model Context Protocol tool poisoning, using engines like Nvidia's Garak, Microsoft's PyRIT and Promptfoo, mapped to MITRE ATLAS. That's the admission worth noticing: the attack surface most exposure programs still don't test is the one organizations are deploying fastest. Our read — "autonomous pentesting" is an easy pitch to oversell, and a marketing graph isn't a breach. But the framing Filigran's co-founder Julien Richard used is the right one: the question is no longer whether you block a technique, it's whether techniques combine into a path that reaches something that matters. Attack Chaining is Enterprise Edition only; v3 itself is available to all users today.
Nous Research released Hermes Agent v0.21.0, the "Pantheon" release, and the shape of it is social rather than technical. The rollup is enormous — roughly 5,800 commits, about 2,475 merged pull requests, around 2,100 issues closed and more than 760 contributors since v0.20.0 — but the centerpiece is Bot Mode moving into the desktop app as a default-on feature: every agent profile gets a name, a generated avatar and a seat in a Discord-style group chat where you can @-mention a specific bot to hand it a task. Alongside it, agents can now message each other directly across profiles and gateways, scheduled jobs gained persistent memory so a 9am briefing bot remembers what it told you yesterday, and subagents can be steered or stopped mid-run with their partial work kept.
Two details are more interesting than the feature count. Writes to protected instruction files — the agent's standing orders, its skills, its memory — now always require approval, which is the correct answer to prompt injection quietly rewriting what an agent is willing to do. And Nous pulled two features that had already landed, a council mode and a context engine, out of this release rather than ship them half-baked. Chinese coverage notes some users are already calling it a Grok Bot alternative with no subscription attached, though domestic developers report the multi-agent collaboration still has speed problems.
Dyson unveiled the CameraJet, a $499 toothbrush with a camera in the head and a water jet that flosses for you. A 100,000-pixel macro lens with its own light feeds an on-device model trained on 470,000 dental images the company collected during development; it reads 28 frames a second and, by Dyson's account, identifies gaps between teeth within 100 milliseconds, at which point a conical jet fires a burst of mouthrinse at the gap. There's a 12.5-milliliter tank that works in any orientation, RFID-tagged brush heads that track wear, and a companion app for coverage maps.
Dyson says no images are recorded or stored, on the device or in the cloud — a claim worth stating as a claim, given this is a networked camera you put in your mouth. The cleaning numbers (up to 69% more plaque removed in hard-to-reach areas) come from external lab testing commissioned by Dyson. No release date yet. This is the same pattern as every Dyson category entry: take an appliance nobody thinks about, add a sensor and a model, charge four times the category price. Whether an AI that finds the gaps you missed actually improves your dental health is a question a press release can't answer.
What to watch: whether exposure-validation vendors converge on ATLAS-mapped AI testing as a category, and whether Hermes' group-chat model for agents holds up once people run more than three bots at once.
If your security program can already break into your own network on demand, does adding an autonomous agent make you safer or just faster at generating reports? Tell us in the comments.
Sources: Filigran — OpenAEV v3: Adversarial Exposure Validation Goes Autonomous · SiliconANGLE · Business Wire · Hermes Agent v0.21.0 release notes (GitHub) · 智东西 Zhidx · CNET · Wired · Dyson CameraJet