A 27B security model from Alibaba beats the frontier labs on CyberGym
A 27-billion-parameter security model that Alibaba fine-tuned itself now sits above Google's and OpenAI's purpose-built cyber models on UC Berkeley's vulnerability benchmark. Separately, the platform vendors started selling the tooling to explain what agents actually do — which is the question the first story is really about.
Alibaba Security's XekRung-1.5-27B-Preview took the No. 1 row on CyberGym's model leaderboard with an 88.9% success rate at reproducing real vulnerabilities — 34.39 points above the 54.51% score of the Qwen3.8-27B base model it was built on, according to the lab's own writeup and the UC Berkeley leaderboard. The entries behind it are the ones the frontier labs built on purpose: Google's Gemini 3.8 Flash Cyber at 86.3%, GPT-5.5-Cyber at 85.6%, Zhipu's GLM-5.3 at 84.5%, DeepSeek-V4-Pro at 83.3%. CyberGym drops an agent into unpatched repositories drawn from 1,507 historical vulnerabilities across 188 large open-source projects and asks it to write a proof-of-concept that actually triggers the flaw, under a restricted-information setting where no fuzzer is installed, no patched diff is visible and network access runs through an allowlist.
The size gap is the point, and so is the caveat. Alibaba attributes the gain to security-specific post-training rather than to a scaffold built for the benchmark: de-identified offense-and-defense trajectories for supervised fine-tuning, then agentic reinforcement learning with build, crash and proof-of-concept validation as the reward signal, and failed attempts recycled into corrective training pairs. But CyberGym's leaderboard is self-reported — every row is submitted by the team that ran it, agent runs are stochastic, and the site itself warns that modest differences may not reflect meaningful capability gaps. There is no independent re-run, and the weights are not public: plain Qwen3.8-27B is apache-2.0, while XekRung returns nothing on Hugging Face.
Nor is 88.9% the top of the benchmark. It is the top of the model track, where a fixed harness does the work; the agent track's leader, Lyrie, sits at 99.2%. We covered the last big CyberGym result in August, when Fudan's Whitzard agent hit 91.2% on that harder track — Fudan's Whitzard agent ranks No. 2 on CyberGym — and the pattern holds across both: on this benchmark the differentiator keeps being training data and harness design, not parameter count.
AWS put CloudWatch Omni into general availability — an AI-first observability workspace that unifies agent, application and infrastructure telemetry in one application-centric view, with a free IDE extension that works without an AWS account. Omni was announced on September 22 and shipped GA the next day in three regions: US East (N. Virginia), US West (Oregon) and Europe (Ireland). It ingests OpenTelemetry, carries existing CloudWatch logs, metrics and traces forward unmodified, and bolts on an agent-specific layer — traces for prompts, model calls and tool invocations, an evaluation workflow, and natural-language investigation backed by AWS DevOps Agent. Frameworks in from day one include LangGraph, CrewAI, the OpenAI Agents SDK, the Vercel AI SDK and Strands, with third-party evaluators such as Braintrust, DeepEval and Ragas plugging in.
The pricing is where this gets interesting. Ingest runs $0.50 per GB for logs and $0.35 per GB for spans, dropping to $0.15 per GB over 30 TB; storage is $0.030 per GB-month on the default tier; agent evaluations bill at Bedrock AgentCore rates. That structure scales with exactly the telemetry agents generate in volume. AWS is betting that "why did the agent do that?" is a product category and not a feature — the analysts InfoWorld quoted are less sure, with HFS Research flagging vendor lock-in on the observability layer and Moor Insights warning that ingestion bills can climb faster than teams expect. The second caveat is the better one: observability that tells you what your agent did is not the same as one that stops it, and nothing in Omni is a runtime control.
What to watch: whether anyone re-runs XekRung independently, and whether Omni's ingestion math survives a fleet of agents that talk all day.
If a self-reported benchmark score is the only public measurement of a security model's capability, what should a security team treat as evidence? Tell us in the comments.
Sources: CyberGym leaderboard (UC Berkeley) · CyberGym leaderboard data · XeKRung Model on CyberGym (Alibaba Security) · CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale (arXiv) · Sina Finance on the CyberGym result · 快科技 on XekRung · Amazon CloudWatch Omni: AI-first observability for agents and applications (AWS What's New) · Introducing Amazon CloudWatch Omni (AWS News Blog) · AWS launches CloudWatch Omni to unify observability for AI agents and applications (InfoWorld) · Amazon CloudWatch Omni pricing