---
title: "Jefferies tests 8 AI office agents — Qwen Office takes the crown" format: brief tags: [daily-brief, arena] section: arena meta_description: "Wall Street analysts hand-tested eight AI agents on real office tasks. Alibaba's Qwen Office scored highest, beating Claude Cowork and OpenAI's Codex." ---
Wall Street is done watching benchmark leaderboards — now it's running its own trials. Jefferies analysts put eight mainstream AI agents through real-world office tasks, and Alibaba's Qwen Office came out on top, ahead of both Claude Cowork and OpenAI's Codex.
Jefferies analysts tested eight global AI agents on five real office tasks and Qwen Office ranked first — the only agent to score above 90 in every category. The test battery was designed to mimic actual desk work: summarizing a company annual report from multiple source documents, searching the web to compare business metrics, controlling a live browser to retrieve information and generate a document, building an English-language presentation from data, and creating a marketing poster from a reference image. Qwen Office was the only product to clear 90 points across all five dimensions, with particular strength in complex multi-step tasks, browser automation, and multimodal content generation. The result matters because it's one of the first times a major Wall Street firm has published a hands-on, task-level comparison of agentic products rather than relying on model-level benchmark scores. For enterprises evaluating which agent to deploy, the implication is that the harness — the engineering scaffolding around the model, including prompt design, tool orchestration, context management, and error correction — can matter as much as the base model itself. Jefferies' analysis found that Qwen Office's "implied Harness score" led all eight products, suggesting Alibaba has invested heavily in the production engineering that turns raw model capability into reliable task completion. The cost angle also favors Qwen: Qwen 3.8 Max's API pricing is significantly lower than several frontier competitors, and because agents require multi-step reasoning and repeated tool calls, the per-task cost advantage compounds over real workflows. That said, the test was narrow — five tasks, eight agents, one analyst team — and real enterprise deployments involve far more complex, domain-specific workflows. But the direction is clear: the agent competition is moving beyond raw model capability toward the full stack of engineering, pricing, and integration that determines whether an agent actually works in production.
MIT Technology Review argues child-monitoring AI apps need a fundamental rethink. In a feature published today, Kelly Clancy examines how AI-powered content-scanning tools like Bark — which scanned eleven billion messages to or from 7.5 million children in the US in 2025 — use machine learning classifiers to flag potentially harmful content including self-harm, bullying, and predatory behavior. The findings are sobering: false positives are rampant, with one parent calling ninety-nine percent of alerts "garbage," and the surveillance dynamic erodes the parent-child trust it's meant to protect. Nearly one in five monitored children reported feeling watched or stripped of privacy, and one in ten said it broke their trust in parents. Researchers led by Pam Wisniewski at UC Berkeley are pushing a resilience-based alternative — training kids to recognize and cope with risk rather than monitoring every message — and early trials show promise, with trained students engaging in far fewer risky online behaviors. The piece also touches on the broader AI safety landscape, noting pending lawsuits against OpenAI and Character.AI over chatbot-related harm to minors, and the Molly Rose Foundation's analysis finding that most major social platforms barely moderate suicide and self-harm content. For the AI industry, the core tension is stark: the same classifier technology that can catch a genuine predator also flags two kids calling a third one "annoying" as bullying.
What to watch: Whether other Wall Street firms publish their own agent benchmarks, and whether the resilience-based approach gains traction as an alternative to surveillance-heavy child safety tools.
Do you think the harness matters more than the model when it comes to real-world agent performance? Tell us in the comments.
Sources: Leiphone (雷峰网) · MIT Technology Review · Bark Annual Report 2025