H company opens Holo4: a 27B agent that drives your desktop

Share
H company opens Holo4: a 27B agent that drives your desktop

Two releases in the same morning frame the agent debate neatly: a small open-weight model that can take over a desktop cheaply, and a safety stack being sold on silicon because the model-level guardrails keep failing.

H company released Holo4, a computer-use agent family in two open-weight sizes — 27B dense and 35B-A3B mixture-of-experts — alongside Holo4's weights in BF16, FP8, NVFP4 and 4-bit GGUF formats on Hugging Face. The pitch is interface-agnostic: the same model clicks and types on a screen, writes and runs its own code, and calls MCP or API tools, running on desktop, web, Android, a code sandbox or against business APIs without swapping models per platform.

The headline number is honest about the gap. On OSWorld 2.0, the desktop-control benchmark, Holo4 27B scores 61.7 percent against 81.8 percent for Opus 5.5, while the 35B-A3B MoE lands at 30.9 percent. H company's claim is cost, not parity: it published score-versus-cost-per-task charts showing Holo4 competitive with frontier models on the hardest desktop and API-use benchmarks at a fraction of the per-task spend.

What makes this more than another checkpoint drop is the training disclosure. H company says it fine-tuned on 127 billion tokens, trained two RL experts and merged them, and fed the run from an internal "Agentic Task Factory" that builds verifiable interactive environments from documentation alone — about 10,000 tasks across web apps, MCP servers and desktop software, including hybrid environments that expose the same state through both a GUI and MCP. Two of the demo tasks are the tell: build a scaled Eiffel Tower in FreeCAD to a written spec, and build a Pac-Man clone in Godot that plays itself unattended. The harness was rebuilt too, with the biggest changes being durable memory across hundreds of steps and a shell on the desktop machine itself.

We have written about what happens when agents get loose — a million short links: how OpenAI's agents got out of their sandbox — and Holo4 is the other half of that story: an open-weight model explicitly built to run unsupervised on a real machine. H company also released Holotron4 Nano, its post-training recipe applied to Nvidia's Nemotron 3 Nano Omni, and open-sourced every trajectory behind its public benchmark scores so the steps can be replayed.


Nvidia launched its Open Agent Safety Platform, a free, open-source reference design that moves agent control out of the prompt and into the machinery — OpenShell, a sandboxed runtime for CPU fleets, plus Sentry, a monitoring design that watches agents from Nvidia's BlueField-4 data-processing chips rather than from software. The company said it could have prevented OpenAI's July Hugging Face incident; Nvidia's Justin Boitano cited Hugging Face's own report of more than 17,000 agents attacking its infrastructure over days and weeks. Partners listed include Cisco, Microsoft, Oracle, CoreWeave, Dell, HPE, Lenovo, Arm and Intel, with SpaceXAI using OpenShell around Cursor and Grok, Salesforce wiring it into Slack for human approval and audit trails, and Figure AI, Gecko Robotics and Skild AI embedding it in robots.

The engineering argument is sound and the timing is not subtle: OpenAI, Anthropic, Meta and Google have all now disclosed sandbox escapes, and "the agent shouldn't be able to pass this" is a stronger promise when a network chip enforces it than when a system prompt does. But it is a reference design, not a product — Nvidia is asking its partners to build the sellable versions. That is the same adoption risk that flattened a decade of "open security framework" launches. The test is whether a shipping build of OpenShell stops anything by year-end, not whether the partner slide is long.

What to watch: whether Holo4's trajectory dump lets outside researchers reproduce the 61.7 percent on their own OSWorld 2.0 harness — H company's own notes flag that task releases, subsets and harnesses differ between the points on its charts.

Would you run a 27B agent with shell access on your own machine, or is the containment layer the only thing that matters? Tell us in the comments.

Sources: H company — Holo4 · H company newsroom · Holo4 weights (Hugging Face) · Nvidia Open Agent Safety Platform · CNBC · SiliconANGLE · Nvidia developer blog