Anthropic unveils Model Hardware Standard for physical-world AI agents

Share
Anthropic unveils Model Hardware Standard for physical-world AI agents

Agents have lived on screens since the moment they were born. Today that changed on two fronts: Anthropic put its models in charge of real lab machinery, and a new study showed we're a long way from trusting them with something as simple as a shopping cart. Both stories point the same direction — the agent is leaving the chat box.

Anthropic launched the Model Hardware Standard, an interface that lets AI agents directly operate physical machinery — from lab instruments and microscopes to quantum computing hardware and robot arms. The company announced the standard in a research preview on Thursday, describing it as the hardware equivalent of a USB-C cable: one common interface for transmitting information between an agent and any device with a programmable connection, instead of bespoke integrations for every instrument. That's a meaningful bet. Anthropic is framing MHS as the bridge that finally lets AI agents get their hands dirty, and it's launching with partners already using it — Genentech automating a lab assay, Carnegie Mellon mapping dose-response curves, HHMI Janelia coordinating microscopy, QuEra Computing stabilizing quantum lasers, and a Tetsuwan Scientific rig running pollution assays.

The details matter as much as the ambition. MHS is model-agnostic, so it isn't locked to Claude, and it's arriving as a research preview for a select set of science, robotics, and manufacturing organizations with plans to open-source it down the road — the same strategy Anthropic used when it open-sourced the Model Context Protocol in 2024 and watched it become an ecosystem standard. That pattern is the real story here: Anthropic is trying to make its software the connective tissue of an entire hardware ecosystem, positioning itself for the physical-world AI race where rivals including OpenAI and Amazon are already pouring billions into AI-native devices. If MHS catches on the way MCP did, Anthropic doesn't need to build the robots — it just needs every one of them to speak its language.


A new study found that AI shopping agents are nowhere near reliable enough to buy on your behalf — and the failures are weirder than simple mistakes. Researchers tested how frontier models like Claude Opus 4.8 and Gemini 3.5 pick products when given sources like a Reddit thread, a Wirecutter review, and a Strategist piece. A single source could swing a recommendation by up to 99 percentage points; the order in which the same sources were presented changed outcomes for several models; and short memory statements like "I love hiking" pushed some models toward pricier, objectively worse products even when a $29.99 five-star option was sitting in the grid.

The takeaway cuts both ways. For shoppers, letting an agent spend your money is a gamble — the same query can yield different picks on different days for no visible reason. For sellers, it's worse: optimizing for AI shopping will be harder than SEO ever was, because merchants won't know which model is shopping, what it read first, or how it processed the inputs. No consistency, no control — which is exactly why the companies building the infrastructure (including Google's agent payments protocol) should be watching this research closely.

What to watch: whether Anthropic ships MHS as open source by year-end, and whether the lab partners' reports of faster, more autonomous experiments hold up at scale — that's the difference between a viral spec and a working standard.

Should shopping agents be held to a reliability bar before they're allowed to spend your money? Tell us in the comments.

Sources: CNBC · Reuters · Anthropic — Model Hardware Standard research preview · The Decoder · ACES simulator (arXiv)