Anthropic's automated researcher beats humans at $4 an hour
Two research stories landed today on the same theme: machines doing work that used to be a human's job. In one corner, an Anthropic fellow shows that an LLM-driven research loop can already out-design its human handlers on alignment post-training. In the other, Meta is testing robots that pull the cables inside its data centers.
Anthropic's Automated Alignment Researcher (AAR) is cheaper, faster, and — in a six-hour window — better than the human researchers it replaced. A paper from Anthropic Fellow Chen Yueh-Han describes a closed-loop system that scans the literature, proposes a method, trains a model against it for half an hour, scores the result, and iterates. Across 10 categories of alignment failure, the system found fixes that improved each benchmark without degrading capabilities, the paper says. "The best AAR method beats what experienced humans propose, on average within six hours," the team writes. "Human guided research directions do not lead to stronger performance." The cost line is the eye-opener: roughly $4 per hour in API inference, against $150 per hour the lab pays its human researchers — a 37× gap. The paper is honest about the limits — the system only improves what the benchmarks measure, and the benchmarks still need humans to maintain them — but the trajectory is the point. Once automated research is on a cost curve, "we have time to do more research" stops being a budget line and starts being a throughput number.
Meta is testing ABB-built robots to swap server cables and reset machines in its data centers, Wired reports, citing unnamed sources familiar with the effort. The push is meant to keep labor costs down as Meta's AI buildout balloons: a faster way to physically maintain the racks an AI fleet depends on, and one that scales with the footprint rather than the headcount. Workers' concerns are the obvious subtext — Bluesky commentary on the story noted that the very "AI factory" jobs the buildout was supposed to create are exactly the ones now being automated away. Meta's footprint is large enough that even a small robotics rollout becomes a meaningful test of whether hyperscalers can keep growing without proportionally growing the people who keep the lights on.
What to watch: whether the AAR pattern holds outside alignment and into capabilities research — and whether Meta publishes numbers on what fraction of data-center tasks robots actually replace, instead of just assists.
Do you think labs should publish automated-research internals the way they publish model weights — open by default? Tell us in the comments.
Sources: TechCrunch · Anthropic research · Wired