JD Logistics' Super Brain 3.0 now runs its warehouses end to end

Share
JD Logistics' Super Brain 3.0 now runs its warehouses end to end

China's biggest e-commerce logistics operation just handed more of itself to AI. Also today: MIT Technology Review published a hands-on set of puzzles that frontier models still flunk — a useful gut-check on how brittle those benchmark scores really are.


JD.com's logistics arm launched "Super Brain" 3.0, an industrial-grade large-model system that now coordinates everything from warehouse shelving to last-mile delivery. Announced Tuesday at China's international transport technology and equipment exhibition, the upgrade claims higher throughput, stronger fault tolerance, and lower operating costs than version 2.0 — with the headline change in route planning: solving optimal paths across hundreds of millions of parcels dropped from minutes to seconds, according to Yicai, citing JD's announcement. The company bills it as the first industrial-scale AI system in logistics to direct the full journey of a parcel — storage, sorting, line-haul, and doorstep — rather than assisting discrete steps, and says it underpins JD's push toward what it calls the world's largest physical-world operations center. The take: this is where agentic AI actually earns money today — not in demos, but in seconds shaved off a routing solver at a scale of hundreds of millions of daily parcels. If the efficiency numbers hold up outside JD's own network, expect every major logistics operator to follow.


Frontier models still fail puzzles that most humans find easy — and MIT Technology Review wants you to try them yourself. Reporter Grace Huckins assembled an interactive gauntlet of seven test types where today's models reliably stumble: mental rotation of 3D objects, trick variants of classic Knights and Knaves riddles, the adversarial SimpleBench questions, ARC-AGI grid puzzles, intuition traps built to bait humans, and river-crossing or logic-grid problems that collapse once complexity rises. The pattern behind the failures is more interesting than any single flub: models ace near-memorized riddles while missing the twist, do far better on visual grids when the image arrives as numbers instead of pixels, and handle planning tasks until roughly six moving pieces — then fall apart. Progress is real, though: models went from solving 18 percent of the New York Times' Connections puzzles in late 2024 to near-perfect accuracy by early 2025. The take: benchmarks keep climbing, but the failure modes are specific and stubborn — spatial reasoning, genuine rule induction, and long planning chains — and a puzzle page you can attempt in ten minutes makes the gap tangible in a way no leaderboard does.

What to watch: whether JD opens Super Brain 3.0 to third-party merchants and overseas warehouses — that would turn an internal cost tool into a product line.

Would you trust a model's benchmark score after failing one of these puzzles yourself — or does your own trial matter more? Tell us in the comments.

Sources: Yicai · Nanfang News · MIT Technology Review · Scientific American