Apple's new Mac Studio and Mac mini are built for local AI

Share
Apple's new Mac Studio and Mac mini are built for local AI

No event, no keynote — just a surprise hardware drop with one very clear audience: people who run AI models on their desks.

Apple refreshed its entire desktop lineup around local AI inference today, debuting the M6 — its first 2nm chip — in a new Mac mini and the M5 Ultra, its first-ever quad-die design, in a new Mac Studio. The M6 brings a 12-core CPU and 12-core GPU with a Dual 16-core Neural Engine, and Apple claims the world's fastest single-threaded performance. The M5 Ultra goes further: it fuses two M5 Max dies with next-generation UltraFusion (over 4.4TB/s of inter-die bandwidth) into up to a 36-core CPU and 80-core GPU, with up to 512GB of unified memory moving at 1.2TB/s — 50 percent more bandwidth than M3 Ultra. That memory pool is the point: Apple says it can hold open-weight LLMs with hundreds of billions of parameters entirely on device. Pricing runs from $899 (Mac mini, M6) and $1,699 (M5 Pro) up to $2,499 (Mac Studio, M5 Max) and $5,499 (M5 Ultra); preorders opened today, shipping September 22.

Why this launch reads differently from a routine spec bump: ever since macOS enabled low-latency Thunderbolt 5 clustering for distributed MLX inference last December, developers have been daisy-chaining Mac minis and Studios to serve models far larger than any single consumer machine could hold — a budget-grade alternative to racks of specialized GPUs. This refresh, as Ars Technica observes, finally designs for that crowd rather than stumbling into it. Our take: with cloud token bills climbing and open models like Qwen and DeepSeek closing the capability gap, Apple is quietly claiming the "your desk is the datacenter" niche before anyone else thinks to contest it.


Researchers at Oasis Security disclosed a flaw in Nvidia's NemoClaw agent toolkit that let a single visit to a malicious website hijack the local model behind a developer's AI agent. NemoClaw launches its Ollama model server bound to every network interface with no authentication, so DNS rebinding let an attacker's page reach it through the victim's own browser — then rewrite the model's prompt template so a hidden instruction rode along beneath every message, invisible to guardrails and untouched by clearing the chat. Oasis reported it to Nvidia's security team before publishing, but the lesson outlives the patch: sandboxing the agent means little when the plumbing underneath is reachable from any tab.


Keenable, founded by former Yandex search chief Andrey Styskin, exited stealth with a $26 million seed round led by Accel to build a web search index designed for AI agents. The company claims more than 100 billion documents indexed, already used in production by several AI labs and inference providers — well timed, given Google and Microsoft have been retiring the search APIs much of the agent ecosystem leaned on. Styskin says the dream is becoming the next Google for AI agents, beating the incumbent on machine-driven queries where human-oriented result pages fall short.

What to watch: whether the top-end 512GB M5 Ultra configuration — which slips to late October — becomes the de facto entry point for running frontier-scale open models outside the cloud.

Would you run your next coding agent on a desk-sized Mac cluster instead of renting tokens? Tell us in the comments.

Sources: Apple Newsroom · Ars Technica · The Verge · SiliconANGLE · Oasis Security · TechCrunch · Keenable

Read more