Apple's new Mac Studio and Mac mini are built for local AI

Share
Apple's new Mac Studio and Mac mini are built for local AI

No event, no keynote — just a surprise hardware drop with one very clear audience: people who run AI models on their desks.

Apple refreshed its entire desktop lineup around local AI inference today, debuting the M6 — its first 2nm chip — in a new Mac mini and the M5 Ultra, its first-ever quad-die design, in a new Mac Studio. The M6 brings a 12-core CPU and 12-core GPU with a Dual 16-core Neural Engine, and Apple claims the world's fastest single-threaded performance. The M5 Ultra goes further: it fuses two M5 Max dies with next-generation UltraFusion (over 4.4TB/s of inter-die bandwidth) into up to a 36-core CPU and 80-core GPU, with up to 512GB of unified memory moving at 1.2TB/s — 50 percent more bandwidth than M3 Ultra. That memory pool is the point: Apple says it can hold open-weight LLMs with hundreds of billions of parameters entirely on device. Pricing runs from $899 (Mac mini, M6) and $1,699 (M5 Pro) up to $2,499 (Mac Studio, M5 Max) and $5,499 (M5 Ultra); preorders opened today, shipping September 22.

Why this launch reads differently from a routine spec bump: ever since macOS enabled low-latency Thunderbolt 5 clustering for distributed MLX inference last December, developers have been daisy-chaining Mac minis and Studios to serve models far larger than any single consumer machine could hold — a budget-grade alternative to racks of specialized GPUs. This refresh, as Ars Technica observes, finally designs for that crowd rather than stumbling into it. Our take: with cloud token bills climbing and open models like Qwen and DeepSeek closing the capability gap, Apple is quietly claiming the "your desk is the datacenter" niche before anyone else thinks to contest it.


Researchers at Oasis Security disclosed a flaw in Nvidia's NemoClaw agent toolkit that let a single visit to a malicious website hijack the local model behind a developer's AI agent. NemoClaw launches its Ollama model server bound to every network interface with no authentication, so DNS rebinding let an attacker's page reach it through the victim's own browser — then rewrite the model's prompt template so a hidden instruction rode along beneath every message, invisible to guardrails and untouched by clearing the chat. Oasis reported it to Nvidia's security team before publishing, but the lesson outlives the patch: sandboxing the agent means little when the plumbing underneath is reachable from any tab.


Keenable, founded by former Yandex search chief Andrey Styskin, exited stealth with a $26 million seed round led by Accel to build a web search index designed for AI agents. The company claims more than 100 billion documents indexed, already used in production by several AI labs and inference providers — well timed, given Google and Microsoft have been retiring the search APIs much of the agent ecosystem leaned on. Styskin says the dream is becoming the next Google for AI agents, beating the incumbent on machine-driven queries where human-oriented result pages fall short.

What to watch: whether the top-end 512GB M5 Ultra configuration — which slips to late October — becomes the de facto entry point for running frontier-scale open models outside the cloud.

Would you run your next coding agent on a desk-sized Mac cluster instead of renting tokens? Tell us in the comments.

Read more

OpenAI busts influence ops that planted fake stories in real media

OpenAI busts influence ops that planted fake stories in real media

The day's AI news runs through one seam: the work is showing up in places nobody planned for — inside real newsrooms, across the whole night sky, and in the M&A column. OpenAI has banned two state-backed influence operations that used ChatGPT to plant fabricated stories inside legitimate news outlets — and rated the Russian one the most disruptive it has seen in two and a half years. In a report dated October 8, OpenAI detailed "Dark Clark," run from Russia across Latin America, which ran a th

Open Source Radar — October 9: plugins, sandboxes, tokens

Open Source Radar — October 9: plugins, sandboxes, tokens

Today's open-source signal is infrastructure rather than hype: Microsoft's code sandbox reaches 1.0, Anthropic's knowledge-worker plugins keep climbing, a beloved token counter flips its default, and LocalLLaMA squeezes a usable 2B model into about 700 MB. knowledge-work-plugins (Python, ~27,900 stars, Apache-2.0) — Anthropic's repository of role-shaped plugins for Claude Cowork is the top AI repository on today's daily trending page, and the stars keep coming: roughly 2,100 more than when we

Deep Dive — The four-token blind spot inside DeepSeek V4

Deep Dive — The four-token blind spot inside DeepSeek V4

ByteDance's Seed research team says it has found the cause of one of the stranger recurring complaints about DeepSeek's models: the same question, asked with nothing changed except a few junk characters bolted onto the front, can flip the model from right to wrong. Their paper, posted to arXiv on September 28, traces the wobble to a memory-saving trick used during long-context inference, and reports that DeepSeek-V4-Flash-Base's retrieval accuracy swings by as much as 40.2 percentage points depe

SoftBank seeks $100B from Gulf investors for an AI fund

SoftBank seeks $100B from Gulf investors for an AI fund

Three moves today point the same direction: the money, the politics, and the price of speed all got more expensive. SoftBank is reportedly seeking up to $100 billion from Gulf investors for a fund that would buy companies and run them with AI. The Financial Times reported the raise, citing people familiar with the matter, and says Masayoshi Son has held discussions in recent weeks with senior figures including in the United Arab Emirates; Reuters and Bloomberg both carried the report but neith