AMD buys Taalas to hardwire AI models into silicon
Inference is where the AI hardware war is being fought, and AMD just made its biggest move yet: a definitive deal to acquire Taalas, the Toronto startup that etches model weights directly into transistors. Also in the mix: a 2-bit Qwen3.6 that fits on a 16 GB GPU, and Anthropic's quiet worry that its money is outrunning its mission.
AMD has agreed to acquire Taalas, the AI-inference startup that bakes a model's weights and dataflow directly into silicon. The deal, announced at market close on Thursday, brings ex-Tenstorrent CEO Ljubisa Bajic's team under AMD's AI group and will integrate Taalas' model-specific integrated circuits (MSICs) into system-level solutions alongside Instinct GPUs, EPYC CPUs, Helios racks and ROCm — terms undisclosed, close expected in Q4 subject to regulatory approval. Taalas' HC1 test chip served Llama 3.1 8B at roughly 17,000 tokens/sec — a claimed 73x an H200 at a tenth of the power — but each chip runs only the model it was built for.
This is AMD's third AI acquisition in nine months, after MK1 and Mext, and its clearest answer yet to Nvidia's reported $20 billion Groq licensing deal: specialized, single-model silicon for the agent-inference era. The tradeoff — a chip that goes obsolete when the model updates — only works if frontier models stop churning, which makes this as much a bet on model stability as on silicon.
EschaLabs released a 2-bit quantized build of Qwen3.6-35B-A3B that squeezes a 256-expert MoE into 12.3 GB and runs on a single consumer GPU. The Escha-W2 checkpoint claims near-lossless quality against its FP8 baseline — 100.2% mean retention across six axes, with MMLU-Pro at 80.9 and LiveCodeBench as the one real gap — at 5.6x compression, serving 225 tok/s single-stream on an RTX 4090 and up to ~2,670 tok/s batched on a 5090 via its SGLang-based runtime. It's Apache-2.0, ships an OpenAI-compatible API, and runs on 16 GB cards like the RTX 5060 Ti — trading context or concurrency, not both.
Sub-2-bit MoE quantization that survives GPQA and MATH-500 is a strong signal that the frontier of "frontier-class model on a mid-range card" keeps moving; the honest caveat is that these are vendor-measured numbers.
Anthropic CEO Dario Amodei has reportedly told associates he's worried new hires are joining for the money rather than the mission. Axios reported the concern, citing a source familiar, amid a broader AI talent war where salaries and pre-IPO equity packages have gone vertical. The optics are awkward: the company pays top-of-market — one events-lead posting lists $320K–$400K, roughly six times the industry average — even as Amodei frets about mercenary motivation. It's the unavoidable tension of a frontier lab valued near $1T: when the mission pays like a lottery ticket, you can't tell the believers from the traders — and as Gergely Orosz put it, the market is already pricing that in.
What to watch: whether AMD ships Taalas silicon as a Llama/Claude-optimized "premium inference" SKU — and whether Anthropic's next round of hiring filters for anything besides comp.
Does hardwiring AI models into silicon make sense — or does it lock in architectures that go stale in a year? Tell us in the comments.
Sources: AMD press release · model card · Axios · The Register · Yahoo Finance