d-Matrix stacks DRAM on its chip and claims 100 TB/s

Share
d-Matrix stacks DRAM on its chip and claims 100 TB/s

Hot Chips 2026 wrapped with a memory bet, and the funding calendar is unusually busy for a Monday.

d-Matrix says stacking memory directly on the compute die beats HBM by 20x on bandwidth density. The startup's Raptor accelerator puts a TSMC N4 logic die face-to-face on top of custom 3D DRAM and claims roughly 100 TB/s per card at 32 GB — against a practical ceiling of about 20 TB/s for today's HBM4 packages, which d-Matrix attributes to limited package "beachfront" and raw physics: driving 100 TB/s through HBM at 2.4 pJ/bit burns about 1.92 kW before any fabric traffic. The pitch numbers: 32.6 GB/s per square millimetre against roughly 1.5 for HBM parts, 13.5x better on power per GB/s, and a claimed 1,000 tokens per second per user while serving a 3-trillion-parameter-class model at 1M context — a 72-card scale-up that fits one rack can host a frontier model such as Kimi K3. That is the whole argument in one line: decode, not training, dominates wall-clock inference time, and decode is memory-bound, so the vendor that fixes the pipe wins the bill. We explained why HBM became the choke point of the AI buildout in AI 101 — What is HBM (high-bandwidth memory)?.


Buildots raised $130 million led by O.G. Venture Partners to sell construction schedules to the AI data-center boom. The Tel Aviv company films job sites with helmet-mounted cameras, turns the footage into a live digital representation of the project, and diffs it against the schedule and the 3D model to classify completed work and forecast delays. Total capital now sits at $297 million, up from a $45 million Series D roughly 16 months ago, with customers growing from about 50 construction companies to more than 100 large firms; Calcalist reports the round values Buildots near $1 billion. The demand story is the interesting part — the company says AI data-center construction is pulling its product, because a slipped substation or shell hands a delay straight to a cluster of GPUs that is already paid for. Eight years of real site footage is the moat it claims, versus general-purpose models trained on scraped internet data.


Apple's new Siri is built so its brain can be swapped for Claude or GPT, and reverse-engineered code shows how deep the hook goes. Sleuth "pdfu" found private frameworks in iOS 27 and macOS Golden Gate with two distinct paths: an extension protocol where "Ask Claude" routes a natural-language request to the third-party model while Siri keeps executing system actions, and a far more consequential inference provider inside "Model Manager Services" that can replace Apple's own server-side Siri model entirely — the outside model receives Apple's Siri planner prompt and tool definitions, requests system actions, gets personal data back, and answers through Siri's voice and interface. Apple has not opened the model-delegation entitlement to third parties and it isn't front-facing; today the "Ask…" path is limited to the ChatGPT extension in the macOS 27 release candidate. The Digital Markets Act, which the European Commission has said applies to Siri, is the obvious reason any of this plumbing exists in the first place.

If Apple is quietly engineered for a swappable assistant, should users pick their model per request — or is that exactly the support nightmare Apple refuses to own? Tell us in the comments.

Sources: ServeTheHome · Tom's Hardware · SiliconANGLE · Buildots · Calcalist · MacRumors · AppleInsider