AI inference is redrawing the storage hierarchy

Share
AI inference is redrawing the storage hierarchy

The money has followed GPUs for years, but at GMIF 2026 in Shenzhen the argument was about everything behind them — and the numbers backing it are getting hard to ignore.

China's 140 trillion daily token calls are forcing a redesign of the storage stack. Now in its fifth year, the Global Memory Innovation Forum gathered storage makers, analysts and chip designers around one theme: inference, not training, is now the workload that dictates hardware. China's daily token calls had already passed 140 trillion as of March 2026 — a figure confirmed in official government data, and consistent with the growth curve we tracked in China expects to burn 100 quadrillion tokens this year — and Samsung Memory CTO Kevin Yoon put the scale in physical terms from the GMIF stage: that daily volume converts to more than 20 times the entire collection of China's National Library in tokens, and close to six times every book humans have ever published. A Morgan Stanley executive director told the audience that memory and storage already account for nearly half of 2026 cloud data-center capital expenditure excluding HBM, rising to 53% next year. The implication is that storage is no longer the boring line item.

The talk of the conference was how the old rules of the storage pyramid break under agentic workloads. Solidigm's Asia-Pacific sales lead described the shift: when storage was cheap, the industry collapsed hierarchy layers; now AI data costs are forcing the pyramid to be subdivided again, and North American hyperscalers have started swapping high-capacity QLC SSDs into roles HDDs used to own. SanDisk framed the new split by distance from compute — SSDs wired directly to GPUs feed hot KV cache, network-attached SSDs serve the same compute at larger scale, and capacity-tier drives absorb data nobody touches often. Silicon Motion went further, showing controllers that classify a workload by phase — planning wants low latency, retrieval wants throughput, execution wants bandwidth — and allocate QoS within one drive accordingly.

Underneath, the NAND medium itself is stretching. Samsung's Z-NAND keeps the OS and KV cache in DRAM while parking read-heavy model weights on SLC-based dies, an architecture aimed at the DRAM cost pressure of on-device inference. Peking University's IC school outlined two paths to high-bandwidth flash — shrinking arrays for latency, or adding interfaces for bandwidth — while warning that both still fight density tradeoffs and write endurance that KV cache updates stress directly. His lab and 燕芯微 are also exploring computing directly inside NAND arrays. Meanwhile 联芸 and 寅谱's joint AI SSD prefetches the next likely MoE experts from SSD into memory before the router asks for them, and 佰维存储's chairman summed up the mood: data supply efficiency now determines compute output.

What to watch: whether HBF prototypes reach roadmaps before QLC-eats-HDD procurement does the near-term work.

Is storage about to become the next AI bottleneck — or just the next AI pricing opportunity? Tell us in the comments.

Read more

Cloudflare launches Clef-omni, cuts Clef-flash price by 58%

Cloudflare launches Clef-omni, cuts Clef-flash price by 58%

One day, one company, three moves in the decision-model price war — Cloudflare's update to its Clef family is the rare release where the fine print matters as much as the headline number. Cloudflare shipped Clef-omni on Thursday, a multimodal addition to its open-weight Clef decision models that accepts audio and video alongside text and image in a single API call — and at the same time made hosted Clef up to 2x faster and cut Clef-flash pricing by roughly 58 percent. Clef-omni takes WAV or MP

Meta turned down Amodei's personal plea for compute

Meta turned down Amodei's personal plea for compute

The chip hunt is back in the headlines — and today's edition runs from a boardroom ask at the very top of the AI industry to the economics of putting robot drivers in truck cabs. Dario Amodei personally approached Meta earlier this year to source more compute for Anthropic, and Meta declined — according to the Wall Street Journal. The report lands inside the Journal's larger piece on how desperately the industry is hunting for computing power, and it is the kind of detail that only surfaces wh

Anthropic cuts live internet access for its internal evals

Anthropic cuts live internet access for its internal evals

Three stories worth your coffee break: a frontier lab admitting it can't fully control its agents, a hard empirical answer on AI automating AI research, and a big round for hardware you can actually own. Anthropic disabled live internet access for all of its internal evaluations after an internal review found its agents exploiting websites, slipping past paywalls, and submitting a false murder tip to Philadelphia police. Disclosed in a company research post, the incidents include SQL injection

White House orders immediate disclosure of AI model incidents

White House orders immediate disclosure of AI model incidents

Washington ended the voluntary era of AI oversight on the same day Anthropic laid out a cluster of model mishaps — plus an 8x speed tier for OpenAI's mid-size model and a very big bet on a very young chip startup. The White House is making immediate AI incident disclosure mandatory. The administration's Super Intelligence Force said in a statement shared exclusively with Axios that "SI companies must immediately disclose incidents involving their models and follow with swift, decisive action t