85% of Nscale's $103B in AI contracts sits with two customers

Share
85% of Nscale's $103B in AI contracts sits with two customers

Microsoft and Anthropic account for 85% of the $103.4 billion in AI compute contracts Nscale disclosed in its IPO filing — and only $2.6 billion of that was active as of late August. The London-based data center developer filed its S-1 on September 18 for a New York Stock Exchange listing under the ticker NSCL, targeting a raise of about $3 billion. Nscale reported $140.6 million in revenue for the first six months of 2026, up from $10.4 million a year earlier, against a $1.02 billion net loss.

The concentration is the story. One customer produced 52% of first-half revenue, and as of August 31 the company had 25,000 GPUs actually generating revenue against 461,000 active and contracted for future delivery. The filing also says Nscale has not secured binding financing for the build-out behind its largest contracts — including the roughly $44.6 billion, 460-megawatt Anthropic agreement at its Monarch campus in West Virginia, where the first two gigawatts of capacity are not expected to commission until the first half of 2028. Nvidia sits on all three sides: chip supplier, holder of more than 5% of the equity, and $1 billion of a $3.1 billion convertible note round signed three days before the filing.

Contracts are not cash. The whole neocloud category is valued on multi-year take-or-pay paper that only becomes revenue after billions in capital spending, and this filing shows exactly how far the paper runs ahead of the metal. We flagged the pre-IPO financing earlier this month — Nscale lines up $3.5B in pre-IPO cash with Nvidia at the table.


Multiverse Computing's answer to model compression: treat block removal as an Ising optimization problem. The company's write-up maps which of a model's transformer blocks to delete onto a spin-glass energy landscape, then searches it with classical and quantum-inspired solvers instead of greedy layer scoring. On Llama-3.3-70B-Instruct, removing 40 of 80 blocks with no retraining left it roughly 23 MMLU points ahead of the standard block-influence baseline; on Qwen3-14B at 12 of 40 blocks the lead was about 10 points. The method also transferred to NVIDIA's Nemotron-3-Nano-30B, a hybrid that interleaves Mamba2, attention, and mixture-of-experts layers.

The counterintuitive part is the bar they set. The solver never has to find the true ground state — the authors only need a handful of good low-energy configurations, which is why cheap solvers suffice and why the ranking is brute-forceable on a single GPU. That matters for anyone shipping 70B-class weights to constrained hardware: at deep compression the coupling between layers is the whole game, and greedy methods miss it.


Apple's M5 Ultra Mac Studio is being reviewed as a local-inference machine first and a workstation second. Reviewers with 256 GB configurations ran Qwen 3.8 Flash-Next at over 100 tokens per second at 4–5 bit quantization, with prefill up 150% against the M3 Ultra, and Tom's Hardware reports it outpacing both the DGX Spark and a Threadripper build. PCMag cites up to 4.3x faster local AI performance from the 36-core CPU / 80-core GPU part. The catch is the price, and every review says so.

The direction of travel is what's interesting. A single desktop now runs frontier-adjacent open weights at interactive speed, which is the hardware half of the argument we laid out in AI 101 — Local LLMs vs cloud APIs: what's the difference?. The gap that remains is memory: 156 GB of weights still does not fit in most machines.

What to watch: whether Nscale's Anthropic agreement converts from a signed contract into committed financing before the roadshow closes — and whether any other neocloud is forced to disclose the same active-versus-contracted gap.

Should regulators and investors treat a $103 billion contract backlog as demand or as risk? Tell us in the comments.

Sources: Bloomberg · Nscale S-1 (SEC) · Techmeme · Hugging Face · arXiv 2602.00161 · MacStories · Tom's Hardware · The Verge