Marin opens a 535B-model training run in live public view

Share
Marin opens a 535B-model training run in live public view

A Chinese lab just started what it bills as the largest fully-open large-model training run ever — and it wants the world to watch it happen in real time, not read about the results after the fact.

Marin has begun training a 535-billion-parameter MoE model with its code, data mixture, and live loss curves published as it goes. Called Marin 535B-A23B, the model activates about 23 billion of its total parameters per token and is being trained on 18.75 trillion tokens across 11 NVIDIA GB200 NVL72 racks — roughly 792 GPUs — over about three months. The plan splits the token budget 80/20 between pretraining and midtraining, and the full effort is pegged at around 2.7×10²⁴ FLOPs. The run is live as of this week: there are no final weights yet, which is exactly the point. Instead of releasing a clean technical report only after a successful run, Marin published its operating plan, engineering risks, scaling methodology, and contingency procedures in an open GitHub issue, with a Weights & Biases board tracking training loss and per-domain data composition along the way.

The openness is the story. Marin describes the effort as an experiment in doing frontier-scale training in public — and as a transparency argument for open-weights AI. Rather than have outsiders take the team's word that the model was trained cleanly on the announced mixture, the project exposes the evidence during the run: sampled documents, the exact pretraining mixture per domain, and forecasts that let anyone check whether the model is tracking its expected loss trajectory in real time. The venture has drawn prominent backing from the research community, including Andrew Ng, who publicly endorsed the approach, alongside senior figures in open-model development.

What makes this more than a publicity stunt is the engineering groundwork. Before launching the full run, Marin trained a four-rung scaling ladder from 1.6 billion to 27.7 billion parameters, costing about 1% of the main run's compute. That ladder doubles as both a loss forecast and an early-warning system — a material deviation in the full run can trigger an investigation before months of compute are wasted. It already paid off once: the ladder surfaced gradient-norm growth above four on longer horizons, which led the team to adopt logit z-loss regularization to keep high-batch configurations from diverging. Marin also documented a custom all-to-all implementation for the expert-parallel transport, and honestly reports the failed designs along the way, not just the wins.

There's a legitimate question of how much can be inferred from a partial record — extrapolating from small runs to a 535B model is one of the experiment's central uncertainties, and a short one-rack gate test doesn't establish the final token-drop rate. But the framing matters. Marin is betting that radical transparency becomes a credibility weapon in the open-weights race, letting the community audit a frontier-scale training run while it's happening instead of trusting a retrospective paper.

What to watch: whether Marin hits its reported ~250,000-tokens-per-second design throughput across all 11 racks over the coming weeks — the first real test of whether the roadmap survives at full scale.

Open training at this scale is a meaningful bet on radical transparency. If it works, does it set a new bar for how every major open model should be built — or will short-term loss wobbles give skeptics ammunition? Tell us in the comments.

Read more

AI 101 — What is a jailbreak?

AI 101 — What is a jailbreak?

A jailbreak is a prompt — or a carefully arranged stack of inputs — engineered to talk an AI system past its own safety rules, so that it produces content or takes actions it would normally refuse. Nothing is broken in the technical sense: the model, the servers, and the locks all keep working. What gets broken is the instruction to say no. Why it matters right now The word "jailbreak" shows up constantly in AI coverage — in stories about chatbots misbehaving, about guardrails, about agents

Samsung projects a 100 trillion won quarter on AI memory demand

Samsung projects a 100 trillion won quarter on AI memory demand

The AI buildout's money keeps landing in the same place — memory — and Samsung just put the biggest number yet on it. Meanwhile, China's leading open-model lab walked through what its next models still can't do. Samsung projected third-quarter operating profit of 107.4 trillion won — roughly $80 billion — which would be the first time any company has cleared 100 trillion won in a single quarter. The preliminary guidance, released Thursday in Seoul, compares with 12.17 trillion won a year ago,

Lei Jun's fund and Huawei's Hubble back DiffuSpace's record round

Lei Jun's fund and Huawei's Hubble back DiffuSpace's record round

Chinese money went after a non-autoregressive architecture this morning, while one of the world's most-watched investors delivered his sharpest warning yet that the AI trade is closer to the exit than the entrance. Shenzhen's DiffuSpace has closed two back-to-back rounds totaling several hundred million yuan — according to reports, the largest funding ever raised by a diffusion language model startup, and the company's first disclosure since it was founded this May. Matrix Partners China, Shun

Kling AI picks banks for a $1B Hong Kong listing

Kling AI picks banks for a $1B Hong Kong listing

The AI money story keeps circling back to Hong Kong — and today it's the video generation business making its move, while defense-tech AI quietly keeps raising. Kling AI has hired CICC, Goldman Sachs and UBS to guide a Hong Kong IPO that could raise at least $1 billion, according to Bloomberg's October 6 report, with a listing targeted as soon as 2027. The Kuaishou-owned video model unit is staging the offering off real revenue: its first half reportedly brought in more than 1.5 billion yuan (