The Frontier
Deep Dive — Latent-space pretraining survived 5.7 trillion tokens
For as long as large language models have existed, one rule has been treated as structural rather than optional: predict the next token, and let structure emerge from the statistics. A technical report posted to arXiv on September 9 breaks that rule at a scale where earlier attempts fell apart