Google's DiffusionGemma rethinks how LLMs write text — and how fast
A new open-weight model from Google trades the one-word-at-a-time habit that has defined language models for a decade, and a separate wave of harder-to-spot AI fakes is spreading on X. Plus, a Y Combinator startup thinks the key to cheap nuclear power for AI data centers is a part nobody talks about.
Google's DiffusionGemma generates whole blocks of text at once, hitting ~1,500 tokens per second on a single H100. Instead of predicting the next token in sequence, the model iteratively refines a 256-token block in parallel — sidestepping the decoding bottleneck that makes conventional autoregressive LLMs slow. The team built it by fine-tuning the MoE Gemma 4 model (3.8B active / 25.2B total parameters) with a two-stage pipeline that used less than 10% of the original model's training-token budget, and it still keeps thinking mode, multimodal input, and long context. It's "experimental," but it sets a new Pareto frontier for the speed-versus-capability trade-off, and it can still fall back to normal autoregressive generation with only minor loss — a hint that hybrid diffusion-AR models are coming.
"Subtlefakes" — slightly altered, nonconsensual AI images — are quietly taking over X. Reporting from 404 Media describes a new class of synthetic media: not the obvious deepfake nudes of past years, but real celebrity photos edited with AI to look more revealing, or entirely generated "post-workout" selfies, designed to read as plausible rather than pornographic. The target named in the piece, actor Xochitl Gomez, shared her own side-by-side showing a parking-lot photo edited to imply she was bending over. Because they skirt the nudity filters most generators enforce, they're trivial to make, and verified engagement-farming accounts earn ad-revenue-share money from the impressions. 404 Media found at least some watermarked with X's own Grok, and image-forensics expert Hany Farid warns the volume and sophistication of fakes is "nothing we've seen before."
Apollo Atomics says the secret to cheaper nuclear power is the steam generator, not the reactor. The YC-backed startup just closed a $26 million seed round (including $5M debt) to commercialize an MIT-born design that shrinks the steam generator — the bulky component that turns reactor heat into turbine-driving steam — from a hand-built, multi-story structure to something the size of a person. That lets Apollo build a reactor roughly 40x smaller than conventional designs, factory-assembled in under 24 months, which the CEO projects at about 3 cents per kilowatt-hour — cheap enough to undercut natural gas. A 40-kilowatt demonstration unit already runs at MIT; commercial deployment is aimed at 2028, with 10–300 MW variants planned.
What to watch: whether DiffusionGemma's speed holds up on real agentic workloads, and whether X's revenue model keeps rewarding the subtlefake farms.
Is "plausible but fake" more corrosive to trust than "obviously fake"? Tell us in the comments.
Sources: arXiv: DiffusionGemma Technical Report · Google DeepMind Blog · 404 Media — Subtlefakes · TechCrunch — Apollo Atomics