Two fixed opening tokens push a base model past its RL version

Share
Two fixed opening tokens push a base model past its RL version

A new paper from MIT, UC Berkeley, Washington University and the Allen Institute for AI says much of what reinforcement learning teaches a model may come down to how it starts an answer. One good prefix, fixed in advance, is enough to close the gap.

Fix the first two tokens and Olmo-3-7B beats its own RL-trained twin. The paper — "Base Models Can Reason By Taking a Cue From Training Data" — tested what happens when researchers pre-fill a model's response opening instead of letting it choose. On MATH-500, the base Olmo-3-7B scores about 42% on its own; force the opening to . followed by a blank line and "Okay," and it jumps to roughly 78%, ahead of the reinforcement-learning version's 75%. Qwen3-14B shows the same pattern with " Alright," as the cue, climbing from 72% to 87% and matching its RL counterpart. No parameters were touched and no worked examples were added — only the starting words changed. The researchers picked the cues by scanning openings the model already produces and measuring which one yields the steadiest answers; on correct responses, about half begin with a period and two newlines, versus 14% of wrong ones.

It reframes what RL is actually buying you. Under an RL-Zero setup, the biggest distribution shift between the base and RL-trained checkpoints sits in the first two tokens of the response: RL raises the odds of the winning cue from 0.14 to 0.65 for Olmo and from 0.04 to 0.58 for Qwen. A model given the RL version's opening, then left to continue on its own, matches the RL score — and a fixed cue is worth roughly 100 steps of RL training for Olmo. The authors' data-rewiring experiment makes the mechanism concrete: swap "Okay" for "Chicken" in the mid-training corpus, retrain, and .\n\nChicken becomes an effective reasoning cue too, lifting MATH-500 from 18% to 77%. "Think step by step" works the same way after "duck duck goose" is substituted. The cue is not magic wording — it is a pointer into associations the training data built.

The safety finding is the part worth watching. Openings also steer refusal behavior: an "I'm sorry" start makes Olmo refuse more, including harmless requests, while "Okay," loosens compliance on dangerous ones — and merely changing punctuation between Okay, and .\n\nOkay flips the tendency. That means alignment behavior can hinge on a prefix a user may be able to influence, not just on learned weights. The caveat is scope: the results cover two open models on math and code benchmarks, so generalization to frontier models is unproven.

What to watch: whether labs test prefix sensitivity as an eval, and whether the cue trick survives at frontier scale.

Does RL mostly teach models how to start talking? Tell us in the comments.

Read more

Senate report: hyperscalers misled the public on data center costs

Senate report: hyperscalers misled the public on data center costs

A yearlong Senate investigation just put Congress's own numbers under the claims AI data centers sell to towns and ratepayers. Also: the Times finally counts what Anthropic's agents submitted to the State Department, and StepFun sets an open-weights date for its top-ranked flagship. A yearlong Senate investigation led by Senators Warren, Van Hollen, and Blumenthal concludes that some of the biggest hyperscalers misled the public about what their AI data centers cost everyone else. Staff querie

Nvidia-backed Firmus pulls its $30 billion IPO after demand collapses

Nvidia-backed Firmus pulls its $30 billion IPO after demand collapses

The biggest AI-infrastructure listing of the year didn't get delayed — it got pulled entirely, after two price cuts in two days found no takers. Nvidia-backed data-center operator Firmus withdrew its Australian IPO on Friday, abandoning what would have been the country's second-largest listing ever after slashing the offer price twice in 48 hours. The company had been marketing shares at A$11 each to raise about $5 billion at a valuation near $30 billion — roughly three times the $10.5 billion

AI inference is redrawing the storage hierarchy

AI inference is redrawing the storage hierarchy

The money has followed GPUs for years, but at GMIF 2026 in Shenzhen the argument was about everything behind them — and the numbers backing it are getting hard to ignore. China's 140 trillion daily token calls are forcing a redesign of the storage stack. Now in its fifth year, the Global Memory Innovation Forum gathered storage makers, analysts and chip designers around one theme: inference, not training, is now the workload that dictates hardware. China's daily token calls had already passed

Cloudflare launches Clef-omni, cuts Clef-flash price by 58%

Cloudflare launches Clef-omni, cuts Clef-flash price by 58%

One day, one company, three moves in the decision-model price war — Cloudflare's update to its Clef family is the rare release where the fine print matters as much as the headline number. Cloudflare shipped Clef-omni on Thursday, a multimodal addition to its open-weight Clef decision models that accepts audio and video alongside text and image in a single API call — and at the same time made hosted Clef up to 2x faster and cut Clef-flash pricing by roughly 58 percent. Clef-omni takes WAV or MP