UkisAI's Swift trims Qwen 3.8's thinking by 63% — and a rival's eval agrees

Share
UkisAI's Swift trims Qwen 3.8's thinking by 63% — and a rival's eval agrees

UkisAI released a family of fine-tunes that cuts the tokens Qwen 3.8 burns on second-guessing itself — 63.4 percent fewer thinking tokens on Swift Flash Next and 58.5 percent on Swift 1.5 27B, with the vendor reporting accuracy within a fraction of a point of the base model.

That is the flattering half. Almost all of these numbers are self-reported, and local-model claims usually die on contact with someone else's harness. This one has a first check: an independent Aider eval suite run at Q8_0 and posted to r/LocalLLaMA measured the earlier Swift 27B at 7,301 median completion tokens against 12,547 for stock Qwen 3.8-27B — same 27.1 percent first-try pass rate, same retry pass rate, roughly 40 percent fewer tokens and half the wall-clock per case. The remaining caution is real, though. The headline 63.4 percent is a median on one model at the xhigh setting; across the 27B's benchmark suite the mean reduction is about 19 percent, and one penalized token hurt math, with AIME 2026 down 4.67 points and a fix promised in the next release.

The method is the interesting part, and it is the same idea we keep seeing in reasoning-token research: UkisAI generated thousands of traces across coding, language, vision and agentic work, found the tokens that recur specifically inside overthinking loops, penalized those during fine-tuning, then restored the accuracy drift with RL and on-policy distillation. It never capped output length. For anyone whose agent bills are dominated by verification loops rather than hard reasoning, that distinction is the product — and the local crowd is already running Flash Next at IQ4_XS on a single 4090 for about 40 tokens per second of decode.


Reuters-confirmed numbers surfaced this week on both sides of the US AI oversight debate: the NSA told lawmakers it is spending billions of dollars this year to evaluate and test advanced AI models, while the CBO scored the leading proposal for a new AI safety center at $20 million a year. The NSA figure comes from two sources described as familiar with classified intelligence estimates, and the Pentagon declined to comment; the exact amount is not public.

Compare the scales. Testing frontier models at defense scale already costs billions, and the body that would give Congress a standing technical assessment of AI risk is priced at $20 million a year, with a separate House AI tracking bill at $36 million over five years. The framing to keep in mind: the US is not short of money for AI evaluation — it is short of an institution that reports the results anywhere the public can see them.


Google's Project Suncatcher will put its first AI payload in orbit on October 1, when a refrigerator-size test satellite built with Planet Labs launches from Vandenberg on a Falcon 9. Google vice president James Manyika says the "MVP" craft should pack enough compute to handle basic AI queries from orbit, running on solar power with no grid connection and no land to license.

The physics has not changed: space has no air for cooling, cosmic radiation chews through chips, and Google manager Travis Beals estimates it would take around 10,000 satellites to match a single 1-gigawatt terrestrial data center. Jeff Bezos has said orbital data centers may need up to 20 years to beat ground-based ones on cost. Still, Google is now hedging in both directions at once — it signed up to supply TPUs for Anthropic's next gigawatt on Earth and, the same week, started buying an option on the sky. Treat it as a research line, not a capacity plan.

What to watch: whether anyone outside UkisAI reruns the Flash Next evals, and whether the October 1 launch slips.

Do you trust a fine-tune's own token-reduction numbers, or do you wait for a third party to rerun them? Tell us in the comments.

Sources: UkisAI Swift release (r/LocalLLaMA) · Swift 1.5 27B collection (Hugging Face) · Independent Aider eval of Swift 27B (r/LocalLLaMA) · The Washington Sun · Techmeme · Google Research · The Decoder