ByteDance founder bans distillation, putting model integrity over speed

Share
ByteDance founder bans distillation, putting model integrity over speed

A major Chinese tech leader just drew a hard line on AI shortcuts, while global model pricing gets bloodier and a viral AI tool gets exposed as fraud.

ByteDance founder Liang Rubo has told staff to avoid using distillation on the company's AI models, even if it means slower development. According to a mid-year meeting memo reported by the Paper and picked up by the Information, Liang acknowledged that ByteDance's large language models are behind competitors — but said the company won't use distillation to game benchmarks or inflate rankings. Distillation, the practice of training smaller models on outputs from larger ones, has become widespread in China's AI ecosystem as companies race to close the gap with frontier labs. Liang's stance is striking because it rejects a shortcut that virtually every major Chinese AI player has used at some point. The statement signals that ByteDance is betting on building genuine model capability rather than chasing leaderboard positions — a bet that could pay off long-term but leaves the company exposed in the near term as rivals ship increasingly capable models.


OpenAI and Anthropic are now locked in a full-blown price war as Chinese AI models flood Silicon Valley's market. The Financial Times reports that token costs are plummeting across the board, with Chinese models offering comparable performance at a fraction of the price. TechCrunch notes that OpenAI priced GPT-5 aggressively enough to spark a broader race to the bottom, while the Washington Post highlights how open-weight models from companies like Zhipu and Moonshot are eroding the pricing power of US labs. The Axios headline says it plainly: AI may never be as cheap as it is today. For developers and enterprises, this is a golden window — inference costs are collapsing faster than anyone predicted. For the labs themselves, it's a margin squeeze that forces a choice between burning cash to maintain market share or ceding ground to cheaper alternatives.


The J-Space "Cognition Suite" — an open-source tool that claimed to make small models rival frontier LLMs — has been exposed as fraudulent, reigniting the debate over whether engineering tricks can replace raw model scale. The tool, which went viral on X, claimed that pairing DeepSeek V4 Flash with an external harness could match or exceed Claude Fable 5 on agent tasks while being 2.53x faster. Independent testing by GitHub user GoForceX found the opposite: performance dropped, token consumption rose, and inference slowed. Jason Wei, the OpenAI researcher who invented chain-of-thought prompting, weighed in with a detailed thread arguing that small models with external tools fundamentally cannot replicate the internalized reasoning of large models. DSPy creator Omar Khattab backed the argument from an academic angle, showing mathematically that recursive tool-calling cannot approximate attention-based global reasoning. The debate cuts to the heart of AI's scaling question: is the path to better performance through bigger models, or through cleverer engineering around smaller ones?

What to watch: Whether ByteDance's anti-distillation stance holds as competitors continue to ship distilled models, and whether the price war forces any US labs to reconsider their pricing strategies.

Tell us in the comments.

Sources: The Information · Memeburn · The Paper · Financial Times · Washington Post · TechCrunch · Axios · Leiphone · 36 Kr · Medium