Mostik's bridge puts a 753B model's thinking inside a 4B phone model
Two stories from the same frontier: a startup arguing that the next capability jump comes from wiring models together rather than scaling them up, and a Chinese robotics lab arguing that a world model is only useful if it knows what the robot's own actions did.
Mostik, a four-month-old startup whose chief scientist is 2010 Fields Medalist Stanislav Smirnov, says it has built a translation layer that lets a 753-billion-parameter GLM-5.2 pass its internal state straight into a 4-billion-parameter Qwen-3.5 — with no text and no fine-tuning on either side. The pair is leading the ARC-AGI 3 Kaggle competition, according to Chinese coverage of the result. Mostik is withholding the technical details until the contest closes, but it has published the headline numbers: the bridged small model closes roughly half the gap to the large one, lifts its own accuracy by 25%, and doubles it on the hardest problem subset where the size gap is widest — at about one-twentieth the large model's inference cost.
The framing is worth more than the score. Every multi-model system today — coding sub-agents, model councils, routers — passes 17 bits of chosen token between models while discarding roughly two megabytes of internal state per token. Mostik's bet is that the discard, not the model, is the bottleneck, and that interpretability work over the last two years says the thrown-away state is where the actual reasoning lives. Both models stay frozen in the demo, which is the strict version of the test: if information survives that, the shared structure came from training, not from the bridge. Skepticism is warranted — these are the company's own numbers on one pair of open-weight models, and Smirnov himself says the field lacks the mathematical language to describe why representations transfer at all. But if it generalizes, it makes every open-weight release a component rather than a product, and moves leverage away from whoever trains the biggest model.
Guangxiang Technology's ActEffect is a bet that robots need to model intervention, not just prediction — what the world does after this gripper closes, not what the next frame looks like. The Beijing lab, spun out of Tsinghua, says its model separates changes the environment causes on its own from changes the robot's own actions cause, then distills that understanding into the policy so nothing extra runs at deployment time. It reports 98.8% average success on the LIBERO manipulation benchmark, 80.3% on the perturbation-heavy LIBERO-PLUS set, and 67.5% across 1,200 trials of RoboCasa-GR1's two-armed dexterity tasks. Its Phi-Bot X1 industrial robot is already validated at real workstations at more than one luxury automaker, with deployments expected before year-end.
Guangxiang's own caution is the most useful part of the release: its researchers say benchmarks still have to be settled on real production lines, measured in task success and customer value. That is the right instinct for a field where at least ten embodied-AI companies shipped a world model between May and August, and where the term is quickly becoming table stakes rather than a differentiator. The dividing line ahead is not who can predict the future but who can answer the counterfactual — a robot that has compressed weight, gravity and friction into reusable rules should adapt to the hundredth unseen object far faster than one that has merely fitted the first fifty.
What to watch: whether Mostik publishes after the Kaggle deadline and lets anyone reproduce the bridge on a second model pair, and whether ActEffect's benchmark numbers survive contact with a real shift.
If a small model can borrow a big one's reasoning at one-twentieth the cost, does the frontier-lab advantage start to erode? Tell us in the comments.
Sources: 量子位 (QbitAI) · AI Chat Daily · 科技区角 (X-TechCon) · 智东西 (Zhidx) · 界面新闻 (Jiemian)