LLM4MIP says it closed 34 open optimisation problems — MIPLIB lists 29 as incumbents
A research site says an LLM workflow closed 34 open optimisation problems. The benchmark that would have to agree lists 29 entries from the same team — and none of them as proven optimal. The gap is the story.
An LLM-assisted workflow now claims 34 of 132 open MIPLIB instances "resolved" — but MIPLIB's own log carries 29 of those submissions, all as incumbent improvements, not proofs. The work is LLM4MIP, published as a research website on 21 September by Yicheng Huang and Wenzhi Gao as co-lead authors with Dongdong Ge, Madeleine Udell and Yinyu Ye — a group with real optimisation pedigree, not a weekend prompt. Their claim splits into 32 instances they say they solved to optimality and two they say they proved infeasible, drawn from a working set of 132 instances chosen out of 217 still open in MIPLIB, the standard public library for mixed-integer programming.
The independent check is more modest, and it is public. MIPLIB's news log records a release on 24 September with 51 better incumbents, two instances moved to "easy", three moved from optimal to hard, and one updated to infeasible. LLM4MIP's solutions appear in the best-known-solution records for 29 instances — 25 improved incumbents plus four instances given a first feasible solution for the first time. That is a genuine contribution to a library where progress is measured in single instances per release, and it is also not 34 resolved problems: an accepted incumbent is a better answer, not a proof of the best answer.
There is no preprint and no peer review — the authors say a technical report is forthcoming — and no third-party replication has surfaced. That matters less than it sounds, because the verifiable layer already exists: MIPLIB's maintainers at Zuse Institute Berlin ran their own feasibility checks before listing the entries, and independent records show no constraint or objective violations. What a reader should take from this is the shape of the result, not the count. Solvers plus a model reasoning about structure is an effective workflow; "LLMs are doing mathematics" is the version that will not survive review.
miHoYo's president says the Genshin Impact maker will be a heavyweight in Chinese large models within two to three years — and it has quietly started building a coding model. Liu Wei made the claim at a Shanghai Jiao Tong University recruiting talk on 23 September, telling students he is "very confident" and inviting them to hold him to it; the quote surfaced widely on 28 September as a 澎湃新闻 exclusive, though QbitAI and GameLook had it days earlier. The substance underneath is what makes it worth noting: miHoYo has run an AI unit called 逆熵 since 2018, filed a generative-AI service named Glossa with Shanghai's cyberspace authority in September 2024, and Liu has put a ceiling of RMB 100 billion over three years on the company's AI spending.
Cloudflare's Matthew Prince expects automated traffic to hit 1,000 times human traffic within five years, and wants HTTP 402 to make bots pay for what they read. In a Decoder interview with Nilay Patel, Prince said bots passed human traffic in May 2026 — months ahead of his own prediction — and extrapolated to a 1,000x ratio while flagging that "I've gone wrong so far in every prediction that I've made on this". His proposed fix revives the original Netscape "payment required" status code, with publishers collecting "a fraction of a penny" per access and the money coming from what people pay for their AI agents, in his framing "similar to how Spotify or Apple Music works". Cloudflare says it is working with Coinbase and Stripe on the rail; its own documentation still describes pay-per-crawl as a closed beta with no published adoption numbers. Identity is the unresolved part — a payment protocol is only as good as the ability to tell a paying agent from a forged name, which is exactly what we found most sites cannot do when seven in ten waved a fake GPTBot straight through.
What to watch: whether the LLM4MIP authors publish the technical report with the two optimality claims their site still carries, and whether MIPLIB upgrades any of those 29 instances from incumbent to solved.
Which number should a benchmark believe — the submitters' count or the maintainers' log? Tell us in the comments.
Sources: LLM4MIP — How much can LLMs help solve MIPs? · MIPLIB 2017 News Log · MIPLIB Changelog · 澎湃新闻 (The Paper) · QbitAI (量子位) · GameLook · The Verge — Decoder: Cloudflare's Matthew Prince · Cloudflare — Introducing pay per crawl