Doubao 2.1 Pro clears 83% of a 387,000-line repo's open issues
ByteDance spent the last 24 hours upgrading three layers of the same stack at once: the model, the workplace suite it lives in, and the agent that ties them together.
Volcano Engine pushed Doubao 2.1 Pro to a 0915 snapshot, and the number it is selling is a repair run: multiple sub-agents worked the open-source game Luanti — roughly 387,000 lines of code and 1,000 real historical issues — in parallel for 36 hours and brought 83% of those issues to merge-ready standard. The API is live on Volcano Ark under both a locked 0915 version and an auto-updating Evolving alias, TRAE has integrated it, and Doubao Work users can pick it from the model menu. On the agent side the company claims stronger evidence tracing, authoritative source retrieval, timeliness judgement and data verification, with a demonstration of more than 500 sub-agents cross-checking over 1,000 web pages — financial reports, production capacity, fleets, hiring — into a traceable due-diligence report that folds in satellite imagery and maritime data as outside evidence. Multimodal coding is the second thread: the model reads design files, drawings and screen recordings and turns them into front-end code and game logic, and ByteDance says image and video inference token consumption fell by more than 30% against the previous generation.
Two caveats belong next to those numbers. They are the vendor's own, and as one Chinese tooling site noted, no third-party evaluation of the 0915 snapshot exists yet. The second is that the 83% figure is a statement about a harness as much as a model — parallel sub-agents on a frozen public repo, with no published cost per fix. The part worth taking seriously is the shift in what labs are selling: not a benchmark score, but an orchestration recipe, and ByteDance now ships the model, the harness and the office suite that consumes it.
Feishu 8.0 opened its data and tools to agents and launched Doubao Work Partner, which ByteDance bills as the first team agent in China — it is added to a group chat like a colleague, carries its own identity and permissions, and remembers what was said in meetings, documents and threads. It can call Feishu documents, multi-dimensional tables, meetings, calendars, cloud drives and enterprise business systems to sync information, review documents, assemble weekly reports and read code, and it is meant to escalate on its own when a project stalls. Administrators configure capability, data scope and usage ceilings the way they would for an employee, and the permission model is inherited rather than bolted on: anything a staffer cannot see, the agent cannot fetch. ByteDance says the product is in targeted co-creation with enterprises before a gradual rollout — we covered the launch of the underlying product in August — ByteDance ships Doubao Work, an AI agent wired into Feishu. The unproven half is adoption: an agent that lives where the work happens is only as useful as the enterprises willing to hand it a seat.
LatticeFlow AI released an independent framework for measuring political bias in language models, built to avoid the two things that make such tests contestable — a human-written neutrality rubric and another model acting as judge. Instead it compares how Chinese and Western models answer the same questions, decomposes each answer into individual claims and measures where they agree and diverge, so the political axis emerges from the data, with every result pinned to a SHA-256 hash of the exact weights tested and provider-side moderation switched off. The first run, about 554 samples per axis, found alignment strengthening with scale: Qwen 3.7 Max sits further toward the Chinese pole than the smaller Qwen3 32B across all six China-politics categories, and on religion and ethnic issues it is the most Chinese-aligned model tested. GLM 5.2, Kimi K2.6, Qwen 3.7 Max, MiniMax M2.7 FP4 and DeepSeek V4 Pro cluster at the same end. The failure mode is reframing rather than refusal — the model answers fluently and the framing shifts. This lands while Western enterprises are adopting Chinese open-weight models precisely for cost — Chinese models were 30–46% of US company token usage on OpenRouter this year, up from 4.5% in the first half of 2025, and they run 60–90% cheaper — a procurement shift we have tracked through buyer behaviour like AT&T swings 40% of its AI calls to open models to cut costs. Keep two things separate: the methodology is more defensible than what came before it, and the publisher is a vendor selling risk-control software.
What to watch: whether the 0915 snapshot's repair claims survive an independent repository-scale benchmark, and whether LatticeFlow's hash-pinned results get reproduced by anyone who does not sell governance tooling.
If an enterprise buys a cheaper model knowing its answers reframe politically, is that an informed trade or a liability nobody has priced yet? Tell us in the comments.
Sources: Leiphone 雷峰网 · IT之家 · QbitAI 量子位 · 时代财经 via 腾讯新闻 · AIBase · LatticeFlow AI — Political Bias in LLMs · Business Wire