UK's AI Security Institute staff signed off work with stress
Two stories worth your attention: the people who test frontier models are taking the damage home, and China's domestic silicon just put a 2.8-trillion-parameter open model on a single machine.
Multiple staff at the UK's AI Security Institute have been signed off work with stress and are undergoing counselling, the Financial Times reported on Tuesday — and the toll reaches past one agency. The FT says top researchers at AISI, OpenAI, Anthropic and Google DeepMind describe burnout and stress from the work itself: testing unreleased systems on schedules set by companies racing to ship them.
Inside AISI the pressure has a specific source. Its cyber-security and bio-chemistry teams have raised alarms about models finding previously unknown software vulnerabilities and, more troublingly, generating novel biological threats, and a former employee described the atmosphere as stressful and at times toxic. A reorganisation in May folded the societal resilience team into the human impacts unit, cutting that group from roughly 15 researchers to three, and the head of the merged-out team resigned in July. The government has said it will protect staff wellbeing. The awkward part is structural: AISI's evaluations are voluntary, so it cannot compel a lab to hand over a model or block a release — the people absorbing the worst of what the tests find have the least leverage to act on it.
This is the second safety-staff story in a week — we covered the last one here: Google DeepMind safety researcher quits, says AI 'could kill us all'.
Inspur launched the Yuanbrain SD200 Ultra at the AICC 2026 computing conference in Hangzhou, a supernode it says runs Moonshot's Kimi K3 — a 2.8-trillion-parameter open model — on a single machine at under 5.85 milliseconds per token. The box packs 128 domestic AI chips into a tightly coupled mesh with 8 TB of unified GPU memory and 64 TB of system memory, at an interconnect latency the company puts at 0.69 microseconds; Inspur says it cut AllReduce communication time by 3.5 times and supports models up to 10 trillion parameters in one machine.
The more interesting engineering is in the tuning: Inspur built a super-operator agent out of Kimi K3 itself to fuse the basic operators across KDA, gated MLA and MoE, which it says cut the operator count tenfold and lifted inference throughput more than three times. A companion HC2000 rack claims ten times the token capacity for the same investment. These are vendor numbers from a vendor stage, with no third-party benchmark behind them yet — but the packaging matters more than the latency claim, because it turns a rack-scale serving problem into a single-box purchase. We covered the last move in this direction in August — Alibaba Cloud's Zhenwu M890 supernode runs 2T-parameter models.
What to watch: whether Westminster gives AISI statutory powers over model access, and whether anyone outside Inspur reproduces that 5.85-millisecond figure.
Should AISI be able to compel access to frontier models, or is voluntary testing the right trade-off? Tell us in the comments.
Sources: Financial Times · Crypto Briefing · Techmeme · ITHome · Jiemian News · Zhidx