Musk wants rivals to run test harnesses on each other's models
Two stories today run the same test on AI's gatekeepers: who gets to check the work before it ships, and what the answer looks like when the checker isn't the government. Then a biology dataset turned into a toy shows what happens when the public gets the keys.
Elon Musk wants the top labs — plus "three or four of the leading Chinese companies" — to let rivals run a "test harness" on their models before release. Speaking at the All-In Summit in Los Angeles on Monday, the SpaceX chief named his own company, OpenAI, Anthropic, Google and Meta as the labs that should open their models to competitors, framing it as a substitute for self-reporting: "instead of grading your own homework, you would at least have competitors grading your homework and raising the alarm if they see concerns." He put the government at the end of the line rather than the front — in his telling, officials should step in only if a lab flags a problem and then declines to fix it.
It is a different mechanism from the one Dario Amodei put on the table last week. Amodei's proposal embeds third-party evaluators with employee-like access inside a lab, while Musk's version gives competitors a week or two of access and lets the market of rivals do the auditing. We covered the industry's move toward that inside-access model when Amodei's essay landed and Altman matched the evaluator pledge within hours — Altman matches Amodei's evaluator pledge, Musk says 'Dario is right'.
The gap between the two plans is the honest part of the story. Musk acknowledged the rival labs haven't agreed to his proposal, and none of them has offered a timeline; labs do not currently hand unreleased frontier models to competitors. Handing your next model to a rival a fortnight before launch is a real business risk, which is exactly why the ask needed a name as loud as Musk's to get heard. He added that he thinks China would sign on to the scheme, which is the one claim in the pitch that no lab or government has confirmed.
In the same 24 hours, President Trump dismissed AI safety worries as a "hoax" and the White House's National Economic Council director Kevin Hassett said the private sector is the "right place" to solve them. So the argument has moved from whether to slow down to who holds the clipboard — and every proposed referee so far is either a competitor or the company itself.
A male fruit fly's connectome has escaped the lab and taken over the internet. Within days of the Janelia and Google map going public, developers dropped the wiring diagram into game engines: a Georgia Tech computer science graduate student, Evan Sinclair Smith, wired the simulated fly brain into Minecraft and posted the result on September 4 — millions of views inside a day — while another developer overfit its motor system to make it "play" Beat Saber, and others ran it through Doom. One modder, seeing the fly repeatedly fed into virtual danger, built a gentler project he calls Fruit Fly Heaven.
The caveat matters more than the videos. The connectome supplies wiring, not experience — developers still choose what sensory data enters the network and how firing turns into movement, and the Beat Saber run reproduces a recorded movement sequence rather than a decision. Smith has said the point is not a conscious fly. This is our second pass at whole-brain emulation in three weeks; the platform ambitions behind it were the story in August — China's DeepSoma simulates whole brains inside real physical worlds. What changed is that the public now has the raw material and the tools, and the results look like entertainment.
Profound raised $180 million at a $1.8 billion valuation to sell brands visibility inside AI answers. Sequoia and Kleiner Perkins led the Series D, seven months after a $96 million round, with Lightspeed, Khosla Ventures and South Park Commons returning. The company says revenue tripled in six months and it now counts more than 1,000 enterprise customers, including Comcast, Estée Lauder and Walmart; it has not disclosed the actual revenue figure, which makes the valuation a bet on a category rather than a number.
The product tracks how often ChatGPT, Perplexity, Gemini and Google's AI Overviews name a brand and in what context — the answer-engine equivalent of a rank tracker, for a search world that no longer shows a list of links. Profound is also pointing the same machinery at agents, monitoring the automated traffic reading customer websites, and it staffs "forward-deployed marketing engineers" alongside its own agents. The uncomfortable dependency is baked in: the whole business assumes models keep mentioning brands in ways that can be measured and nudged, and Profound has no vote in how they do that.
What to watch: whether any lab actually agrees to open an unreleased model to a rival on Musk's terms — the first yes is the story, not the proposal.
Should competitors, not regulators, be the ones who see your model first? Tell us in the comments.
Sources: CNBC — Musk urges labs to test each other's models · Techmeme — Musk on test harnesses and Chinese labs · TeslaNorth — Musk's peer-review pitch at All-In Summit · 404 Media — A digital fly brain has taken over the internet · IBTimes UK — developers run the fly brain in Minecraft and Beat Saber · Google Research — the male fruit fly brain map · TechCrunch — Profound hits unicorn valuation with $180M Series D · TNW — Profound's $180M round for visibility in AI answers · Bloomberg — Profound hits $1.8 billion value