DeepMind alumni's Inherent says its Faraday agent beats GPT-5.5 at reproducing research

Share
DeepMind alumni's Inherent says its Faraday agent beats GPT-5.5 at reproducing research

A London lab founded by Google DeepMind alumni says its new agent, Faraday, outperformed Claude Opus 4.8 and GPT-5.5 at independently reproducing the findings of published scientific papers — and it did it on a 27B-parameter Qwen 3.6 model, a fraction of the frontier systems' size. Inherent emerged from stealth just weeks ago with a $50 million seed round, and this benchmark claim is its first real proof point.

The company is upfront that the task was narrow: replication, not discovery. But that's the right training ground — reproducing a paper without knowing the answer in advance demands reading methods, forming hypotheses, designing experiments and interpreting results, which is roughly how human PhD students cut their teeth. Chief scientist Edward Hughes told TechCrunch the interesting part was not beating frontier agents but how they got there: rather than training primarily on studies of how science is conducted, Inherent leans on reinforcement learning, rewarding good experimental outcomes so the behavior generalizes toward its actual goal — agents that can discover new knowledge, not just verify old results.

The claim deserves skepticism until independent runs confirm it, since it comes from the company itself. Still, the shape of the result matters more than the score: if research taste can be taught to small models through rewards instead of scale, the moat around frontier labs gets thinner. We covered a related signal earlier today — Mystery model Ox Alpha tops GPT-5.6 — where a smaller challenger again embarrassed bigger systems.


Your local LLM isn't dumber than the API version — your stack is quietly sabotaging it. A viral Level1Techs forum post (now on the Hacker News front page) ran controlled experiments on Qwen3.6-27B showing that changing only the attention backend inside an inference engine — same GPU, same weights, same prompt — flips enough next-token probabilities to derail tool calls entirely: one run dialed the wrong Cisco interface, then executed the wrong command twice trying to recover.

The rot compounds as context grows. With an int4 KV cache the model's "IQ drops like a rock" past roughly 40k tokens and a recoverable mistake becomes unrecoverable; across five weight formats, Nvidia's own NVFP4 release came dead last with about half of sampled tokens flipped by 88k of context, while a community INT8 build beat the official FP8 checkpoint. The author even found tensor parallelism flipping results — two GPUs failed a call that one and four GPUs passed.

The takeaway for anyone running local inference: treat quantization and runtime configuration as accuracy decisions, not just speed knobs, and benchmark your own stack on long, tool-heavy workloads — not three zero-shot prompts. Labs publishing "impossibly low" divergence numbers for quants without methodology disclosure deserve the side-eye too.

What to watch: whether anyone replicates Inherent's Faraday numbers independently — and whether the Level1Techs team ships its promised distributable testing rig so you can measure your own stack's drift.

Running local models for real work? Have your own token-flip horror stories — or a quant you swear by? Tell us in the comments.

Read more

Underdog launches a private on-device AI assistant, backed by a16z

Underdog launches a private on-device AI assistant, backed by a16z

The privacy split in consumer AI got a new entrant tonight, GitHub turned code-review benchmarks into a vendor-bias argument, and OpenAI opened its oddest API to everyone. Underdog launched in invite-only beta: an on-device AI assistant that keeps your data on your machine and never charges you a subscription. Self-taught coder and Thiel Fellow Sigil Wen — who moved to Silicon Valley at 17 and lived in an AI hacker house with Andrej Karpathy — built his own inference engine, Husky, to run a 2

Today in AI — October 6, 2026

Today in AI — October 6, 2026

A day of second-order moves: the labs are buying task data from software vendors instead of scraping the web, Waymo is borrowing to fund the robot world, and the biggest bank in the US just put a number on what Anthropic's latest model cost it in risk. Models & Research * OpenAI is training GPT-6 Astra on Ironclad's real contracting work. Ironclad staff helped turn 11 tasks across legal, commercial and procurement work — setting up NDAs, approval workflows, clauses that change by jurisdict

Meta, Walmart and Stripe publish the Personal Agent Protocol

Meta, Walmart and Stripe publish the Personal Agent Protocol

The agent economy is writing its rulebook tonight: one open standard for AI bots at the checkout, a cheaper image model from Google, and Anthropic turning bug-hunting into a tiered product. Meta, Walmart, Stripe and Sierra are publishing an open "personal agent protocol" — a standard that defines how personal AI agents interact with businesses online. The group behind it reads like a cross-section of agentic commerce: Meta, Sierra, Walmart, Stripe, Shopify, Genesys, Rocket, NiCE, Decagon and I

Lambda raises up to $4B from Blackstone ahead of its IPO

Lambda raises up to $4B from Blackstone ahead of its IPO

The neocloud money is consolidating fast, and today's inbox shows both ends of the market: a heavyweight pre-IPO round on one side, and a Google open model you can run on a phone on the other. Lambda is raising up to $4 billion led by Blackstone and Coatue at a $14.5 billion pre-money valuation — its last private round before a planned IPO. The Wall Street Journal reported the scoop from a letter to limited partners, and Reuters independently confirmed the headline terms: the round is led by t