Black Forest Labs' 7B robot model beats models twice its size
Black Forest Labs took its video model into robotics and put it on top of NVIDIA's robot benchmark — with a model small enough to run on the machine itself. Also today: Databricks buys its way into spreadsheets, and India's Sarvam claims a document-AI record.
Black Forest Labs released FLUX 3 Action, an open world-action model for robots that tops NVIDIA's RoboLab-120 leaderboard at 42.9% success with only seven billion parameters. The model is a robotics sibling of the FLUX 3 video stack: it takes multi-camera video from a robot's workspace and predicts both the next action and how the environment will change as a result. On NVIDIA's own leaderboard it sits first at 42.9%, across 515 of 1,200 episodes, ahead of Cosmos3-Nano-Policy at 36.8% — a 16-billion-parameter model more than twice its size. Black Forest Labs also claims up to 3.95x faster inference than that predecessor, which is the number that matters when a policy has to keep up with a moving arm. The weights are on Hugging Face in three checkpoints — base, SO-101 and DROID — under a custom licence rather than a standard open-source one.
Why it matters: robotics is the one branch of AI where the scaling playbook has visibly underdelivered, because a policy has to run inside a control loop on hardware with a power budget. A seven-billion-parameter model winning a third-party benchmark is a bet on training data and inference efficiency instead of parameter count — the same trade HiDream-O1-Embodied made when it topped RoboColiseum's robustness chart this month at six billion. Two caveats are worth keeping. RoboLab-120 is simulation, not a robot floor, and the 42.9% belongs to the guidance-distilled checkpoint; the plain single-step version scores 38.3%. And the "less than half the size of the previous best open model" framing compares only against 16B Cosmos while a smaller open model already sits in the same table. The record is real; the size story is curated.
Databricks acquired Row Zero, a cloud spreadsheet startup whose sheets scale past the million-row ceiling of a normal one, and says more deals are coming. Terms were not disclosed. The tell is who found it first: Databricks' own finance team was already pairing Row Zero with Genie, the company's natural-language data agent, and the deal folds that workflow into the product — an agent answering business questions inside the spreadsheet analysts already know, rather than a new business-intelligence tool they have to learn. Axios reports it is the first acquisition attached to the Genie platform. Databricks has now closed five deals this year — Quotient AI and SiftD.ai in March, Panther in June, Electric in August — after raising $5 billion in August at a $190 billion valuation on a $7 billion annualized revenue run rate. Ali Ghodsi's line to TechCrunch, that the company intends to do "many more acquisitions like this," is the strategy in one sentence: buy the interfaces agents will act through.
Sarvam AI released Vision 2.1 and says its three-billion-parameter document model now sits on the Pareto frontier of OCR. The Bengaluru lab reports 87.3 on olmOCR-Bench, 94.97 on OmniDocBench v1.6 — second behind PaddleOCR-VL 1.6 at 96.01 — and 87.39 word accuracy across 22 Indian languages on its own Indic benchmark, where it claims wide margins over Gemini 3.6 Flash and GPT 6 Astra. What is checkable is the price: 0.50 rupees per page for digitisation and 1 rupee per page for structured extraction, roughly a third of what Sarvam charged earlier this year. What is not checkable is the ranking — none of the numbers have been independently reproduced, and the only non-Sarvam write-up of the release adds technical specifics that appear nowhere in the primary source. The weights are not open either; Vision 2.1 is API-only, with just the benchmark data published. Treat the scores as the vendor's own and the pricing as the story.
What to watch: whether anyone outside Black Forest Labs reproduces the RoboLab-120 score on real hardware, and whether Databricks' next acquisition is also an interface.
Would you put a seven-billion-parameter policy in charge of a robot arm today, or wait for someone outside the lab to reproduce the number? Tell us in the comments.
Sources: The Decoder · Black Forest Labs · FLUX 3 Action (Hugging Face) · NVIDIA RoboLab-120 leaderboard · VentureBeat · Databricks · GeekWire · Axios Pro · TechCrunch · Sarvam AI · Sarvam Indic OCR Bench (Hugging Face) · Indian Express