Black Forest Labs' 7B robot model beats models twice its size

Share
Black Forest Labs' 7B robot model beats models twice its size

Black Forest Labs took its video model into robotics and put it on top of NVIDIA's robot benchmark — with a model small enough to run on the machine itself. Also today: Databricks buys its way into spreadsheets, and India's Sarvam claims a document-AI record.

Black Forest Labs released FLUX 3 Action, an open world-action model for robots that tops NVIDIA's RoboLab-120 leaderboard at 42.9% success with only seven billion parameters. The model is a robotics sibling of the FLUX 3 video stack: it takes multi-camera video from a robot's workspace and predicts both the next action and how the environment will change as a result. On NVIDIA's own leaderboard it sits first at 42.9%, across 515 of 1,200 episodes, ahead of Cosmos3-Nano-Policy at 36.8% — a 16-billion-parameter model more than twice its size. Black Forest Labs also claims up to 3.95x faster inference than that predecessor, which is the number that matters when a policy has to keep up with a moving arm. The weights are on Hugging Face in three checkpoints — base, SO-101 and DROID — under a custom licence rather than a standard open-source one.

Why it matters: robotics is the one branch of AI where the scaling playbook has visibly underdelivered, because a policy has to run inside a control loop on hardware with a power budget. A seven-billion-parameter model winning a third-party benchmark is a bet on training data and inference efficiency instead of parameter count — the same trade HiDream-O1-Embodied made when it topped RoboColiseum's robustness chart this month at six billion. Two caveats are worth keeping. RoboLab-120 is simulation, not a robot floor, and the 42.9% belongs to the guidance-distilled checkpoint; the plain single-step version scores 38.3%. And the "less than half the size of the previous best open model" framing compares only against 16B Cosmos while a smaller open model already sits in the same table. The record is real; the size story is curated.


Databricks acquired Row Zero, a cloud spreadsheet startup whose sheets scale past the million-row ceiling of a normal one, and says more deals are coming. Terms were not disclosed. The tell is who found it first: Databricks' own finance team was already pairing Row Zero with Genie, the company's natural-language data agent, and the deal folds that workflow into the product — an agent answering business questions inside the spreadsheet analysts already know, rather than a new business-intelligence tool they have to learn. Axios reports it is the first acquisition attached to the Genie platform. Databricks has now closed five deals this year — Quotient AI and SiftD.ai in March, Panther in June, Electric in August — after raising $5 billion in August at a $190 billion valuation on a $7 billion annualized revenue run rate. Ali Ghodsi's line to TechCrunch, that the company intends to do "many more acquisitions like this," is the strategy in one sentence: buy the interfaces agents will act through.


Sarvam AI released Vision 2.1 and says its three-billion-parameter document model now sits on the Pareto frontier of OCR. The Bengaluru lab reports 87.3 on olmOCR-Bench, 94.97 on OmniDocBench v1.6 — second behind PaddleOCR-VL 1.6 at 96.01 — and 87.39 word accuracy across 22 Indian languages on its own Indic benchmark, where it claims wide margins over Gemini 3.6 Flash and GPT 6 Astra. What is checkable is the price: 0.50 rupees per page for digitisation and 1 rupee per page for structured extraction, roughly a third of what Sarvam charged earlier this year. What is not checkable is the ranking — none of the numbers have been independently reproduced, and the only non-Sarvam write-up of the release adds technical specifics that appear nowhere in the primary source. The weights are not open either; Vision 2.1 is API-only, with just the benchmark data published. Treat the scores as the vendor's own and the pricing as the story.

What to watch: whether anyone outside Black Forest Labs reproduces the RoboLab-120 score on real hardware, and whether Databricks' next acquisition is also an interface.

Would you put a seven-billion-parameter policy in charge of a robot arm today, or wait for someone outside the lab to reproduce the number? Tell us in the comments.

Sources: The Decoder · Black Forest Labs · FLUX 3 Action (Hugging Face) · NVIDIA RoboLab-120 leaderboard · VentureBeat · Databricks · GeekWire · Axios Pro · TechCrunch · Sarvam AI · Sarvam Indic OCR Bench (Hugging Face) · Indian Express