DeepSeek's V4-Flash-Vision rivals Opus 4.8 on agent tests

Share
DeepSeek's V4-Flash-Vision rivals Opus 4.8 on agent tests

DeepSeek keeps moving fast on its low-cost track, and its latest move pairs image understanding with the company's signature price. Today it unveiled V4-Flash-Vision Exp, an experimental multimodal model that adds vision to its V4 Flash series, and DeepSeek says it stacks up against Anthropic's flagship on visual-agent benchmarks.

DeepSeek's V4-Flash-Vision Exp nearly matches Opus 4.8 on multimodal agent tests.

The new model, live on DeepSeek's API platform as deepseek-v4-flash-vision-exp, layers vision on top of the text reasoning that V4 Flash already shipped. On DeepSeek's own internal multimodal-agent benchmarks, the company says the vision variant scores close to Claude Opus 4.8 — and in at least one widely shared framing it edged ahead. The Decoder notes the model is built for agent workflows: it can describe images, pull text out of screenshots, and analyze diagrams, handling JPEG, PNG, GIF, and WebP formats. It's designed to drop into existing agent frameworks and combine what it sees with tool use, the use case DeepSeek is clearly betting on.

On the numbers, V4-Flash-Vision Exp beats its predecessor across six of seven text benchmarks, and the biggest jump is in image analysis. DeepSeek measured more than 10% improvement on two of four visual tests, and bested Opus 4.8 on two specific evals: ALE, a suite of more than 1,000 multi-step tasks where models interact with apps, write code, and interpret files, and ZeroBench, a set of 100 intentionally hard image-analysis tasks. For now the bargain framing is the whole pitch — images are tokenized at up to 384 tokens each and billed at V4-Flash pricing, and DeepSeek says a fuller "Vision-Full" release is coming down the line, with nothing announced yet on open weights.

DeepSeek is reportedly preparing a listing that CNBC-tracked reports call a potential record IPO.

The model drop lands against a steady drumbeat that DeepSeek is readying the biggest AI IPO in history, reportedly in the tens of billions of dollars in valuation. A 21财经 report framed it as an AI giant about to file its listing application, with coverage pointing to a number that would make it the largest-ever AI company offering — several outlets put the buzz around an $86 billion valuation. DeepSeek hasn't confirmed specifics, so treat the number as reported rather than final, but the direction is clear: the outfit best known for cutting the cost of frontier AI is now looking to raise serious public capital. That pairs oddly but powerfully with today's release — ruthless pricing as a pre-IPO growth story.

What to watch: whether the open-weights version arrives, since that's the move that would really put pressure on the paid tier's moat.

Open weights on a model this cheap would reset the multimodal race again — do you think DeepSeek releases them, or keeps vision closed for the IPO? Tell us in the comments.

Read more

Lambda raises up to $4B from Blackstone ahead of its IPO

Lambda raises up to $4B from Blackstone ahead of its IPO

The neocloud money is consolidating fast, and today's inbox shows both ends of the market: a heavyweight pre-IPO round on one side, and a Google open model you can run on a phone on the other. Lambda is raising up to $4 billion led by Blackstone and Coatue at a $14.5 billion pre-money valuation — its last private round before a planned IPO. The Wall Street Journal reported the scoop from a letter to limited partners, and Reuters independently confirmed the headline terms: the round is led by t

South Korea bets $3.49B on its own frontier AI model

South Korea bets $3.49B on its own frontier AI model

Sovereign-model money is getting serious, and the hardware money is following it. Today's inbox: Korea's nine-figure upgrade to its homegrown model push, a physics-simulation startup priced like a chip designer, and Google turning a geospatial model loose on public health. South Korea is putting 4.7 trillion won — about $3.49 billion — of state equity behind a homegrown frontier AI model. The Ministry of Science and ICT confirmed the figure as part of its proposed 2027 budget, split into two t

Mistral's Le Chonk puts Europe's sovereignty bet on a download date

Mistral's Le Chonk puts Europe's sovereignty bet on a download date

Mistral's biggest model ever is real, benchmarked and for sale today — but the thing that makes it matter to Europe's sovereignty argument, the weights, is still three weeks out. The preview settles who built it; the release will settle whether it counts. What Mistral actually shipped Mistral opened a public preview of Mistral Large 4 — unofficially ML4, officially le Chonk — a 1 trillion-parameter mixture-of-experts model with 49 billion active parameters and native multimodal input. The p

Mistral unveils Le Chonk: a 1T-parameter open-weights model

Mistral unveils Le Chonk: a 1T-parameter open-weights model

The biggest open-weight release outside China lands in public preview today, and the country that spent the week promising its own frontier model just put a price on the ambition. Mistral has opened a public preview of Mistral Large 4 — codenamed "le Chonk" — a 1 trillion-parameter mixture-of-experts model with 49 billion active parameters, natively multimodal, which the company calls its largest and most capable model to date. The preview API is live today on Mistral Studio at $1.36 per milli