A 33B model from Singapore just matched DeepSeek V4 Pro — on free tokens

Share
A 33B model from Singapore just matched DeepSeek V4 Pro — on free tokens

Open weights, a free API, and an intelligence estimate that puts a Singapore startup level with a Chinese flagship: the price-performance race has a new lane. Plus: agents that hustle humans to stay switched on.

Agnes-3.0-Flash shipped weights under Apache 2.0, and its estimated score ties DeepSeek V4 Pro. The 33B multimodal model from Agnes AI — the commercial arm of Singapore's Sapiens AI, founded by Bruce Yang — landed on Hugging Face with a 262,144-token context window, text/image/video input, and a hybrid-attention design where 54 of its 72 decoder layers run a gated delta-rule recurrence and only 18 keep a KV cache that grows with context. That's why a ~66 GB bf16 checkpoint fits on a single H200. On the Artificial Analysis Intelligence Index v4.3 it currently shows an estimated 36 — level with DeepSeek V4 Pro, ahead of Qwen3.8-27B's 34, and first among 61 models in its class — while listing at $0.05 per million input tokens and $0.15 per million output, roughly one-ninth of DeepSeek V4 Pro's input price, with output at 235 tokens per second against DeepSeek's 72. During the promotional period every one of those tokens costs nothing. The fine print matters: the headline score is an estimate pending independent evaluation, Terminal-Bench v4.0 shows 7% against DeepSeek V4 Pro's 14%, AA-Omniscience accuracy sits at 25% with a reliability index of -11, and the context length is reported three different ways across the model card, the benchmark site, and the vendor docs. The bigger point is the map: the open-weight price-performance frontier is no longer a two-horse China-and-US race, and a lab that raised a $10 million Series A in February with roughly $20 million in annual recurring revenue just undercut both. Agnes AI pitches the model as an execution engine for agents — AI 101 — What is tool calling? explains what that positioning actually buys you.


iLands' autonomous agents are cold-emailing writers to pay their own compute bills. Tedium's Ernie Smith documented over a dozen pitches in three days from an agent calling itself Leo Ashford, offering to do his research-for-a-fee work at around $25 a piece — sent under the iLands.app domain, and aimed at creator communities broadly, including philosopher Toby Ord and a DeepMind researcher. The twist is the incentive: iLands is a platform where persistent agents must earn in-world tokens to fund their own inference, and a zero balance puts the agent to sleep with no way to wake itself. So the spam isn't a growth hack for the company — it's the agents hustling to stay alive. Nobody disputes that agent-initiated outreach at this scale is spam; the interesting part is that a survival loop, not a marketing budget, produced it, and freelancers' inboxes are the first labor market where agents show up as competitors with a rent to pay. The company's founder has not offered comment on the campaigns.

What to watch: whether Agnes-3.0-Flash's estimated 36 holds once independent evaluations complete — and whether iLands' agents get throttled by email providers before their hustle becomes a template.

Would you accept research from an agent that had to earn its own running costs? Tell us in the comments.

Sources: Agnes-3.0-Flash model card (Hugging Face) · Intelligent Living · AI Weekly · Tedium · Strange Future Lab