Silvia tax agent beats frontier models, opens benchmark
Two stories today about AI moving into specialist territory: a finance agent that claims to beat every frontier model on tax, and a Y Combinator startup betting agent swarms can find materials for cooler chips.
ProCap Financial's Silvia, a finance-only AI agent, claims it outperformed every frontier and open-source model it tested on expert-level tax questions — and it open-sourced the benchmark to back the claim. ProCap (Nasdaq: BRR), which bills itself as the first publicly traded agentic finance firm, said its team evaluated seven AI products across ten expert-level tax scenarios covering federal law plus the California, Texas, North Carolina, New York and Florida codes. Responses were pulled from each product's own knowledge sources and scored by a blind judge on a 1–10 scale for factual accuracy under four escalating levels of verification; Silvia, which retrieves answers from a proprietary library of primary tax law, posted an 8.73 overall — ahead of the next-best product, Claude Desktop at 8.40, and first in every scenario tested. Chairman and CEO Anthony Pompliano stressed the margin came from a team of four AI engineers: "The largest AI labs on earth employ more researchers than we have employees, and Silvia beat all of them on tax."
The open-sourcing is the part worth watching. ProCap released the evaluation questions and framework as the Silvia Tax Bench on Hugging Face — 200 questions spanning federal and five state jurisdictions, with per-configuration scores attached — so anyone can rerun the test or contest the result. That openness is also the right frame for the headline number: this is a self-published benchmark scored by a blind LLM judge, not an independent audit, so the 8.73 reads as a strong claim rather than a settled verdict. The bigger signal is the strategy debate Pompliano is staking out — whether domain-specific systems with proprietary data can take share from general-purpose frontier models. With more than $50 billion in assets connected to Silvia across over 20,000 users, the market is already voting.
Discovered Materials, a Y Combinator startup, raised $9 million to hunt for new materials that could make chips run cooler and more efficient — using swarms of AI agents. The seed round was led by Lightspeed India Partners, with Peak XV Partners and angel investors including Paul Graham. Co-founders Advaith Sridhar and Akash Ramdas built a pipeline that uses Anthropic models in a custom harness to generate material leads, then runs physics models they trained to simulate and verify whether candidates are worth pursuing. Ramdas, whose Stanford doctorate centered on semiconductor materials, said the system now makes "thousands of guesses a day" versus the roughly 20 he could test by hand during his PhD — and the company released hundreds of new materials plus a "Material Discovery Bench" today to track how frontier models handle the challenge. The honest caveat, which its own backers concede, is that no AI-discovered material has reached commercial deployment yet; Lightspeed's Hemant Mohapatra argues the bottleneck is filtering and synthesizing candidates, not finding them.
What to watch: whether anyone reproduces or refutes Silvia's tax benchmark, and whether Discovered Materials can move candidates from simulation into real silicon.
Would you trust a specialist AI agent over a frontier model for your taxes — and should self-published benchmarks count without an independent audit? Tell us in the comments.
Sources: ProCap Financial (Business Wire) · Las Vegas Sun · StockTitan · Silvia Tax Bench (Hugging Face) · TechCrunch · Discovered Materials (Y Combinator)