Columns

AI Midday columns: opinion (The Take) and practical How-to guides.

How to — tell a real benchmark from a marketing one

The Frontier

How to — tell a real benchmark from a marketing one

Every model launch ships a benchmark table. Almost every one of them is, in some sense, true. None of them are telling you the same thing. The job is not to find a fake benchmark — most aren't fake — but to find the one whose number you can carry into a decision. Here is the routine. Benchmarks are the scoreboard the field runs on, and the scoreboard is under strain. Terminal-Bench 4.0 spent recent releases removing saturated tasks and fixing broken ones; SWE-bench had to be split into a "Verif

The Take — Uber's war on robotaxis is a toll booth, not a conscience

The Arena

The Take — Uber's war on robotaxis is a toll booth, not a conscience

Uber is not trying to stop robotaxis. It is trying to make sure that when they arrive, they arrive on Uber's platform and pay Uber a cut. Every piece of the company's new labor politics — the union alliances, the workforce-transition language, the 85 percent rule its lobbyists are pushing in New Jersey — is aimed at that outcome, not at slowing the technology down. I think the distinction matters more than it sounds, because it changes who we should root for. If this were a company making commo

The Take — SoftBank's $5.5B warrants prove AI infrastructure is funding itself

The Arena

The Take — SoftBank's $5.5B warrants prove AI infrastructure is funding itself

SoftBank just handed OpenAI $5.5 billion in stock warrants to stay in a data center lease, and the most important word in that sentence is "just." This is not a customer buying compute. This is a landlord paying its tenant to remain a tenant, then preparing to sell the resulting revenue stream to public market investors as if it were organic demand. The AI infrastructure boom is real. The financing underneath it is more circular than anyone with an IPO to file wants to admit. The facts, as we c

How to — read a model launch without getting spun

The Frontier

How to — read a model launch without getting spun

The first 90 minutes after a frontier model drops, every claim in the launch post is competing for your attention with another claim that contradicts it. The way to read a model launch is not to chase the headline number; it is to walk through the same five things the people who actually use these models look at, in the same order, every time. Here is the routine. A "model launch" in 2026 is rarely one model. It is a small stack of claims, each with a different shelf life: a new checkpoint, a n

The Take — The Anthropic–Pentagon ruling isn't a win for one lab. It's a red line for all of them

The Guardrails

The Take — The Anthropic–Pentagon ruling isn't a win for one lab. It's a red line for all of them

A federal judge just told the US government it cannot punish an AI company for refusing to let its model kill people without a human in the loop. That sentence is bigger than Anthropic, bigger than the $200 million contract at the center of the fight, and bigger than Pete Hegseth's bruised ego. It is the first time a court has converted an AI lab's stated safety principle into something the Constitution actually defends — and the next two quarters will reveal whether OpenAI and Google treat that

How to — know if RAG is actually doing anything

The Stack

How to — know if RAG is actually doing anything

Judging a RAG system by reading its answers tells you almost nothing. That is the trap: a fluent, confident, well-cited answer can come from a model that never retrieved a single document — and a genuinely grounded one can still fail because the right text never surfaced. The way to know if retrieval is actually doing anything is to stop grading final outputs and start asking two separate questions. Here is how to do that without building a research lab. RAG stands for retrieval-augmented gener

The Take — Unitree's rout isn't a bubble. It's the brain lagging the body

The Arena

The Take — Unitree's rout isn't a bubble. It's the brain lagging the body

Unitree did not lose half its value because investors suddenly hated robots. They stopped paying for a general-purpose humanoid whose intelligence, by the builders' own admission, is still stuck in its "GPT-2 era." The rout is the market finally pricing the one thing the last funding cycle chose to ignore: the gap between a robot body that can walk and a robot brain that can work. I think we should stop calling this a bubble popping and call it what it is — a repricing from promise to proof. An

The Take — Lethal AI just crossed a line no one was watching

The Guardrails

The Take — Lethal AI just crossed a line no one was watching

If Ukrainian accounts hold, an AI-guided Russian drone killed three people in Zaporizhzhia on its own — no human in the loop to pick the target — as our morning brief laid out this morning. Months of institutional talk about "meaningful human control," about drawing the line before machines decide to kill, all dissolve into a single unverified strike that nobody could stop, reported like a war statistic. My take: this is not the moment autonomous warfare began. It is the moment we ran out of roo

How to — run a local LLM

The Stack

How to — run a local LLM

Running a large language model on your own machine means your prompts don't go to a cloud provider, there's no per-token bill, and the assistant still answers when the network drops. Here's how to go from zero to a working local model in six moves, even if you've never done it before. 1. Size the model to the memory you actually have The single decision that decides the whole project is memory — the model's files have to sit in RAM (or your graphics card's memory) while it runs, alongside eve

The Take — Sutton is right that frozen AI has a ceiling

The Frontier

The Take — Sutton is right that frozen AI has a ceiling

Richard Sutton, the reinforcement-learning pioneer who wrote "The Bitter Lesson," is one of the few people in AI whose contrarianism is worth taking seriously. But his most cited claim this week isn't his strongest. He is wrong to call synthetic data the industry's biggest mistake; he is right — and the field can't answer it — that an AI which stops learning the moment it ships has a hard ceiling on what it can ever become. Sutton, a 2024 Turing Award winner speaking on Sequoia Capital's Traini

The Take — The data center backlash is the new NIMBY tax on AI

The Everyday

The Take — The data center backlash is the new NIMBY tax on AI

In just twelve months, public opposition to AI data centers in the United States went from a coin-flip to a near-veto. Three out of four Americans now say they oppose one being built near them, up from 42 percent a year ago, and 61 percent are "strongly opposed." This is not a PR problem the industry can outspend or litigate away. It is a structural constraint that will shape where AI infrastructure gets built and how fast — and almost everyone building that infrastructure is still treating it l

The Take — Why Anthropic's revenue flip means the safety premium is real

The Arena

The Take — Why Anthropic's revenue flip means the safety premium is real

Anthropic just passed OpenAI on quarterly revenue for the first time. I think this is the strongest evidence yet that enterprise customers are willing to pay a premium for AI they trust — and that the "safety-first" label, long dismissed as a marketing angle, is now a business strategy with teeth. The numbers that tell the story Anthropic generated more than $11.5 billion in revenue in Q2 2026, more than doubling the $4.73 billion it posted in Q1 and growing 14-fold year over year from $787 m