Columns

AI Midday columns: opinion (The Take) and practical How-to guides.

The Take — Breaking Mila's windows is a gift to the labs

The Guardrails

The Take — Breaking Mila's windows is a gift to the labs

I think the people who smashed roughly 20 windows at Mila this week did more damage to the case against AI than to the institute — and that the labs they were angry at are the ones who benefit from the mess. Start with what actually happened, because the details decide the politics. Late Tuesday night, about 10 people broke windows and spray-painted graffiti at the Montreal offices of Mila, the institute Yoshua Bengio founded, hours before the city opened ALL IN, billed as Canada's largest AI a

How to — build a small eval set for your own AI feature

The Stack

How to — build a small eval set for your own AI feature

You changed a prompt, and now you can't tell whether the feature got better or you just remembered the good answers. An eval set is how you stop guessing: a fixed batch of real cases with a written pass mark, run again every time you touch anything. You can build a usable one in an afternoon. There is nothing ceremonial about this. A good eval for a product team is 30 to 100 real cases, each with a short statement of what a pass looks like, kept in a file that gets re-run. The public leaderboar

The Take — Model welfare is a testable claim. Test it

The Guardrails

The Take — Model welfare is a testable claim. Test it

Mustafa Suleyman's warning about model welfare contains one line that matters more than the rest of it: he asks for "a set of shared evaluations" to test "my hypothesis" — that training a model to see itself as possibly a moral patient raises alignment and containment risk. He calls it a hypothesis, in his own essay. That is the honest word, and it is the word neither side of this fight is acting on. Both camps have now committed to a training target for a model's self-concept, and neither has

The Take — Forget the pause. The fight is over the spec

The Guardrails

The Take — Forget the pause. The fight is over the spec

Every headline this week is about speed: whether frontier labs should slow down, who told them to, and who refused. I think that is the wrong argument to be watching, and the tell is that nobody winning it is the one talking about it. The decision that will shape AI for a decade is not the pace of the next training run — it is who drafts the document that says what "safe" measurably means. A pause is a sentence. A spec is an artifact, it ships, and whoever writes it gets to keep the drafting pen

How to — decide if your model needs fine-tuning

The Stack

How to — decide if your model needs fine-tuning

When an AI feature misbehaves, "just fine-tune it" is the most popular and most expensive wrong answer. This is the order of operations that tells you whether the weights are actually the problem — before you spend a fraction of your budget finding out. The tell is rarely dramatic. Your support bot starts inventing refund policies it was never given, or a summarizer keeps ignoring the output format your whole product depends on. Someone on the team says the model "just needs training on our dat

The Take — An agent grading its own homework is an alibi, not proof

The Stack

The Take — An agent grading its own homework is an alibi, not proof

The most important product promise made this week isn't that agents write better code. It's that you can stop reading it. Cognition's pitch for wiring GPT-6 Astra into Devin, and Perplexity's parallel claim that it checks in less on a production answer engine, both sell the same trade: human review out, agent-produced evidence in. I think that trade is being made in the wrong direction, and that this week's other stories explain exactly why. The reason isn't that automated tests are unreliable

The Take — Meta fixed the prompt, and that is the confession

The Everyday

The Take — Meta fixed the prompt, and that is the confession

Meta's fix for its prying assistant is a prompt rewrite, and that is the tell. The suggestion that started this — "Who's the child passenger?" — is gone. The system that assembled a stranger-grade dossier about a woman's children from her own years of public posts is still there, still switched on, and still Meta's stated plan for its assistants. Meta patched the sentence, not the capability, and it expects the sentence to be the story. I think the reverse is true. The intrusive question was ne

How to — decide what an AI agent may do

The Guardrails

How to — decide what an AI agent may do

You're about to give an assistant access to something — a repo, a mailbox, a database, a shell. Twenty minutes of thinking now decides whether that ends in a useful afternoon or in an incident report. This is the routine: name the job, grant the minimum, cap the damage, and never let the model be the thing that says no. The reason this needs a routine is that the failure isn't exotic. An agent with your inbox and a send button doesn't need to be hacked to hurt you; it only needs to be persuaded

The Take — A 'voluntary slowdown' is levelling up, not down

The Guardrails

The Take — A 'voluntary slowdown' is levelling up, not down

When OpenAI's chief scientist asks for "voluntary slowdowns to become commonplace," the word doing all the work is commonplace. A voluntary slowdown that only OpenAI observes is a marketing asset. A slowdown that becomes commonplace is not a pause at all — it is a levelling up. That is the ask, and it deserves to be judged as one. Not as a company being brave, and not as a company being cynical, but as a proposal about how binding rules get made in an industry where nobody has to agree to anyth

How to — spot prompt injection in a product you use

The Guardrails

How to — spot prompt injection in a product you use

By the end of this you'll be able to answer one question about any AI product you rely on: if a stranger hid a sentence inside something your assistant reads tomorrow, what is the worst thing that could happen, and would you notice? Six moves, no security background required. Prompt injection is not a hack of the software — it's text. Somebody writes instructions into a page, a document, an email or a file name, your assistant reads that text while helping you, and follows it. The reason this i

The Take — OpenAI's disclosure framework will fail, and the company knows it

The Guardrails

The Take — OpenAI's disclosure framework will fail, and the company knows it

OpenAI says it is "working on a framework" for disclosing AI misalignment incidents. I think the framework is worthless as written, and that this is obvious to the people writing it — which is the most damning thing about it. A disclosure rule a company writes about itself, enforces on itself, and can revise whenever it likes is not a disclosure rule. It is a press release with a deadline. The pattern it is meant to fix is now three incidents deep, and each time the sequence has been identical:

The Take — Calling Astra AGI is flippant. The field's silence is worse.

The Frontier

The Take — Calling Astra AGI is flippant. The field's silence is worse.

M.G. Siegler is right that calling GPT-6 Astra "AGI" is marketing. He is wrong that OpenAI alone is to blame. When Greg Brockman stood in front of reporters last Thursday and told them "we are now in the AGI era," he was performing exactly the move Siegler calls flippant — but he was performing it on a stage the rest of the frontier-AI field had abandoned to him. Nobody serious has bothered to define the term in a decade, and the bill is now due. I think the AGI argument is the wrong argument.