The Frontier
GPT-6 saturated ARC-AGI-3. The next benchmark grades invention
ARC Prize has concluded its flagship exam is no longer hard enough, and the next version will stop asking models to solve puzzles and start asking them to invent things. Also today: a UC Berkeley and Arena study measures how much of a coding agent's bill comes from