Grok 4.7 ships at 4.6's price — and a hardened safety stack

Share
Grok 4.7 ships at 4.6's price — and a hardened safety stack

xAI's latest flagship landed this morning with the price left untouched, and New York turned its AI safety law into an operational deadline with dates attached.

SpaceXAI released Grok 4.7 today, priced identically to Grok 4.6 at $2 per million input tokens and $6 per million output tokens. The company says the model was built on a larger base and put through longer-horizon reinforcement learning aimed at tasks that take hours rather than minutes, with the gains showing up in self-verification — the model checking its own work — and in managing very long context. It is live through Cursor, Grok Build, the Grok API, and the major cloud platforms and model routers. Pricing is the headline the industry keeps watching: xAI shipped Grok 4.6 in August at the same $2/$6 rate and matched GPT-5.6 Sol on the Artificial Analysis Intelligence Index, and it has now held that line for a second flagship in six weeks.

On xAI's own published benchmarks, the long-task numbers moved more than the headline scores. In the xhigh configuration, Grok 4.7 scores 71.0% on DeepSWE v1.1 — behind GPT-5.6 Sol max at 72.7% and ahead of Fable 5.1 max at 70.0% — alongside 46.3% on CursorBench 4.0, 1,657 on AA Briefcase, and 19.6% on the Harvey legal agent benchmark. The number worth pausing on is Terminal-Bench 4.0: 38.0%, up from 20.3% on Grok 4.6. Terminal-Bench measures multi-hour work inside a shell, which is exactly the regime where agents stop being impressive demos and start being expensive mistakes. Nearly doubling that score while holding the price is xAI's actual competitive move — not the coding score that trails the leader by 1.7 points.

The safety framing is the part that inverts xAI's usual position. SpaceXAI says Grok 4.7 runs a new safety stack that it describes as its strongest yet at refusing improper requests and resisting jailbreaks, and its HackerBench v0.3 result — 3.3% of malicious cyber prompts allowed through, with few false blocks on legitimate defensive work — is the kind of number labs publish when they want enterprise security buyers to read it. xAI has also opened invitation-only red-team access to Grok 4.7 for selected cybersecurity partners. Treat all of it as vendor-reported until external evals land; the model is a day old.


New York has fixed the dates on its AI safety law: developers register with the state from November, and from January 1, 2027 they must file quarterly catastrophic-risk assessments and report critical safety incidents within 72 hours. Governor Kathy Hochul announced the rollout schedule for the RAISE Act and named Marc Gilman as deputy director for the law — the first full-time hire for the Office of Digital Innovation, Governance, Integrity and Trust (DIGIT), the new unit inside the Department of Financial Services that will administer it. Large frontier developers will also have to publish safety and transparency frameworks on their websites, register with DIGIT, file a disclosure statement at least every other year, and pay assessments; the public can file suspected incident reports, and DIGIT will publish an annual summary with recommended changes to the law.

The mechanism is closer to aviation-style mandatory reporting than to a licensing regime — the state wants a paper trail of failures, not veto power over releases. That makes the 72-hour clock the load-bearing clause: an incident reporting duty creates the first real dataset of frontier-model harms that regulators can act on, and it starts accumulating two months from now.

What to watch: whether the January filing deadlines slip the way the model calendars did, and whether other states copy New York's structure — mandatory incident reporting is the piece that travels politically, since it regulates disclosure rather than capability.

Is 72 hours the right clock for reporting a frontier-model safety incident — too slow for serious harms, or too fast to produce anything but lawyer-vetted boilerplate? Tell us in the comments.

Sources: SpaceXAI — Introducing Grok 4.7 · Grok on X — Grok 4.7 is out today · Chaincatcher — SpaceXAI launched Grok 4.7, speed doubled and price halved · BlockTempo — SpaceXAI releases its strongest coding model, Grok 4.7 · WNYT NewsChannel 13 — New York requires AI companies to report safety incidents in 72 hours · Newsday — NYS' new AI oversight office gets its first full-time hire · NYS Focus — Alex Bores on New York's role in regulating AI