xAI's Grok 4.6 matches GPT-5.6 Sol on intelligence index

Share
xAI's Grok 4.6 matches GPT-5.6 Sol on intelligence index

Two frontier releases landed within hours of each other: xAI shipped Grok 4.6 with a claim of parity with OpenAI's GPT-5.6 Sol on a composite intelligence index, and DeepSeek pushed its flagship V4 Pro to general availability. A 3-billion-parameter vision model built to run on a phone rounds out the afternoon.

xAI released Grok 4.6, its most capable model yet, and says it matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index — a composite of nine benchmarks where both score 61, behind only Anthropic's Fable 5 Max at 62. Arriving just over a month after Grok 4.5, the new model is tuned for long-running agents and ambitious interactive and visual work: researching topics, working across a codebase, or taking a product idea to a working first version. xAI says a longer supplemental training run with curated model-generated data and a reworked optimizer produced the jump, and that Grok 4.5 was used to regenerate the supervised fine-tuning trajectories.

The benchmark picture is more mixed than the headline suggests. Grok 4.6 leads the field on GDPVal-AA (1753 vs. GPT-5.6 Sol's 1728) and posts a strong CursorBench 3.2 score of 69.9%, but trails GPT-5.6 Sol on DeepSWE (65.9% vs. 73%) and Terminal-Bench (26% vs. 34.6%). The commercial move is the aggressive one: $2 per million input tokens and $6 per million output, with a fast variant at double the price, and 2x included usage in Cursor and Grok Build for the first week. It is live today in Cursor, Grok Build, the API, and through partners like OpenRouter, Vercel, and Cloudflare.


DeepSeek pushed its flagship V4 Pro to general availability as DeepSeek-V4-Pro-0813, pairing a 1-million-token context window with a full agent-developer toolkit. The release ships with structured JSON output, tool calls, the Responses API, Anthropic-API compatibility, and beta support for conversation-prefix continuation and fill-in-the-middle completion — with OpenAI- and Anthropic-compatible endpoints on the wire, so teams can switch without re-plumbing. Pricing formalizes a deliberate two-track lineup: Pro runs 3 yuan per million cache-miss input tokens and 6 yuan per million output (roughly three times Flash), with concurrency capped at 500 requests versus Flash's 2,500. The message is clear — Flash for high-frequency scale, Pro for heavy reasoning, long-context coding, and agent workloads. We covered the independent verification of the Flash tier's agent scores last week — Independent run confirms DeepSeek V4 Flash's 82.7% score.


Liquid AI released LFM2.5-VL-3B, a 3.1-billion-parameter vision-language model designed to run entirely on-device. The multimodal variant of LFM2.5 pairs a 2.6B language backbone with a SigLIP2 NaFlex vision encoder and posts genuinely useful edge numbers: 228 tokens per second on an Apple M5 Max and 116 on an AMD Ryzen AI Max+ 395, all inside 3.3 GB of memory, with a phone (Galaxy S26 Ultra) managing 20 tokens per second. Liquid says the model substantially improves grounding and full-page OCR with layout annotation over its predecessor, and it supports tool use — aimed at single-turn, high-throughput jobs like scanning documents or near-real-time object detection rather than long-context reasoning. It lands with GGUF, ONNX, and MLX exports, so it runs on llama.cpp, vLLM, and Apple Silicon frameworks today.

What to watch: whether the pricing pressure Grok 4.6 and DeepSeek V4 Pro are applying to the frontier's top tier starts showing up in OpenAI and Anthropic's API prices.

Grok matching GPT-5.6 Sol on the composite index at a fraction of the price — does the benchmark race still mean anything? Tell us in the comments.

Sources: SpaceXAI · Techmeme · TradingKey · Pandaily · r/LocalLLaMA (DeepSeek V4 Pro 0813) · Hugging Face · Liquid AI · r/LocalLLaMA (LFM2.5-VL-3B) · AI Midday (DeepSeek V4 Flash coverage)