Upstage ships Solar Pro 4, an agentic model with 512K context

Share
Upstage ships Solar Pro 4, an agentic model with 512K context

Korean AI lab Upstage rounded out its 2026 lineup this week with Solar Pro 4, a closed-API flagship aimed squarely at document-heavy agent work, roughly a week after the open-weights Solar Open 2. It is the company's clearest bet yet that the next AI battleground is finishing office jobs, not passing chat tests.

Upstage released Solar Pro 4, a proprietary agentic model for multi-document, multi-tool workloads, with a 512K-token context window and up to 128K output tokens. It is the commercial sibling to Solar Open 2, which ships open weights for self-hosting; Solar Pro 4 is API-only, targets jobs that chain tool calls across many steps, and handles English, Korean, and Japanese. Reasoning is on by default with an effort dial for latency, and the API is OpenAI-compatible, so Upstage's migration pitch is that swapping the endpoint and model name is the whole integration. The model is live on Upstage's console, SolarChat, OpenRouter, and Upstage Studio, with dedicated or on-premises deployments available through sales.

Pricing is built for agent workloads that burn many calls per task: $0.30 per million input tokens, $0.06 cached, and $1.20 for output — with a 90% launch discount through September 10 that cuts the OpenRouter price to $0.03 in and $0.12 out. Upstage's benchmark claims center on agentic evals, where the gap over Solar Open 2 is largest: 57.0 vs 43.2 on Terminal-Bench v2.1, 71.0 vs 62.7 on the long-document AA-LCR, 49.2 vs 37.3 on BrowseComp, and 23.0 vs 18.1 on multi-turn banking (τ³-Banking). Knowledge and coding scores are roughly level (GPQA Diamond 89.0). Several of those numbers are vendor-cited or in-house, so independent verification is still pending; the model has already appeared on the Agent Arena leaderboard.

The product story is as deliberate as the benchmark story. Upstage says Solar Pro 4 was trained on finished work via OfficeVerse, a pipeline that synthesizes office tasks across 11 industry domains and grades them pass or fail on the final deliverable, and the launch demo shows the model turning three prompts into an Excel workbook, a review report, and a slide deck in a single session. A companion agent cookbook ships seven work agents, and the model is explicitly designed to label answers as grounded, not-in-document, or mismatch rather than fabricate when the evidence runs out — the "say it can't verify" behavior enterprise buyers keep asking for.

What makes this worth watching isn't just another long-context API. A mid-tier lab is betting its flagship on finish-the-job reliability at aggressive prices, in a market where OpenAI and Anthropic still charge a premium for agentic access. If independent evals hold up, Solar Pro 4 becomes a credible budget option for the office-automation workloads that are quietly the biggest real-world agent market — and another sign Korea's AI push is producing exportable products, not just chip rallies.

What to watch: independent benchmark runs before the September 10 promo ends.

Do you trust vendor benchmarks for agentic models, or do you wait for independent runs? Tell us in the comments.

Sources: Upstage · LLM Stats · OpenRouter · Arena AI Agent Leaderboard