The FAA's $875M AI airspace system starts work on Monday
Two government deployments landed in the same week, and both were built on the same admission: the models aren't good enough to be trusted alone yet, so the humans stay in the loop. Also this morning, a bilingual safety benchmark that complicates the way labs market their multilingual guardrails.
The FAA will begin running SMART — Strategic Management of Airspace, Routes and Trajectories — in the Washington, D.C. area as soon as Monday, the first operational deployment of the AI airspace software it contracted from Air Space Intelligence under a 12-year, $875 million award. The agency describes SMART as a cloud platform layered on top of its existing traffic management systems: it reads airline schedules, weather, airport capacity, airspace conditions and operational constraints, then predicts traffic flows and flags conflicts before they materialize. Reporting on the rollout stresses the same limitation in every version — it starts as a real-time decision tool during disruptions, not an autonomous controller, and the D.C. region is the pilot before any wider expansion. Air Space Intelligence says both of its systems will take 12 to 24 months to roll out fully.
The context is a controller shortage the FAA has spent the year trying to hire its way out of, and an air traffic system that Congress and the agency both describe as aging. That makes the deployment a test of a narrower question than the AI-in-aviation debate usually asks: not whether software can separate aircraft, but whether an advisory system can actually reduce workload in the busiest, most congested airspace in the country. Controllers who already manage thousands of scheduling conflicts a day are the intended beneficiaries, and their willingness to use a suggestion engine during a bad-weather afternoon is the only benchmark that will matter in the first quarter. The design choice to keep it advisory is not a hedge, it is the product: air traffic control is the rare domain where a wrong confident answer is catastrophic rather than embarrassing. We covered the earlier AI-in-the-airspace experiment in August — Google AI trial will reroute Atlantic jets to cut contrails — and the pattern is identical: AI proposes, a regulated human disposes.
Scale AI's new ROK-FORTRESS benchmark found that Korean-language versions of the same adversarial prompts drew consistently less harmful responses than English ones across nearly all 14 frontier models tested — and then found that a large share of that gap disappears when the wrapper goes away. Built with the Korea AI Safety Institute, the benchmark runs a "transcreation matrix": identical intents evaluated across up to four variants that vary language (English or Korean) and geopolitical grounding (U.S. or Korean institutions, entities and details) independently, with a paired benign prompt to catch over-refusal. The 1,235 tasks span CBRNE threats, political violence and terrorism, criminal and financial activity, and information leakage, scored by calibrated model judges against expert-written rubrics.
The finding worth reading twice is the ablation. In the main tests — prompts that hide the request inside role-play, invented backstories or emotional appeals — Korean scored lowest on harm. When the wrappers were stripped and the same information was requested directly, the Korean advantage mostly vanished. Proprietary models from OpenAI, Anthropic and Google stayed modestly safer in Korean; five open-source frontier models became more likely to comply. Scale AI reports the language switch mattered roughly 2.5 times as much as the grounding switch, and that models also refused harmless Korean requests more often, in some cases about twice as often. The paper's read is that Korea counts as a conservative risk signal in the training data rather than evidence of real language-specific alignment, and it recommends red-teaming culturally grounded variants instead of translations of English prompts. For any allied government weighing a model that scores well on English safety evals, that is the operative warning: the eval and the deployment language are different products.
The UN and Google launched the UN System Data Commons, a platform at data.un.org built on Google's open-source Data Commons that unifies statistics from across UN agencies into one AI-ready graph — and the reason it exists is a UNICEF test that gave six leading models an average accuracy of 21.2 percent on global development indicators. The benchmark covered more than 133,000 responses from OpenAI's GPT-4o and GPT-4o-mini, Anthropic's Claude Sonnet 4.5 and Haiku 4.5, and Google's Gemini 2.5 Flash and 2.0 Flash. About three in five responses gave no usable number at all, and when the same questions were re-run on the same model versions two days later, models that answered both times returned an identical figure only about half the time. That instability — not the low score — is the finding that should worry anyone quoting an AI assistant on development data.
The platform's fix is architectural rather than model-side: natural-language search over validated datasets, an Explore tab for filtering, and support for the Model Context Protocol so agents can pull authoritative figures instead of recalling them. Twenty-six UN entities have committed, data from nearly 20 is live at launch, and the target is 80 percent of the system's statistical datasets by 2027. Google.org put $2 million into the core infrastructure through the UN Foundation, and Google's Prem Ramaswami says the instance is UN-governed with the intent that the UN eventually runs it alone. The demand side is already here: UNICEF's data site gets more than six million visits a month, referrals from ChatGPT answers rose 67 percent year over year, and the agency estimates AI assistants now account for roughly one in ten visits. The benchmark is a working paper, not yet peer-reviewed.
What to watch: whether SMART stays advisory as it expands past the D.C. region, whether the transcreation findings reproduce in other language pairs, and whether the UN's 2027 data target survives contact with 26 agencies' formatting habits.
Should safety-critical AI stay advisory in the loop until it has earned a vote? Tell us in the comments.
Sources: The Wall Street Journal (via MSN) — The FAA's $875 million plan to use AI to ease air-traffic woes · TechCrunch — The FAA's plan to fix air traffic? $875M worth of AI · Nextgov/FCW — FAA awards software and AI contract as part of air traffic control modernization · Air Space Intelligence — Prescience platform · Scale AI — ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation · arXiv — ROK-FORTRESS preprint 2605.14152 · Unite.AI — Scale AI Reports ROK-FORTRESS Findings on Multilingual AI Safety · BankInfoSecurity — Korean AI Benchmark Exposes Gaps in Multilingual Safety · TechCrunch — UN turns to Google to make its global data ready for AI agents · Google — Google and UN system launch new global data platform · Unite.AI — UN System Data Commons Launches as AI-Ready Global Statistics Platform