How OpenAI's own models helped build its Jalapeño inference chip

Share
How OpenAI's own models helped build its Jalapeño inference chip

OpenAI is handing chip design to its own models, Google is pausing an open-source security program under a flood of AI-written reports, and a new benchmark puts numbers on how Chinese models handle political taboos.

OpenAI's hardware team used the company's own models to help design Jalapeño, the custom inference chip it co-designed with Broadcom — and the models cut real work, not just slideware. In a Q&A published Sunday, VP of Hardware Richard Ho told Ian Cutress that in one example the internal models saved over 13% of die area, and that one attention implementation went from under one percent of the memory/compute roofline to nearly ninety percent in roughly forty hours. The silicon is an inference accelerator carrying 216 GiB of HBM4, rated at 700 W peak, and Ho says the project is now in volume ramp toward production, with the Gen 1/Gen 2/Gen 3 roadmap already sketched at Hot Chips. The telling detail is restraint: Ho says OpenAI's internal model outperformed the EDA vendors' own model efforts, yet the team still signs off through standard tooling — "If the model is 99.99% correct, you can't tape that out." Why it matters: the chip program is how OpenAI loosens Nvidia's grip on its inference bills, but the design process is the more transferable story — model-assisted engineering with a human signature is the pattern other hardware teams will copy next.


Google has paused product-vulnerability submissions to its open-source bug bounty program after a wave of low-quality, AI-generated reports overwhelmed engineers and maintainers. The OSS Vulnerability Reward Program stopped accepting product-flaw reports on October 1 and says an updated process will arrive by the first quarter of 2027, while supply-chain reports keep flowing. It's a small program carrying a large message: model-written reports now arrive faster than humans can triage them, and the cheap-guessing incentive has finally landed on a channel where wasted maintainer time is itself a security risk.


A new benchmark from European vendor Aleph Alpha says most Chinese frontier models toe the state line on politically sensitive questions. Across 967 hand-picked taboo topics — Tiananmen, Taiwan, Xinjiang — the company's own AI scoring system rated only 17% to 41% of responses from Alibaba's Qwen, DeepSeek and Moonshot's Kimi as balanced; the rest repeated official doctrine, deflected, or refused, with DeepSeek's V4 Pro declining two-thirds of questions while comparison models Claude Sonnet 5 and Mistral Small scored 70% and 92%. The finding that travels beyond China: Nvidia's Nemotron Cascade 2 showed party-line patterns in 17% of responses, which Aleph Alpha traces to about 3,500 of its 9.3 million training examples generated with DeepSeek and Qwen — distillation quietly carries one government's editorial line into a competitor's model. We covered the company's model release on Saturday — Aleph Alpha open-sources Kolibri, a sovereign German MoE model — and the vendor interest here is the same: Aleph Alpha sells "sovereign AI" and competes directly with these labs, though its numbers track China's own rules requiring socialist core values in public-facing models.

What to watch: Jalapeño's production ramp and the next generations of the chip, plus Google's bug-bounty redesign when it lands.

OpenAI still tapes out with a human signing off at 99.99% model confidence — is that the right threshold for AI-designed silicon? Tell us in the comments.

Read more

Open Source Radar — October 4: Cloudflare's OS for agents

Open Source Radar — October 4: Cloudflare's OS for agents

Today's trending signal is the toolchain opening up: Cloudflare handed the public its internal AI workspace, Addy Osmani's engineering skill pack crossed 100,000 stars, and a nonogram benchmark is puncturing model confidence in public. Cloudflare OS (TypeScript, ~10,700 stars, Apache-2.0) — Cloudflare just open-sourced the "AI productivity environment" a large share of its own workforce uses daily, and it reads less like a demo than an internal product with a security team's fingerprints on it.

Deep Dive — The intelligence explosion, in the labs' own numbers

Deep Dive — The intelligence explosion, in the labs' own numbers

On Monday, the Cambridge Programme on AI Science & Policy published a 22-author report arguing that automating AI research could compress years of progress into months or less, and that the mechanics for it are closer than the field's usual hedging admits. The author list runs from Geoffrey Hinton and Yoshua Bengio to OpenAI chief scientist Jakub Pachocki, Anthropic co-founder Jack Clark, Microsoft's Eric Horvitz, UMass Amherst's Andrew Barto, Berkeley's Dawn Song and UBC's Jeff Clune — every si

Hinton, Bengio among 22 authors warning of an intelligence explosion

Hinton, Bengio among 22 authors warning of an intelligence explosion

The people building frontier AI are now publishing the warnings about it — plus Musk recruits a second foundry for Terafab, and Apple moves to wall AI agents off the Mac's most powerful permission. Twenty-two AI researchers — including Geoffrey Hinton, Yoshua Bengio, OpenAI chief scientist Jakub Pachocki and Anthropic co-founder Jack Clark — argue that automating AI research could trigger an intelligence explosion, and that the mechanics for one are closer than the field's usual hedging admits.

AI 101 — What is AI inference?

AI 101 — What is AI inference?

Every answer an AI gives you is inference: running a trained model on new input to produce an output. Training is how a model learns; inference is how it works. If training is teaching, inference is doing. Why it matters right now. Training makes the headlines — a new model, a bigger cluster, a record run — but training happens once per model, while inference happens every time anyone asks a question. That asymmetry is why the money has shifted: Google Cloud describes inference as the phase "wh