OpenAI publishes first Jalapeño inference chip benchmarks
OpenAI released initial results for Jalapeño, its custom inference silicon: 1.5–1.9× better efficiency and 1.7–3.6× lower latency versus NVIDIA GB200/GB300 in internal tests, running at 700W TDP but staying under 550W in practice. Deployment begins by year-end, with Gen 2 and Gen 3 already in development. Model-assisted kernel optimization via GPT-Astra + Codex added another 1.5–1.8× on top.
Take vendor-published silicon benchmarks with salt, but the direction matters: if OpenAI can serve its own workloads on Jalapeño at claimed efficiency, per-token pricing pressure on frontier APIs accelerates through 2027. Model your inference budgets assuming another 30–50% price cut cycle, not stability.