2026-06-24 · OpenAI

OpenAI and Broadcom unveil LLM-optimized inference chip

researchinfrastructure

read at source ↗ openai.com

OpenAI and Broadcom unveil LLM-optimized inference chip

Source: OpenAI Date: 2026-06-24 URL: https://openai.com/index/openai-broadcom-jalapeno-inference-chip

Summary

OpenAI and Broadcom announced a custom LLM inference chip codenamed “Jalapeno,” designed specifically for transformer inference workloads rather than general-purpose GPU compute. The chip is built around the memory-bandwidth and interconnect requirements that dominate inference at scale — the bottleneck profile is fundamentally different from training, where raw FLOP throughput is the primary constraint. The announcement positions this as OpenAI’s step toward owning its inference silicon stack, reducing dependence on NVIDIA for production serving. (The announcement page returned 403 to direct fetch; details below are drawn from the announcement title and contemporaneous coverage.)

Implications

  • Inference hardware economics. A purpose-built inference ASIC from a hyperscaler-backed partnership is a direct challenge to NVIDIA’s H100/H200 inference business. Broadcom has prior experience building custom AI accelerators for Google (TPUs were co-designed with Broadcom’s supply chain); “Jalapeno” follows the same playbook of vertically integrating inference cost into the operator’s own stack rather than paying GPU spot prices.
  • Closed-frontier infrastructure. This is the infrastructure-layer expression of the same dynamic visible in OpenAI’s model strategy: lock in capability advantages not just through weights but through the stack underneath them. A chip optimized specifically for their model architectures creates a cost moat that external API competitors can’t replicate without equivalent silicon access.
  • NVIDIA exposure. Any inference-optimized ASIC that ships at production scale compresses the total addressable market for NVIDIA’s inference-oriented products. Watch whether other frontier labs (Anthropic, Google, Meta) accelerate their own silicon programs in response — Google’s TPU lineage already hedges this, but the OpenAI/Broadcom move raises the stakes for labs without custom silicon.
  • Watch: announced production timeline, wafer allocation details, and whether “Jalapeno” feeds only ChatGPT/API inference or becomes available to third parties via Azure.

← all signals