OpenAI says Jalapeño chip surpasses Nvidia GB200 and GB300 in inference tests

OpenAI says Jalapeño chip surpasses Nvidia GB200 and GB300 in inference tests

OpenAI’s Broadcom-built inference ASIC posts higher work per watt and lower latency than Nvidia Blackwell systems, with executives pointing to cheaper tokens even as Jim Cramer dismisses the threat.

Summary

OpenAI published InferenceX benchmark results for Jalapeño, its custom AI inference ASIC developed with Broadcom, saying the chip delivered 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency than Nvidia’s GB200 and GB300 Blackwell systems across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, with minimum token-to-token times 2.7 to 4.1 times lower. Hardware vice president Richard Ho said performance per watt could run 1.8x to 4x better than existing chips and that customer token prices should fall as production scales, while CEO Sam Altman simply noted the chip is fast. OpenAI plans small-volume deployment by end-2026 and a larger 2027 ramp, is already advancing second- and third-generation designs, and will keep buying Nvidia and other accelerators for training and inference rather than treating Jalapeño as a full replacement. Jim Cramer remained skeptical that any custom chip will seriously challenge Nvidia’s dominance. Nvidia shares closed at $213.05, up 2.19%, and edged higher in after-hours trading.

Terms & Concepts
  • AI inference: The process of running a trained model to generate outputs, distinct from the training phase that builds the model.
  • ASIC: An application-specific integrated circuit engineered for a narrow set of workloads rather than general-purpose computing.
  • Performance per watt: How much useful AI compute a processor delivers for each unit of electrical power it consumes.