
OpenAI’s Broadcom-built inference ASIC posts higher work per watt and lower latency than Nvidia Blackwell systems, with executives pointing to cheaper tokens even as Jim Cramer dismisses the threat.
OpenAI published InferenceX benchmark results for Jalapeño, its custom AI inference ASIC developed with Broadcom, saying the chip delivered 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency than Nvidia’s GB200 and GB300 Blackwell systems across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, with minimum token-to-token times 2.7 to 4.1 times lower. Hardware vice president Richard Ho said performance per watt could run 1.8x to 4x better than existing chips and that customer token prices should fall as production scales, while CEO Sam Altman simply noted the chip is fast. OpenAI plans small-volume deployment by end-2026 and a larger 2027 ramp, is already advancing second- and third-generation designs, and will keep buying Nvidia and other accelerators for training and inference rather than treating Jalapeño as a full replacement. Jim Cramer remained skeptical that any custom chip will seriously challenge Nvidia’s dominance. Nvidia shares closed at $213.05, up 2.19%, and edged higher in after-hours trading.