Nvidia has revived its Rubin CPX AI inference accelerator after the project appeared to disappear from the company’s roadmap at GTC 2026, according to TF International Securities analyst Ming-Chi Kuo. Production is expected to begin in the first quarter of 2027. The redesigned chip targets prefill, the input-processing stage of inference that generates the KV cache and can account for more than 50% of current inference workloads. Each CPX GPU is expected to deliver performance approaching a standard Rubin GPU, with maximum power consumption of 2,300 watts and 168GB of HBM4 memory. That compares with 288GB of HBM4 in a standard Rubin GPU and 128GB of GDDR7 in the earlier CPX plan. An eight-card CPX compute tray would provide about 1.34TB of combined HBM4 capacity. The new design will use an independent MGX ETL rack, allowing customers to deploy 64, 128, 192 or 256 CPX GPUs. Each 64-GPU module will contain eight eight-GPU compute trays and one switch tray. NVLink will connect GPUs within trays, while Spectrum-6 Ethernet will link trays and modules, with OSFP fiber used between rack modules. Nvidia recommends pairing CPX with Vera Rubin NVL72 systems at a 1:1 ratio: CPX would handle prefill and send the resulting KV cache to Rubin GPUs over Ethernet RDMA for decoding. The revival is complementary to Nvidia’s Groq 3 LPU and LPX products, which focus on low-latency decoding. Nvidia reported second-quarter revenue of $96.22 billion in August, up 106% year over year and above the $92.18 billion market expectation, and said Vera Rubin had entered full mass production with CoreWeave, Nebius, Microsoft Azure, Google Cloud and Oracle Cloud among its partners. Nvidia had not responded to requests for comment as of press time.