TLDR: NVIDIA has introduced the Rubin CPX GPU, a specialized processor designed to dramatically accelerate AI inference for massive-context workloads exceeding one million tokens. Unveiled at the AI Infra Summit, this GPU features 30 petaFLOPs of NVFP4 compute and 128 GB of GDDR7 memory, offering a 3x improvement in attention acceleration over previous systems. It integrates into the Vera Rubin NVL144 CPX platform, promising significant performance gains and economic returns for advanced AI applications like software development and generative video.
NVIDIA has announced a significant leap in artificial intelligence infrastructure with the introduction of the Rubin CPX GPU, a purpose-built accelerator designed to tackle the escalating demands of long-context AI workloads. Unveiled at the AI Infra Summit on September 9, 2025, the Rubin CPX is engineered to enhance inference performance and efficiency for systems processing over one million tokens, a critical requirement for modern agentic AI systems engaged in multi-step reasoning, persistent memory, and long-horizon context tasks.
The complexity of AI inference has grown exponentially as models evolve to handle intricate applications such as reasoning over entire codebases in software development, maintaining cross-file dependencies, and generating long-form video content. These workloads introduce unprecedented challenges in compute, memory, and networking, necessitating a fundamental re-evaluation of inference scaling and optimization. The Rubin CPX directly addresses these bottlenecks, particularly in the compute-intensive ‘context phase’ of inference.
Key specifications of the Rubin CPX GPU include 30 petaFLOPs of NVFP4 compute performance and 128 GB of GDDR7 memory. It also features integrated hardware support for video decoding and encoding, streamlining multimedia AI workflows. NVIDIA reports that the Rubin CPX delivers a remarkable three times faster attention acceleration compared to the NVIDIA GB300 NVL72 systems, significantly boosting an AI model’s ability to process extended context sequences without performance degradation. Notably, the Rubin CPX utilizes a single, monolithic die, a departure from the multi-chip module designs characteristic of other GPUs in the upcoming Rubin family.
This new GPU is a cornerstone of NVIDIA’s broader strategy for disaggregated inference, which is optimized through the NVIDIA SMART framework. This approach separates the context and generation phases of inference, allowing for targeted optimization of compute and memory resources. By processing these phases independently, the architecture improves throughput, reduces latency, and enhances overall resource utilization. The Rubin CPX is specifically designed to excel in the context prefill operations, which can span massive datasets like enterprise chatbot sessions with 256,000 tokens or comprehensive code analysis exceeding 100,000 lines.
The Rubin CPX integrates seamlessly into the new NVIDIA Vera Rubin NVL144 CPX platform. This integrated NVIDIA MGX system combines 144 Rubin CPX GPUs, 144 Rubin GPUs, and 36 Vera CPUs in a single rack, delivering an astounding 8 exaFLOPs of NVFP4 AI compute power and 100 TB of high-speed memory. This represents a 7.5x increase in AI performance compared to NVIDIA GB300 NVL72 systems. NVIDIA projects that this platform can generate substantial economic returns, with a potential for $5 billion in token revenue for every $100 million invested.
Jensen Huang, founder and CEO of NVIDIA, stated, “The Vera Rubin platform will mark another leap in the frontier of AI computing — introducing both the next-generation Rubin GPU and a new category of processors called CPX.” He further emphasized that just as RTX revolutionized graphics, the Rubin platform aims to redefine AI computing.
Also Read:
- Nebius Group Secures Landmark $17.4 Billion AI Infrastructure Agreement with Microsoft
- Google Unveils EmbeddingGemma: A Powerful On-Device AI Model for Enhanced Privacy and Efficiency
The Rubin CPX will be supported by NVIDIA’s comprehensive AI software stack, including the NVIDIA Dynamo platform for efficient AI inference scaling and compatibility with the latest NVIDIA Nemotron family of multimodal models. AI innovators such as Cursor, Runway, and Magic are already exploring the potential of Rubin CPX to accelerate their applications. NVIDIA anticipates the Rubin CPX to be available at the end of 2026, following the launch of the regular Rubin GPU earlier that year.


