Hardware

AMD Just Took the Gloves Off: 2nm GPU, 34x Faster Inference, and a Direct Shot at Nvidia

CRAZE CRAZE Summary 3 things to know
  • AMD's MI455X GPU on 2nm offers 34x faster inference and 4x more PFLOPS, a generational leap.
  • Helios rack beats Nvidia's Vera Rubin in compute, memory, and tokens per dollar using open standards.
  • OpenAI, Anthropic, Meta, and others committed, but ROCm software still lags behind CUDA's maturity.
Emon Editorial | · 3 min read
AMD Just Took the Gloves Off: 2nm GPU, 34x Faster Inference, and a Direct Shot at Nvidia

On July 23, in San Francisco, AMD CEO Lisa Su made a statement. She unveiled the Instinct MI455X — the first GPU built on TSMC's 2nm process, packing 320 billion transistors and 432GB of HBM4 memory at 23.3 TB/s bandwidth. She also announced Helios, a 72-GPU rack-scale system designed to go head-to-head with Nvidia's Vera Rubin NVL72.

AMD Just Took the Gloves Off: 2nm GPU, 34x Faster Inference, and a Direct Shot at Nvidia
AMD

The headline numbers are aggressive. AMD claims Helios delivers 15% more compute than Vera Rubin, 50% more memory capacity, and up to 30% better tokens per dollar.

The MI455X is not an incremental upgrade. It is a generational leap. Performance: 40 PFLOPS of MXFP4 compute and 20 PFLOPS of MXFP8 compute — both four times higher than the previous generation MI355X. FP32 performance hits 315 TFLOPS, exceeding Nvidia's Rubin in certain matrix operations.

Memory: 432GB of HBM4, a 50% increase over MI355X, with 23.3 TB/s bandwidth. In a 72-GPU Helios rack, that aggregates to 31 TB of unified memory — 50% more than Nvidia's 20.7 TB.

AMD Just Took the Gloves Off: 2nm GPU, 34x Faster Inference, and a Direct Shot at Nvidia
The MI455X is not an incremental upgrade

Architecture: the chip uses a chiplet design with eight compute dies on TSMC 2nm and two base dies on TSMC 3nm. It replaces the Infinity Cache with a shared 192MB L2 cache, expands local data storage per compute unit, and introduces a new "Work Group Processor" architecture that shifts from 64-bit to 32-bit wavefront execution.

Performance claim: AMD says MI455X delivers up to 34x higher token throughput than MI355X in high-concurrency inference scenarios, reducing token cost by up to 18x. These figures are based on pre-production testing and pre-production software.

Helios integrates 72 MI455X GPUs, 18 sixth-generation EPYC "Venice" CPUs, and Pensando networking into a single 225-245 kW liquid-cooled rack.

The networking approach is deliberate. Helios uses UALink over Ethernet for GPU-to-GPU communication and Ultra Ethernet for scale-out — open standards rather than Nvidia's proprietary NVLink. This positions AMD as the "open alternative" in an industry increasingly wary of vendor lock-in.

The customer list is a direct challenge to Nvidia's dominance: OpenAI, Anthropic, Meta, Microsoft, and Oracle have all committed to Helios. OpenAI plans to bring Helios online in Q4 2026, with deployments accelerating through 2027. Anthropic signed a deal to deploy up to 2 GW of AMD MI450 GPUs starting in early 2027, with the first gigawatt coming online in the first half of the year. AMD also committed to invest up to $5 billion in Anthropic equity as part of the deal.

The catch is software. Nvidia's CUDA ecosystem has a 15-20 year head start. ROCm has improved significantly, but "there are still gaps in terms of latest optimizations, and configuration is more complicated."

AMD is addressing this through its partnership with Anthropic — Anthropic said Claude spent a weekend autonomously adapting AMD's entire rack setup — and by building ROCm.AI, a developer platform announced alongside Helios. But the gap remains real.

AMD Just Took the Gloves Off: 2nm GPU, 34x Faster Inference, and a Direct Shot at Nvidia
AMD GPU

The AI chip market is shifting. With 95% market share, Nvidia has been the only game in town. AMD is now offering a credible alternative — not just a chip, but a full system, backed by open standards and major customer commitments.

Lisa Su framed it as a broader trend: "Agentic AI is changing how infrastructure is planned and deployed." AMD estimates that by 2030, the AI accelerator market will reach $1.4 trillion, and the total addressable market for compute will reach $2 trillion.


P.S. If you are an enterprise AI buyer, you now have a real choice — AMD Helios offers more memory, better tokens per dollar, and open standards, but only if you're willing to work with ROCm. The hardware gap is closing; the software gap is next.

Advertisement

CRAZE

Use CRAZE to turn this article into a faster answer: pull the summary, surface the key term, or jump straight to the next story in this thread.

Article