All posts·[ Industry ]

AMD’s Helios moment: what Microsoft’s rack-scale bet means for a 30-GPU operator

Microsoft just became the first hyperscaler to commit to AMD’s Helios racks at scale — 72 MI455X GPUs, 31.1 TB of HBM4, inference-first. Here’s the honest read from an operator whose Phase 1 is contractually NVIDIA.

YJ
Yasir Jahangir
Co-founder & COO
Jul 21, 2026
7 min read

On July 20, Microsoft became the first hyperscaler to publicly commit to deploying AMD’s Helios rack-scale AI platform at scale on Azure. Helios is not a GPU — it is an integrated rack: 72 Instinct MI455X GPUs with 31.1 TB of HBM4 across the system, sixth-generation EPYC Venice CPUs, Pensando networking, and the ROCm software stack, shipping in the second half of 2026. Azure is positioning it inference-first, with new ND MI455X v7 instances, squarely against NVIDIA’s Vera Rubin NVL72. AMD stock moved about 5% on the announcement, which follows its gigawatt-class deals with OpenAI and Meta.

Comparison of NVIDIA rack-scale systems and AMD Helios rack specifications
The unit of competition shifted from chip to rack — and as of this week, there are two racks

Why this matters more than another GPU launch

For three years, buying AI infrastructure meant buying NVIDIA — the only question was allocation. A credible second rack changes procurement math everywhere: hyperscalers gain pricing leverage, supply pressure on NVIDIA parts eases at the margin, and the industry’s single-vendor risk finally gets a release valve. The open question is software: whether ROCm closes enough of the CUDA gap for production inference at frontier scale. Microsoft betting its own Azure services on it is the strongest ROCm endorsement to date — but a forward commitment, not shipped silicon; Helios hardware arrives in H2 2026. Meanwhile the same week brought a reminder of the other bifurcation: Tom’s Hardware reported a gigawatt-scale Chinese data center running entirely on domestic accelerators — zero NVIDIA — which, if it scales, shifts Chinese demand off the parts US buyers compete for.

72×
MI455X GPUs per Helios rack · 31.1 TB HBM4
H2 2026
Helios ship window — forward commitment today
1st
Microsoft: first hyperscaler committed at scale

What a 30-GPU operator actually does with this news

Honestly: nothing rash. Our Phase 1 is 30 NVIDIA B200s behind battery storage in Mead, Oklahoma — contracted with anchor tenants, and unchanged by a rack that ships after our power-on. What changes is Phase 2 planning. We are opening a formal multi-vendor evaluation track, and MI455X-class inference is on it, with three gates before AMD silicon earns a place in Building 2: availability outside Azure through OEM channels at our scale; ROCm maturity in the serving stack our tenants actually run (vLLM-class, day-one model support); and delivered dollars-per-token versus B200 at our power rate — measured, not marketed.

The deeper stack: Venice, UALink, and what it means for Mead

Helios is more than the GPU. The rack pairs MI455X with sixth-generation EPYC “Venice” CPUs — Zen 6 on TSMC’s 2 nm process, the first 2 nm data center silicon shown publicly — and with UALink, the open scale-up interconnect standard backed by AMD, Microsoft, Meta and others as the industry’s alternative to NVIDIA’s proprietary NVLink. If UALink adoption holds, switching costs between vendors fall — which is worth more to a small multi-vendor operator than any single spec. The per-GPU numbers frame the inference case: 432 GB of HBM4 is 2.25× a B200’s memory, meaning large models need fewer GPUs to hold weights at all. The costs are real too: MI455X is a liquid-cooled part with early reports around 2–2.5 kW per GPU — roughly double a B200 — so Helios-class hardware presumes the direct-liquid-cooling and 480 V power work already on our Phase 2/3 roadmap. Our 3 MVA transformer handles Helios-class racks with room to spare; the gating item is the mechanical and electrical buildout, not the utility feed.

2.25×
MI455X memory vs B200 (432 GB vs 192 GB)
2 nm
EPYC Venice — first 2 nm data center CPU shown
~2–2.5 kW
Early-reported MI455X power — liquid cooling required
The small-operator advantage in a two-vendor world

Hyperscalers must bet billions years ahead. A 30-GPU operator can wait for shipped hardware, real benchmarks, and street pricing — then buy whichever rack wins on measured economics. Vendor competition is the one force in this market that structurally favors the small buyer.

[ FAQ ]
What is AMD Helios?

Helios is AMD’s rack-scale AI platform: an integrated system of 72 Instinct MI455X GPUs with 31.1 TB of HBM4 memory across the rack, sixth-generation EPYC Venice CPUs, Pensando networking, and the ROCm software stack, sold as one unit. It ships in the second half of 2026 and competes with NVIDIA’s Vera Rubin NVL72.

Is Microsoft replacing NVIDIA with AMD?

No — it is adding a second supplier. Microsoft committed to deploying Helios at scale on Azure for inference workloads (new ND MI455X v7 instances), alongside its existing NVIDIA fleet. It is the first hyperscaler to make a public at-scale Helios commitment, and it is a forward commitment: hardware ships H2 2026.

Will SmartTec offer AMD GPUs?

Not in Phase 1A, which is 8× HGX B200 — 64 GPUs installed, 60 rentable — plus Cerebras CS-3 capacity at the Mead, Oklahoma site. SmartTec has opened a multi-vendor evaluation track for Phase 2 that includes MI455X-class inference, gated on OEM availability at small-operator scale, ROCm serving-stack maturity, and measured dollars-per-token versus B200.

Does AMD’s rise help GPU buyers?

Structurally, yes. A credible second rack-scale vendor gives every buyer pricing leverage, eases allocation pressure on NVIDIA parts at the margin, and reduces single-vendor risk. Smaller operators benefit most, because they can wait for shipped hardware and measured benchmarks rather than committing capital years ahead.

Do you need certifications or licenses to buy AMD Instinct GPUs?

No. A US company deploying domestically needs no certification to buy AMD Instinct hardware — it ships inside servers from OEMs like Dell, Supermicro, HPE, and Lenovo through a normal business account. Vendors run standard export-control screening, which US domestic buyers pass automatically. The real prerequisites are physical: MI455X-class parts are liquid-cooled at roughly 2–2.5 kW per GPU per early reports, so the facility needs direct liquid cooling and 480 V power distribution before the purchase order matters.

What is UALink?

UALink is an open scale-up interconnect standard — backed by AMD, Microsoft, Meta and others — designed as an industry alternative to NVIDIA’s proprietary NVLink for connecting accelerators inside a rack. Broad UALink adoption would lower the cost of switching between GPU vendors, which particularly benefits multi-vendor operators.