There is a
second rack now.
For three years, buying AI infrastructure meant buying NVIDIA and the only question was allocation. AMD Instinct has closed enough of the gap that the question is now a real one. Here is the reference sheet — measured results where they exist, vendor claims labelled as vendor claims.
HBM3E per MI355X today — 50% more memory per GPU than the B200's 192 GB
MI355X throughput versus B200 on Llama 2 70B, MLPerf Inference 6.0 — measured, not vendor peak
MI355X versus B200 on GPT-OSS-120B offline and server, same MLPerf round
HBM4 in one Helios rack — 72× MI455X at 432 GB each, volume 2H 2026
| Specification | AMD Instinct MI355X | NVIDIA B200 SXM |
|---|---|---|
| Architecture | CDNA 4 | Blackwell |
| Memory | [✓]288 GB HBM3E | 192 GB HBM3e |
| Memory bandwidth | 8 TB/s | 8 TB/s |
| FP16 / BF16 tensor (peak) | 5.0 PFLOPS | ~4.5 PFLOPS |
| FP8 tensor (peak) | 10.1 PFLOPS | ~9 PFLOPS |
| FP4 / MXFP4 (peak) | 10.1 PFLOPS | ~9 PFLOPS |
| FP64 vector/matrix | [✓]78.6 TFLOPS | ~40 TFLOPS |
| Compute units | 256 CUs | — |
| Scale-up fabric | Infinity Fabric | [✓]NVLink 5 · 1.8 TB/s |
| Serving stack | ROCm · vLLM, SGLang, PyTorch, HF | [✓]CUDA · TensorRT-LLM, vLLM, SGLang |
Peak figures as published by each vendor. AMD MI350-series numbers from AMD's own product page; B200 figures are the commonly published SXM specifications. Peak FLOPS across vendors are not directly comparable — see the sparsity note below. Compiled July 25, 2026.
Spec sheets are written by marketing departments. MLPerf submissions are run on identical models and datasets under audited rules, which makes them the only apples-to-apples comparison available. AMD submitted on MI355X; the relative figures below are against NVIDIA submissions in the same round.
| Benchmark | MI355X result | Relative |
|---|---|---|
| Llama 2 70B — single node, FP4 | 100,282 tokens/sec | 97% of B200 · 93% of B300 |
| Llama 2 70B — 11 nodes, 87 GPUs, offline | 1,042,110 tokens/sec | 3.1× the prior MI325X generation |
| GPT-OSS-120B — 12 nodes, 94 GPUs, offline | 1,031,070 tokens/sec | 111% of B200 |
| GPT-OSS-120B — 12 nodes, 94 GPUs, server | 900,054 tokens/sec | 115% of B200 |
| Wan-2.2-t2v — text-to-video, single stream | Official submission | 93% of B200 · 108% after post-deadline tuning |
Source: AMD MLPerf Inference 6.0 results blog. Partner submissions in the same round also covered MI300X, MI325X and MI350X.
| Specification | AMD Helios · 72× MI455X | NVIDIA GB200 NVL72 |
|---|---|---|
| Accelerators per rack | 72× Instinct MI455X | 72× Blackwell + 36 Grace CPUs |
| HBM per accelerator | [✓]432 GB HBM4 | ~186 GB HBM3E |
| Total HBM per rack | [✓]31 TB | 13.4 TB |
| Memory bandwidth per GPU | [✓]19.6 TB/s | ~8 TB/s |
| FP4 compute | 2.9 EF (convention unstated) | 1,440 PFLOPS (sparse) |
| FP8 compute | 1.4 EF (convention unstated) | 720 PFLOPS (dense) |
| Scale-up bandwidth | [✓]260 TB/s | 130 TB/s |
| Scale-out networking | 43 TB/s · Pensando Vulcano 800G NICs | ConnectX / BlueField |
| Host CPU | 6th-gen EPYC Venice · up to 256 cores | Grace (Arm Neoverse) |
| Availability | Volume deployments 2H 2026 | [✓]Shipping |
NVIDIA states GB200 NVL72 as 1,440 PFLOPS FP4 sparse and 720 PFLOPS FP8 dense— the convention is explicit. AMD publishes Helios at 2.9 EF FP4 and 1.4 EF FP8 without saying whether structured sparsity is included. A 2:1 sparse figure is exactly double its dense equivalent, so comparing one vendor's sparse number against another's dense number can be wrong by 100% while looking perfectly precise. We are not going to mark a winner on a row we cannot verify, and neither should anyone's procurement model.
“A credible second supplier changes procurement math everywhere — even for operators who never buy the second one. Single-vendor risk was the industry's quietest line item.”
— SmartTec Engineering
This is vendor-neutral reference material, not a product page. SmartTec does not offer AMD Instinct capacity, and none is scheduled. Phase 1A at Mead is 8× HGX B200 — 64 GPUs installed, 60 rentable — behind battery storage, and unchanged by anything on this page. AMD Instinct sits on our open multi-vendor evaluation track for Phase 2/3 — you can already model MI455X fit in our calculator, where it is deliberately fit-checkable but not priced, because no street pricing exists yet. We publish this because buyers deserve to know what the alternatives actually measure, including the ones we do not sell.
On measured MLPerf Inference 6.0 results the two are close enough that workload matters more than badge. AMD's MI355X delivered 100,282 tokens per second on single-node Llama 2 70B in FP4, which is 97% of the NVIDIA B200 result and 93% of the B300. On GPT-OSS-120B at multi-node scale the ordering reverses: MI355X reached 1,031,070 tokens per second offline and 900,054 in server mode, or 111% and 115% of B200 respectively. The MI355X also carries 288 GB of HBM3E against the B200's 192 GB — 50% more memory per GPU, which matters when a model must fit in as few accelerators as possible.
Helios is AMD's rack-scale answer to NVIDIA's NVL72: 72 Instinct MI455X accelerators paired with 6th-generation EPYC Venice CPUs in a single rack, with volume deployments expected in the second half of 2026. AMD publishes 432 GB of HBM4 per GPU for 31 TB per rack and 19.6 TB/s of bandwidth per GPU, against 13.4 TB of HBM3E per rack for GB200 NVL72. Scale-up bandwidth is 260 TB/s on Helios versus 130 TB/s of NVLink on NVL72, and Helios adds 43 TB/s of scale-out through AMD Pensando Vulcano 800G NICs. The compute headline figures — 2.9 exaFLOPS FP4 for Helios against 1,440 PFLOPS sparse FP4 for NVL72 — are not directly comparable, because the two vendors do not state sparsity conventions the same way.
Because the published numbers use different conventions. NVIDIA labels GB200 NVL72 as 1,440 PFLOPS FP4 sparse and 720 PFLOPS FP8 dense, making the convention explicit. AMD's Helios figures of 2.9 exaFLOPS FP4 and 1.4 exaFLOPS FP8 do not state whether structured sparsity is included. A 2:1 sparse figure is exactly double its dense equivalent, so a reader comparing one vendor's sparse number to another's dense number can be off by 100% without noticing. This is why MLPerf submissions — run on identical models and datasets under audited rules — are the only apples-to-apples comparison available, and why peak FLOPS should be treated as a marketing ceiling rather than a procurement input.
For mainstream inference serving, largely yes. AMD's ROCm stack now supports PyTorch, Hugging Face, vLLM and SGLang, which covers the majority of production LLM serving deployments, and AMD's MLPerf Inference 6.0 submissions were produced on that stack rather than on custom research code. The remaining gap is breadth rather than capability: CUDA has a decade-plus lead in long-tail kernels, profiling tooling, and third-party library support, so bespoke or research workloads are more likely to hit friction on ROCm than a standard vLLM deployment is.
It changes how many accelerators a model needs, which changes the economics. A model and its KV cache must fit in aggregate GPU memory, so 288 GB per GPU against 192 GB means roughly a third fewer GPUs to hold the same working set, and at the MI455X's 432 GB it is a factor of more than two against a 192 GB B200. Fewer GPUs per replica means less tensor-parallel communication, better utilization, and lower cost per served token — which is why memory capacity per accelerator, not peak FLOPS, is frequently the binding constraint in large-model inference.
- AMD — Helios rack-scale solution, official specifications
- AMD — Instinct MI350 series (MI350X / MI355X) specifications
- AMD — MLPerf Inference 6.0 results
- AMD Investor Relations — Advancing AI 2026 announcements
- NVIDIA — GB200 NVL72 rack specifications
All figures from vendor primary sources and published MLPerf results. Last reviewed July 25, 2026.