[ REFERENCE · AMD INSTINCT · JULY 2026 ]

There is a
second rack now.

For three years, buying AI infrastructure meant buying NVIDIA and the only question was allocation. AMD Instinct has closed enough of the gap that the question is now a real one. Here is the reference sheet — measured results where they exist, vendor claims labelled as vendor claims.

288 GB

HBM3E per MI355X today — 50% more memory per GPU than the B200's 192 GB

97%

MI355X throughput versus B200 on Llama 2 70B, MLPerf Inference 6.0 — measured, not vendor peak

111–115%

MI355X versus B200 on GPT-OSS-120B offline and server, same MLPerf round

31 TB

HBM4 in one Helios rack — 72× MI455X at 432 GB each, volume 2H 2026

[ SHIPPING TODAY · PER ACCELERATOR ]
SpecificationAMD Instinct MI355XNVIDIA B200 SXM
ArchitectureCDNA 4Blackwell
Memory[✓]288 GB HBM3E192 GB HBM3e
Memory bandwidth8 TB/s8 TB/s
FP16 / BF16 tensor (peak)5.0 PFLOPS~4.5 PFLOPS
FP8 tensor (peak)10.1 PFLOPS~9 PFLOPS
FP4 / MXFP4 (peak)10.1 PFLOPS~9 PFLOPS
FP64 vector/matrix[✓]78.6 TFLOPS~40 TFLOPS
Compute units256 CUs
Scale-up fabricInfinity Fabric[✓]NVLink 5 · 1.8 TB/s
Serving stackROCm · vLLM, SGLang, PyTorch, HF[✓]CUDA · TensorRT-LLM, vLLM, SGLang

Peak figures as published by each vendor. AMD MI350-series numbers from AMD's own product page; B200 figures are the commonly published SXM specifications. Peak FLOPS across vendors are not directly comparable — see the sparsity note below. Compiled July 25, 2026.

[ WHAT WAS ACTUALLY MEASURED · MLPERF INFERENCE 6.0 ]

Spec sheets are written by marketing departments. MLPerf submissions are run on identical models and datasets under audited rules, which makes them the only apples-to-apples comparison available. AMD submitted on MI355X; the relative figures below are against NVIDIA submissions in the same round.

BenchmarkMI355X resultRelative
Llama 2 70B — single node, FP4100,282 tokens/sec97% of B200 · 93% of B300
Llama 2 70B — 11 nodes, 87 GPUs, offline1,042,110 tokens/sec3.1× the prior MI325X generation
GPT-OSS-120B — 12 nodes, 94 GPUs, offline1,031,070 tokens/sec111% of B200
GPT-OSS-120B — 12 nodes, 94 GPUs, server900,054 tokens/sec115% of B200
Wan-2.2-t2v — text-to-video, single streamOfficial submission93% of B200 · 108% after post-deadline tuning

Source: AMD MLPerf Inference 6.0 results blog. Partner submissions in the same round also covered MI300X, MI325X and MI350X.

[ RACK SCALE · 2H 2026 ]
SpecificationAMD Helios · 72× MI455XNVIDIA GB200 NVL72
Accelerators per rack72× Instinct MI455X72× Blackwell + 36 Grace CPUs
HBM per accelerator[✓]432 GB HBM4~186 GB HBM3E
Total HBM per rack[✓]31 TB13.4 TB
Memory bandwidth per GPU[✓]19.6 TB/s~8 TB/s
FP4 compute2.9 EF (convention unstated)1,440 PFLOPS (sparse)
FP8 compute1.4 EF (convention unstated)720 PFLOPS (dense)
Scale-up bandwidth[✓]260 TB/s130 TB/s
Scale-out networking43 TB/s · Pensando Vulcano 800G NICsConnectX / BlueField
Host CPU6th-gen EPYC Venice · up to 256 coresGrace (Arm Neoverse)
AvailabilityVolume deployments 2H 2026[✓]Shipping
Why the FLOPS rows have no winner marked

NVIDIA states GB200 NVL72 as 1,440 PFLOPS FP4 sparse and 720 PFLOPS FP8 dense— the convention is explicit. AMD publishes Helios at 2.9 EF FP4 and 1.4 EF FP8 without saying whether structured sparsity is included. A 2:1 sparse figure is exactly double its dense equivalent, so comparing one vendor's sparse number against another's dense number can be wrong by 100% while looking perfectly precise. We are not going to mark a winner on a row we cannot verify, and neither should anyone's procurement model.

“A credible second supplier changes procurement math everywhere — even for operators who never buy the second one. Single-vendor risk was the industry's quietest line item.”

— SmartTec Engineering
[ WHAT THIS PAGE IS, AND IS NOT ]

This is vendor-neutral reference material, not a product page. SmartTec does not offer AMD Instinct capacity, and none is scheduled. Phase 1A at Mead is 8× HGX B200 — 64 GPUs installed, 60 rentable — behind battery storage, and unchanged by anything on this page. AMD Instinct sits on our open multi-vendor evaluation track for Phase 2/3 — you can already model MI455X fit in our calculator, where it is deliberately fit-checkable but not priced, because no street pricing exists yet. We publish this because buyers deserve to know what the alternatives actually measure, including the ones we do not sell.

[ FAQ ]
How does the AMD Instinct MI355X compare to the NVIDIA B200 for inference?

On measured MLPerf Inference 6.0 results the two are close enough that workload matters more than badge. AMD's MI355X delivered 100,282 tokens per second on single-node Llama 2 70B in FP4, which is 97% of the NVIDIA B200 result and 93% of the B300. On GPT-OSS-120B at multi-node scale the ordering reverses: MI355X reached 1,031,070 tokens per second offline and 900,054 in server mode, or 111% and 115% of B200 respectively. The MI355X also carries 288 GB of HBM3E against the B200's 192 GB — 50% more memory per GPU, which matters when a model must fit in as few accelerators as possible.

What is AMD Helios and how does it compare to NVIDIA GB200 NVL72?

Helios is AMD's rack-scale answer to NVIDIA's NVL72: 72 Instinct MI455X accelerators paired with 6th-generation EPYC Venice CPUs in a single rack, with volume deployments expected in the second half of 2026. AMD publishes 432 GB of HBM4 per GPU for 31 TB per rack and 19.6 TB/s of bandwidth per GPU, against 13.4 TB of HBM3E per rack for GB200 NVL72. Scale-up bandwidth is 260 TB/s on Helios versus 130 TB/s of NVLink on NVL72, and Helios adds 43 TB/s of scale-out through AMD Pensando Vulcano 800G NICs. The compute headline figures — 2.9 exaFLOPS FP4 for Helios against 1,440 PFLOPS sparse FP4 for NVL72 — are not directly comparable, because the two vendors do not state sparsity conventions the same way.

Why can't you compare AMD and NVIDIA peak FLOPS directly?

Because the published numbers use different conventions. NVIDIA labels GB200 NVL72 as 1,440 PFLOPS FP4 sparse and 720 PFLOPS FP8 dense, making the convention explicit. AMD's Helios figures of 2.9 exaFLOPS FP4 and 1.4 exaFLOPS FP8 do not state whether structured sparsity is included. A 2:1 sparse figure is exactly double its dense equivalent, so a reader comparing one vendor's sparse number to another's dense number can be off by 100% without noticing. This is why MLPerf submissions — run on identical models and datasets under audited rules — are the only apples-to-apples comparison available, and why peak FLOPS should be treated as a marketing ceiling rather than a procurement input.

Is ROCm mature enough to serve production inference?

For mainstream inference serving, largely yes. AMD's ROCm stack now supports PyTorch, Hugging Face, vLLM and SGLang, which covers the majority of production LLM serving deployments, and AMD's MLPerf Inference 6.0 submissions were produced on that stack rather than on custom research code. The remaining gap is breadth rather than capability: CUDA has a decade-plus lead in long-tail kernels, profiling tooling, and third-party library support, so bespoke or research workloads are more likely to hit friction on ROCm than a standard vLLM deployment is.

Does more HBM per GPU actually matter for inference?

It changes how many accelerators a model needs, which changes the economics. A model and its KV cache must fit in aggregate GPU memory, so 288 GB per GPU against 192 GB means roughly a third fewer GPUs to hold the same working set, and at the MI455X's 432 GB it is a factor of more than two against a 192 GB B200. Fewer GPUs per replica means less tensor-parallel communication, better utilization, and lower cost per served token — which is why memory capacity per accelerator, not peak FLOPS, is frequently the binding constraint in large-model inference.