Why we're starting with NVIDIA Cloud Partner architecture
Reference architectures aren't glamorous, but they save months of integration work. Here's why we picked NVIDIA's NCP framework for our base deployment and what it unlocks for customers.
When you're building a new AI cloud, you have a choice: design your own architecture from scratch, or adopt NVIDIA's Cloud Partner (NCP) reference architecture. We chose NCP. Here's why.
What NCP gives you
NVIDIA Cloud Partner is a reference architecture for building AI clouds on NVIDIA hardware. It covers everything from rack-level topology to network fabric to storage to base operating system. It's the same architecture NVIDIA uses for its own DGX Cloud.
The practical benefits:
- →NVIDIA NGC software stack pre-configured. Containers, drivers, NCCL, all tested and working.
- →InfiniBand NDR topology validated. No debugging fabric issues for months.
- →Multi-node training workloads (Megatron, NeMo) work out of the box.
- →Storage integration (NetApp, Pure, VAST) pre-validated.
The tradeoffs
NCP is opinionated. If you want to do something outside the reference architecture — different storage, different scheduler, different OS — you're on your own. We've found this is rarely a problem in practice, but it's worth knowing.
What it means for customers
When you bring a workload to SmartTec, you can run it the same way you'd run it on DGX Cloud, AWS, GCP, Azure, or CoreWeave. Your existing tooling, your existing orchestration, your existing scripts — they all work. Less migration friction, faster time-to-production.