Everything in the stack.
NVIDIA and Cerebras compute on z1power megawatt batteries, orchestrated by AURA. Every layer — silicon, power, cooling, AI, observability — designed and shipped together.
Six layers.
One platform.
NVIDIA reference architecture
Reference-architecture NVIDIA deployments: H100 / H200 / B200 / GB200 on InfiniBand NDR, 8 GPUs per node, full fat-tree topology. Bare-metal or hypervisor-bounded.
Cerebras wafer-scale inference
Dedicated CS-3 systems for committed inference workloads. The fastest tokens-per-dollar on earth, on the same fabric as your NVIDIA compute.
AURA orchestration
Predictive 72-hour load forecasting. Coordinates BESS charge cycles, GPU thermal load, and grid events so your workloads run uninterrupted.
z1power-based BESS
Grid-independent power engineered in-house on z1power LFP modules. Sub-10ms failover from grid to battery by design. We control the full integration: storage, inverters, and AURA.
Live observability
Real-time telemetry for every node, GPU, and battery cell. Per-job cost tracking. SOC 2 Type II. Public status page at /status.
Bare-metal isolation
Full hardware access for HPC and security-sensitive workloads. No noisy neighbors. Single-tenant nodes available on every cluster.
For training and dense inference.
H100 / H200 / B200 / GB200 built to NVIDIA reference-architecture practices. 8 GPUs per node, InfiniBand NDR fabric, fat-tree topology. Bare-metal or orchestrated with Kubernetes or Slurm.
Explore NVIDIA compute →For the fastest tokens on earth.
Wafer-scale CS-3 systems for committed inference workloads. Lowest latency inference available, on the same fabric as your NVIDIA compute. Per-token billing on shared endpoints.
Explore Cerebras inference →Plays well with your stack.
Bring your own scheduler, observability, model hub, and secrets. SmartTec doesn't lock you in.