[ DGX SPARK · THE DEV BOX · THE DEPLOY ]

NVIDIA gave you a $3K desk box.
We give you the AI factory.

The DGX Spark is the most exciting dev box of the year. It's also where SmartTec starts. Same GB10 architecture. Production scale.

NVIDIA DGX Spark desktop supercomputer on a desk
[ NVIDIA ]
DGX Spark · GB10
PRICE
$3,000
one-time
FORM FACTOR
Desktop
240 W
MEMORY
128 GB
unified

Best-in-class for prototyping. 200B params at your desk, 405B with two linked over ConnectX-7. But it's one developer, one model, one power outlet.

NVIDIA GB200 NVL72 rack-scale system at SmartTec
[ SMARTTEC AI FACTORY ]
GB200 · CS-3 · H200
PRICE
$2.40+
/GPU-hr
FORM FACTOR
Rack → farm
120 kW/rack
MEMORY
13.5 TB
/rack HBM3e

Same NVIDIA silicon, deployed behind megawatt-scale batteries we built. Hundreds of concurrent users. Sub-10ms failover. SLA-backed. Token-billed or reserved.

[ SIDE-BY-SIDE ]

Same silicon. Different scope.

Both run NVIDIA. Both run Grace Blackwell. Only one is an AI factory.

Dimension
DGX Spark
SmartTec
Scale
1 chip · 1 PFLOPS FP4 · 128 GB unified
1,000+ GPUs · 1.4 exaFLOPS/rack (GB200 NVL72) · 13.5 TB/rack HBM3e
Form factor
Sits on your desk
Sits in our Mead, OK AI factory
Power
Standard wall outlet (240W)
Megawatt-scale LFP batteries + grid fallback, no wait
Max model size
Up to 200B params solo · 405B with two linked
Any model. Llama 405B, DeepSeek V3, your custom — all fit
Concurrent users
1 developer · 1 model
Production fleet · hundreds of concurrent requests
Uptime
Until your power goes out, or you restart
99.9% SLA · battery-backed · sub-10ms failover
Cost
$3,000–$4,000 one-time + electricity
From $2.40/GPU-hr or $48K/mo flat for a CS-3
[ TOKENS PER WATT ]

The metric Jensen just made the headline.

At GTC 2026, NVIDIA said “tokens per watt” is the new efficiency benchmark. We agree. SmartTec scores higher because our power comes from a battery that arbitrage-charges off-peak — we don't run on the gas peakers your cluster hits during a hot summer afternoon.

[ GRID-POWERED ]
0.62
TOKENS / WATT · 24×7 MIXED LOAD

Typical hyperscaler running on retail grid. Hits dirty peakers during peak. Average emissions intensity ~0.4 kg CO₂/kWh.

[ SMARTTEC ]
0.81
TOKENS / WATT · +30% VS GRID

Behind-the-meter BESS charges off-peak (cheapest, cleanest hours), discharges during compute bursts. Same silicon, 30% more tokens per joule.

[ WHY THIS MATTERS ]

At 1 GW of sustained AI load, the difference between 0.62 and 0.81 tokens/watt is roughly $140M/year in power cost and ~600,000 tons of CO₂— for the same inference output. We're not chasing the metric; we're already winning it.

[ THE WORKFLOW ]

Build on the desk. Deploy on the floor.

This is how the next generation of AI teams ship. Prototype locally. Production at scale. Same models, same toolchain, zero rewrite.

[ 01 · DEVELOP ]

On your DGX Spark

Run a 70B or 200B model at your desk. Fine-tune. Eval. Iterate in seconds. vLLM and the NVIDIA stack work identically to SmartTec — same commands, same containers.

[ 02 · FIT-CHECK ]

On SmartTec /inference

Plug in your model + target throughput. We show VRAM fit, max concurrent sequences, tokens/sec, and a copy-pasteable vLLM deploy command. 30 seconds.

[ 03 · SHIP ]

On SmartTec compute

Same model, same command. Now running on a GB200 NVL72 rack behind 25 kW of LFP batteries per CS-3. SLA-backed. Sub-10ms failover. Priced per token.

[ READY TO SHIP ]

From prototype to production in one deploy.

Run the fit-check on your model. We'll show you the exact vLLM command, the rack you'd land on, and the per-token cost. Q4 2026 power-on for design partners.