NVIDIA Makes the Financial Case for Gigawatt-Scale AI Infrastructure
Highlighting a $60 million per megawatt price tag for modern AI data centers, NVIDIA argues that long-term returns rely on multi-generation hardware durability, workload flexibility, and massive token-per-watt efficiency gains.

The Economics of Multi-Megawatt AI Infrastructure
With modern AI data centers scaling to megawatt and gigawatt capacities, building out compute capacity has turned into an immense capital undertaking. Each megawatt of factory capacity costs approximately $60 million, pushing operators to justify investments through long-term return on investment (ROI). On Oct. 1, NVIDIA outlined a detailed financial framework for what it terms "AI factories," arguing that commercial returns hinge on three interdependent variables: earning capacity, useful hardware life, and workload fungibility.
According to NVIDIA, hardware profitability cannot rely solely on peak speed or short-term demand spikes. Instead, sustainable infrastructure must maximize output within fixed power envelopes, maintain secondary market value across multi-year depreciation cycles, and run both AI and non-AI workloads via a unified software stack. The company published the framework ahead of founder and CEO Jensen Huang's keynote at GTC Berlin, scheduled for Oct. 21 at 11 a.m. CEST.
Benchmarks, Fleet Longevity, and Enterprise Deployment Data
At the core of earning capacity is power efficiency, as available electricity remains the operational ceiling for hyperscalers and cloud providers. Citing data from SemiAnalysis AgentX, NVIDIA highlighted that its upcoming Vera Rubin NVL72 systems achieve over 30 times higher throughput per megawatt than GB300 NVL72 systems, alongside an up to 45-fold decrease in cost per million tokens when running the DeepSeek V4 Pro model. The firm also pointed to AgentX analysis evaluating GB300 NVL72 performance on the GLM 5.3 model.
Countering concerns that rapid generational improvements quickly render older silicon obsolete, NVIDIA presented data showing data center fleets retaining utility well past their original accounting horizons:
- Extended lifespans: NVIDIA’s A100 GPU, launched in 2020, remains in active commercial service six years later. Cloud provider CoreWeave recently extended customer reservations for hardware first rolled out in 2020 through 2029.
- Secondary values and contracts: Independent evaluations from Barkr estimate the useful economic life of an eight-GPU H100 system at five to six years, while projecting nine to 10 years for the GB300 NVL72 based on equipment resale dynamics. Silicon Data found that six-year-old A100 processors still fetch 25% of their original acquisition price, despite standard five-year depreciation schedules writing them down to zero. Market metrics from Ornn Data indicate five-year rental contracts for A100 systems retain 80% of the pricing seen on short-term one-month leases.
- Depreciation schedules: A September 2026 report by Sprout, titled "The Productive Life of a Data Center GPU," noted that major cloud operators have systematically pushed back their server depreciation timelines. In one historical example, Microsoft operated its NVIDIA V100 clusters for 8.4 years against a six-year accounting book life.
NVIDIA attributes this longevity to the backward-compatibility of its CUDA ecosystem, supported by over 1,000 CUDA-X libraries and 10 million developers. Because newer algorithmic patterns run on existing silicon, operators avoid stranded assets.
Current deployments reflect this broad hardware utility:
- Eli Lilly operates a 1,016-GPU on-premises cluster supporting small-molecule, protein, and genomics modeling alongside internal agentic assistants.
- Pinterest runs post-training and inference for vision-language models across a 14,000-GPU cloud footprint mixing NVIDIA Blackwell, Hopper, and legacy architectures.
- Revolut relies on the cuDF library to process billions of banking records before training and deploying foundation models.
- Runway splits its pipeline across generations, training video world models on Hopper systems and serving live user inference on Blackwell.
- Texas A&M University maintains a 95% to 98% utilization rate across 26 research initiatives using supercomputing clusters for drug discovery and molecular simulations.
- Outside core machine learning, Dassault Systèmes applies GPU acceleration to digital twin simulations for aircraft certification at Wichita State and vehicle engineering at Lucid Motors, while consumer goods giant Unilever cut marketing production costs by 50% by generating product visuals entirely through digital twins.
Power Caps Drive the Hardware Fungibility Push
The financial case outlined by NVIDIA reflects a broader shift across the data center industry, where access to grid power—rather than floor space or server rack availability—dictates operational growth. Under strict megawatt caps, operators must extract more revenue per unit of electricity while insulating themselves against sudden shifts in model architectures.
NVIDIA argues that making tokens cheaper expands aggregate demand rather than dampening it, as reduced inferencing costs unlock new consumer and enterprise applications. By positioning its programmable GPUs and CUDA stack against dedicated custom ASICs, the company is betting that data center builders will prioritize flexible systems that can pivot between training, inference, and general enterprise parallel computing across the better part of a decade.


