▮▮Coloprice
← Glossary

GPU cloud

A GPU cloud is a cloud service that rents GPU compute capacity — for AI training, fine-tuning, inference, and rendering — billed per GPU-hour or via reserved contracts. The market spans hyperscalers (AWS, Azure, Google Cloud) and specialized neocloud providers (CoreWeave, Lambda, Nebius, Crusoe). Pricing varies widely: in 2026, on-demand NVIDIA H100s range from roughly $2-3 per GPU-hour at specialist providers to $4-14 at hyperscalers, down sharply from about $8 peaks in 2023; multi-year reserved clusters price far lower per hour. Modern deployments cluster thousands of GPUs with high-speed interconnects (NVLink within racks, InfiniBand or RoCE Ethernet between them). GPU clouds are among the largest lessees of wholesale colocation and build-to-suit capacity, since each 1,000-GPU cluster needs roughly 1.5-2.5 MW of liquid-cooled infrastructure.