Renting an Nvidia H100 for a training run costs anywhere from $0.57 an hour to $14.90 an hour depending on where you click “deploy” — GetDeploying’s June 2026 cloud GPU index put the average across providers at $3.14 an hour — and that gap is the single most confusing number in cloud computing right now. AWS, Microsoft Azure, and Google Cloud all sell access to the same silicon, yet their sticker prices, discount structures, and even GPU-to-instance ratios differ enough to swing a year-long training budget by six figures. If you’re choosing where to run inference for a production app or a multi-month foundation-model training job in August 2026, the “which cloud” question isn’t about brand loyalty anymore. It’s a spreadsheet problem.
This comparison breaks down what AWS, Azure, and Google Cloud actually charge for H100 and B200 GPU instances, how their committed-use discounts stack up, what the newest AI model pricing looks like on each platform, and where each cloud wins outright. We pulled figures from cloud GPU pricing indexes, hyperscaler rate cards, and FinOps cost-optimization reports published between June and August 2026, cross-checked against each provider’s own published rate cards.
Don't miss new tech stories on Google
Add Tech Insider once in the Google app and our stories appear in your news suggestions.
Why Cloud GPU Pricing Became the Center of the AWS vs Azure vs Google Cloud Debate
Cloud market share used to be the headline metric in every AWS vs Azure vs Google Cloud comparison. AWS still leads with roughly 28-31% of global infrastructure-as-a-service spend, Azure sits around 21-25%, and Google Cloud trails at 12-14%, according to 2026 analyst estimates cited across multiple cloud comparison guides. Those numbers barely moved this year. What changed is the reason buyers pick a cloud in the first place.
AI and GPU compute now decide contracts that used to hinge on storage egress fees or database licensing. Enterprises running large language model workloads, fine-tuning jobs, or inference at scale are comparing H100 and B200 GPU pricing line by line, because the delta between clouds on a single GPU-hour compounds fast across a training cluster running for weeks. A team renting an 8-GPU H100 node for a month pays roughly the same $98,000-plus regardless of provider at list price, but the per-GPU economics diverge wildly once you look past the sticker and into spot pricing, reserved discounts, and normalized per-GPU rates on smaller instance shapes. Buying the hardware outright isn’t necessarily the cheaper path either: CloudZero’s August 2026 retail data puts a single H100 80GB card at roughly $31,000, while Mercatus AI’s June 2026 pricing sheet lists new H100 SXM5 OEM units at $25,000-$30,000 per GPU, PCIe cards at $22,000-$27,000, and a full 8-GPU HGX server at $250,000-$320,000, with GPUpoet clocking the average H100 PCIe street price at $29,025 that same August. Those numbers are a big part of why most teams still rent rather than capitalize the silicon themselves.
The other shift is model access. Azure still holds exclusive-ish early access to OpenAI’s GPT-5.x family through Azure OpenAI Service. AWS Bedrock leans on Anthropic’s Claude models plus its own Titan and Nova lines. Google Cloud pairs Vertex AI with Gemini and its in-house TPU hardware, which sidesteps Nvidia GPU pricing entirely for some workloads. Picking a cloud today means picking a GPU price curve and a model roster at the same time, and the two don’t always agree.
H100 and B200 GPU Pricing: AWS vs Azure vs Google Cloud Specs Table
The table below normalizes on-demand pricing to a per-GPU hourly rate where instances bundle multiple GPUs, since that’s the only way to compare apples to apples across three different instance families.
| Metric | AWS | Microsoft Azure | Google Cloud |
|---|---|---|---|
| H100 instance family | P5 (p5.48xlarge) | ND H100 v5 | A3 (a3-highgpu-8g) |
| H100 on-demand, per GPU | ~$6.88-$12.29/hr | ~$12.25-$14.50/hr | ~$3.00-$11.50/hr |
| 8-GPU H100 node, on-demand | ~$98.32/hr | ~$98.46/hr | ~$98.32/hr |
| H100 spot/preemptible, per GPU | ~$1.95-$2.50/hr | Low-priority VM pricing varies | ~$2.10-$2.80/hr |
| 1-year commitment discount | ~30% | ~38% | Committed use discounts available |
| 3-year commitment discount | ~50% | ~62% | Committed use discounts available |
| B200 availability | P6 instances (rolling out) | ND B200 v6 (rolling out) | A4 (rolling out), median ~$16.11/hr on-demand |
| Native AI chip alternative | Trainium2/Trainium3 | Maia 200 (limited) | TPU v6/v7 (Trillium/Ironwood) |
| Flagship first-party model | Claude via Bedrock | GPT-5.5 via Azure OpenAI | Gemini 3.1 Pro via Vertex AI |
| Global IaaS market share (2026) | ~28-31% | ~21-25% | ~12-14% |
| Kubernetes-managed offering | EKS | AKS | GKE |
| Object storage entry price | S3 Standard, $0.023/GB | Blob Hot, ~$0.018/GB | Cloud Storage Standard, ~$0.020/GB |
| Free tier for compute | 750 hrs t2.micro/mo (12 mo) | $200 credit, 12 mo free tier | $300 credit, always-free tier |
Two things jump out. First, at the 8-GPU node level, list pricing for H100 clusters is nearly identical across all three clouds, sitting within 15 cents of each other at roughly $98.32-$98.46 an hour. Second, once you normalize down to smaller instance shapes or look at Google’s A3 pricing in certain regions, GCP’s per-GPU H100 rate can run as low as $3.00-$3.35 an hour, a fraction of Azure’s $12.25-plus. That’s not a rounding error. It’s the difference between a $2,200 monthly bill and a $9,000 one for the same chip.
AWS Cloud GPU and AI Pricing in Detail
AWS runs H100 GPUs through its P5 instance family, with p5.48xlarge bundling eight H100 80GB GPUs at roughly $98.32 an hour on-demand, or about $6.88 per GPU when normalized. That’s the number you’ll see quoted most often as “AWS H100 pricing,” but it understates cost for teams that don’t need a full 8-GPU node. Smaller P5 configurations and spot capacity push the effective per-GPU rate anywhere from $1.95 to $12.29 depending on availability and instance size, since AWS calculates spot discounts against demand in each Availability Zone rather than a flat percentage.
Where AWS pulls ahead is reserved and committed-use economics at scale. A June 2025 price cut trimmed H100 on-demand rates by roughly From their 2024 peak, AWS EC2 effective rates are not documented as having fallen by a precise **44%**, and current Savings Plans guidance shows 3‑year discounts roughly in the **56–72%** range depending on payment option, rather than being capped at “up to 50%.” For large language model inference specifically, AWS Bedrock’s Claude Sonnet-based stack came out dramatically cheaper than rival flagship-model bills in at least one June 2026 cost analysis: an estimated $72,000 a month (about $864,000 a year) for a workload of 1 billion input tokens and 200 million output tokens monthly, versus roughly $3.25 million a year on Azure OpenAI with GPT-5 or Google Cloud with Gemini 2.5 Pro at list pricing for the same volume. That gap says more about model choice than raw GPU pricing, but it’s the number finance teams actually see on the invoice.
AWS also ships its own Trainium and Inferentia silicon as an alternative to Nvidia GPUs, aimed at teams willing to trade some flexibility for lower training and inference costs. It’s a smaller ecosystem than CUDA, but AWS has been pushing it hard as a hedge against Nvidia’s GPU supply constraints and pricing power. For teams already standardized on AWS for everything else, from RDS-managed databases to Secrets Manager, staying in-ecosystem for GPU compute avoids cross-cloud egress fees that can add up fast on multi-terabyte training datasets.
Microsoft Azure GPU and AI Pricing in Detail
Azure’s ND H100 v5 series is consistently the most expensive H100 option among the three hyperscalers when normalized per GPU, landing between $12.25 and $14.50 an hour on-demand in most 2026 pricing surveys — ComputeTape’s July 2026 index measured Azure’s ND96isr H100 v5 instance specifically at $12.29/hr on-demand, right in that band. At the 8-GPU node level it’s within a rounding error of AWS and GCP ($98.46 an hour), but Azure’s smaller instance shapes and non-committed pricing tend to run higher than either competitor. Azure’s discount structure partly compensates: AWS EC2 Instance Savings Plans typically offer roughly 40-60% off for 1-year terms and 56-72% for 3-year terms, while Azure and Google Cloud each run their own discount ladders, so a single “38% off on-demand” or “62% at three years” figure understates how much those bands vary by instance family and region. Azure also offers low-priority (spot-equivalent) VMs for interruptible workloads, though published rates for those vary more by region than AWS’s or GCP’s spot pricing.
What Azure sells that the others can’t fully match is first-party access to OpenAI’s frontier models through Azure OpenAI Service. GPT-5.5 pricing on Azure runs around $15 per million input tokens and $60 per million output tokens for the flagship tier, with a “Global” GPT-5.5 Pro tier reaching as high as $30 input / $180 output per million tokens in some August 2026 pricing breakdowns, reflecting premium positioning for the top-end model. For enterprises whose AI roadmap is built around GPT-5.x specifically, rather than an open model or Claude, Azure is functionally the only option, which is why Microsoft’s Azure earnings backlog hit $678 billion in 2026, up 84% year over year. That backlog number matters for buyers too: it signals Microsoft has locked in enough enterprise AI commitments that GPU capacity constraints, and by extension pricing pressure, aren’t going away soon.
Azure’s Maia 200 custom AI accelerator exists but remains a limited-availability option rather than a mainstream alternative to Nvidia GPUs, unlike Google’s TPU lineup or AWS’s Trainium chips, which are both available as standard SKUs. That leaves Azure customers more exposed to Nvidia’s own pricing and supply decisions than AWS or Google Cloud customers who have a mature in-house alternative to fall back on.
Google Cloud GPU and AI Pricing in Detail
Google Cloud is the price leader on raw H100 GPU rates in most 2026 comparisons, with on-demand per-GPU pricing on its A3 instance family ranging from roughly $3.00 to $3.50 an hour in some analyses, and up to $9.00-$11.50 an hour in others depending on region and instance configuration. Preemptible (Google’s version of spot) H100 capacity can drop as low as $1.57-$2.80 an hour per GPU. At the full 8-GPU a3-highgpu-8g node level, GCP lands at the same roughly $98.32 an hour as AWS’s p5.48xlarge, so the “Google is cheapest” claim depends heavily on which instance shape and region you’re pricing against.
Google’s real structural advantage is TPUs. Its sixth and seventh-generation Tensor Processing Units (marketed under names like Trillium and Ironwood in various 2026 materials) give Google Cloud customers a path to train and serve models without touching Nvidia GPU pricing at all, something neither AWS’s Trainium nor Azure’s Maia 200 matches in maturity or ecosystem support. That’s a major reason a 2026 “same app, three bills” cost comparison named Google Cloud the outright winner for AI and machine learning workloads, with Azure as runner-up, largely because of TPU economics plus Vertex AI’s tooling. Google’s flagship model pricing supports that lead too: Gemini 3.1 Pro runs around $2 per million input tokens and $12 per million output tokens, and Gemini 2.5 Flash-Lite drops as low as $0.10/$0.40 per million tokens for lightweight tasks, undercutting both Claude Opus and GPT-5.5 on list price.
B200 pricing is where Google Cloud loses its edge. A July 2026 GPU rental index put Google Cloud’s B200 on-demand rate at roughly $16.11 an hour, the highest of any provider tracked, against a market median around $6.25 and specialized marketplace lows near $3.75. If your workload needs the newest Blackwell-generation silicon rather than H100, Google Cloud is currently the expensive option, not the cheap one.
Cloud AI Model Pricing Table: GPT-5.5 vs Claude vs Gemini
GPU rental rates only tell half the story. Most teams don’t manage raw GPU clusters, they call a managed model API and pay per token. Here’s how the three clouds’ flagship-model pricing compared as of August 2026.
| Cloud / Model | Input ($/1M tokens) | Output ($/1M tokens) | Notes |
|---|---|---|---|
| AWS Bedrock: Claude Haiku 4.5 | $1.00 | $5.00 | Budget tier, fast inference |
| AWS Bedrock: Claude Sonnet 4.5 | $3.00 | $15.00 | Mainstream production tier |
| AWS Bedrock: Claude Opus 4.6 | ~$5.00 | ~$25.00 | Flagship reasoning tier |
| Azure OpenAI: GPT-5.2 Global | $1.75 | $14.00 | Mid-tier global deployment |
| Azure OpenAI: GPT-5.5 Global | $5.00 | $30.00 | Flagship general model |
| Azure OpenAI: GPT-5.5 Pro | $30.00 | $180.00 | Top-end reasoning tier |
| Google Vertex AI: Gemini 2.5 Flash-Lite | $0.10 | $0.40 | Cheapest production-grade tier |
| Google Vertex AI: Gemini 2.5 Flash | $0.30 | $2.50 | Balanced cost/performance |
| Google Vertex AI: Gemini 3.1 Pro | ~$2.00 | ~$12.00 | Flagship reasoning tier |
The pattern holds across every published 2026 comparison: Google’s Gemini tiers are the cheapest at every comparable quality level, AWS’s Claude tiers sit in the middle, and Azure’s GPT-5.x lineup is priced at a premium, especially at the “Pro” reasoning tier where output tokens cost $180 per million, roughly seven times Claude Opus’s rate and fifteen times Gemini 3.1 Pro’s. That premium buys access to a model family many enterprises have already built workflows around, so it’s not automatically the wrong choice, but it is the most expensive one on a pure cost basis.
Benchmarks: What Three Independent Cost Analyses Found
We cross-referenced three separate 2026 cost studies rather than relying on one source, since GPU and model pricing shifts month to month and vendors rarely publish head-to-head comparisons themselves.
- nOps FinOps report (August 13, 2026): Normalized per-GPU H100 pricing put AWS at $12.29/hr on-demand with spot rates of $1.95-$2.50/hr and 1-year/3-year commitment discounts of 30%/50%. GCP came in at $9.00-$11.50/hr on-demand with preemptible rates of $2.10-$2.80/hr. Azure ran $12.25-$14.50/hr on-demand with low-priority VM discounts plus Typical commitment tiers discussed in current material are closer to **40–60%** for 1‑year and **56–72%** for 3‑year AWS EC2 Instance Savings Plans, and roughly **mid‑30s to low‑60s** discount bands for Azure Reserved VM Instances depending on payment option.
- Spendark ML cost benchmark (March 2026, updated August 2026): At the full 8-GPU H100 node level, AWS’s p5.48xlarge, Azure’s NCads/ND H100 v5, and GCP’s a3-highgpu-8g all landed within 14 cents of each other, between $98.32 and $98.46 an hour, showing near-total price parity for full-node training clusters at list price.
- Spheron and Thunder Compute GPU indexes (mid-2026): Per-GPU normalized rates showed the widest divergence, with GCP’s A3 series as low as $3.35/hr, AWS’s P5 around $6.88/hr, and Azure’s ND H100 v5 at roughly $12.29/hr, a nearly 4x spread between the cheapest and most expensive hyperscaler for the identical Nvidia H100 chip.
The takeaway across all three: full-node, on-demand H100 pricing is essentially commoditized between the big three clouds. The real savings live in three places, specialized marketplace rentals (Vast.ai listed H100 as low as $1.65/hr in one August 2026 index), long-term commitment discounts, and smaller/fractional instance shapes where GCP consistently undercuts Azure. That volatility isn’t new, either: SemiAnalysis data cited by Compux shows blended H100 rental rates fell roughly 64% from about $8/hr at the chip’s 2023 launch to near $1.70/hr by mid-2025, then rebounded about 40% to around $2.35/hr by March 2026 as demand outpaced new capacity, and kept climbing from there, with Spheron’s index tracking Hyperbolic’s H100 SXM listing up to $3.19/hr by August 26, 2026, a reminder that this year’s cheap spot rate can look expensive again within twelve months.
Storage, Networking, and Egress Costs Beyond GPU Compute
GPU rental rates get all the attention, but for most AI workloads, storage and networking costs quietly add 15-Top‑tier cloud support surcharges are generally **up to around 10% of monthly spend with minimum floors** (for AWS and Google Cloud) rather than a flat 30% on top of the compute bill; AWS Enterprise Support, for example, is tiered starting at 10% of the first usage band with a fixed monthly minimum. Training a large model means staging terabytes of data close to the GPUs, checkpointing model weights repeatedly during long runs, and serving inference from object storage or a vector database. None of that is free, and none of it is priced the same way across AWS, Azure, and Google Cloud.
Object storage entry pricing favors Azure on paper, with Blob Storage’s Hot tier running around $0.018 per GB versus roughly $0.020 on Google Cloud Storage Standard and $0.023 on AWS S3 Standard. That gap narrows or reverses once you factor in egress. AWS charges for data leaving its network to the public internet or to another cloud provider, and those fees can dominate a bill for teams that regularly move checkpoints or datasets between regions or providers. Google Cloud and Azure both have their own egress schedules, and none of the three make cross-cloud data movement cheap, which is a large part of why multi-cloud GPU brokering strategies only make financial sense for teams with genuinely portable, well-compressed datasets.
Network throughput between GPUs matters just as much as raw compute price for large training jobs, since a cluster bottlenecked on interconnect bandwidth wastes expensive GPU-hours waiting on data. AWS’s P5 instances use Elastic Fabric Adapter (EFA) networking, Azure’s ND H100 v5 series relies on NVLink plus InfiniBand, and Google’s A3 instances pair with its custom Jupiter network fabric. All three now support the 3.2 Tbps-class interconnect bandwidth needed for efficient multi-node training on H100 clusters, so this is no longer a meaningful differentiator the way it was a couple of years ago. Where it still matters is B200 and next-generation clusters, where interconnect maturity varies more by provider and region.
Vector database and retrieval infrastructure costs are the newest line item on AI cloud bills, driven by the rise of retrieval-augmented generation pipelines. AWS, Azure, and Google Cloud all now offer managed vector search (OpenSearch Serverless, Azure AI Search, and Vertex AI Vector Search respectively), priced per index size and query volume rather than per GPU-hour, and teams building a RAG pipeline should budget for this separately from compute, since it scales with document corpus size rather than model usage.
Enterprise Support Plans and SLA Comparison
Support tier pricing rarely shows up in GPU cost comparisons, but it’s a real line item for any team running production AI infrastructure, and the three clouds structure it differently enough to matter. AWS’s Business support tier starts at the greater of $100 a month or a percentage of monthly spend (roughly AWS Enterprise Support historically used a $15,000/month minimum with a tiered 10%/7%/5%/3% usage structure, but current 2026 guidance shows minimums as low as **$5,000/month** for some enterprise tiers, still using a similar banded percentage schedule. Azure’s equivalent Professional Direct plan starts around $1,000 a month scaling with usage, while Unified Support (Azure’s enterprise tier) is typically negotiated based on total Microsoft spend across the account. Google Cloud’s Enhanced Support starts near $500 a month or Azure Premium/Unified support is not priced as a flat **3% of monthly spend** across all customers; instead, Azure uses banded and often negotiated enterprise support structures rather than a uniform percentage rate.
SLA commitments for GPU instance availability are broadly comparable across the three: all three providers commit to roughly AWS EC2, Azure VMs, and Google Compute Engine commonly advertise regional or multi‑zone SLAs around **99.95–99.99%** monthly uptime for standard compute, with tiered credits where roughly **99.0–99.99%** availability yields about **10%** credit and worse tiers can yield **25–30% or more**, so a blanket “99.9% monthly uptime” baseline does not reflect the default SLAs of the major providers. None of the three offer meaningfully differentiated SLA terms specifically for GPU capacity versus general compute, which means the practical difference between providers during a capacity crunch, like the kind that pushed on-demand H100 pricing up during 2025’s supply shortage, comes down to how quickly each provider’s account team can source available capacity, not what the SLA document says.
For teams evaluating total cost of ownership rather than just the hourly GPU rate, it’s worth pricing support tiers alongside compute discounts, since a company spending $50,000 a month on GPU infrastructure could pay anywhere from $500 to $15,000 a month in support fees depending on the tier and provider, a swing large enough to change which cloud comes out cheaper overall.
Real-World Use Cases: Which Cloud Wins for What
Pricing tables only matter in context. Here’s how the numbers play out for five common workload types.
Startup Fine-Tuning an Open-Weight Model on a Budget
A seed-stage team fine-tuning a 70B-parameter open model for a few weeks wants the lowest per-GPU rate with no long-term commitment. Google Cloud’s A3 preemptible instances, at roughly $2.10-$2.80 per GPU hour, or third-party marketplaces layered on top of any cloud’s spot capacity, beat Azure’s on-demand ND H100 v5 rate by a wide margin. This is the scenario where shopping around actually pays off, since a 4x per-GPU spread on a six-week training run is the difference between a $15,000 bill and a $50,000 one.
Enterprise Standardized on Microsoft 365 and Azure AD
A company already running its identity, productivity, and internal tooling stack on Microsoft’s ecosystem usually stays on Azure for AI too, even paying the GPT-5.5 premium, because the integration cost of standing up a parallel AWS or GCP AI stack outweighs the per-token savings. Azure’s Current 3‑year commitment discounts for AWS EC2 Instance Savings Plans often reach into the **mid‑60s to low‑70s percent off** range, while Azure 3‑year Reserved VM discounts typically fall in the **mid‑50s to low‑60s percent** range, so a fixed “62% three‑year commitment discount” is not the most accurate single figure.
High-Volume Inference API for a Consumer App
An app serving millions of chatbot or summarization requests a day is extremely token-cost sensitive. Google’s Gemini 2.5 Flash-Lite at $0.10/$0.40 per million tokens is roughly 10x cheaper than AWS’s Claude Haiku 4.5 and dramatically cheaper than any Azure GPT-5.x tier. For workloads where model quality differences don’t matter much at the margin, this is close to a clear-cut decision in Google’s favor.
Large-Scale Foundation Model Pretraining
A lab training a frontier-scale model from scratch needs thousands of GPUs for months, which makes committed-use discounts the dominant factor, not on-demand or spot pricing. Azure’s 62% three-year discount and AWS’s 50% three-year Savings Plan both beat any spot strategy at this scale, since spot capacity for thousands of GPUs simultaneously is rarely available anyway. AWS’s Trainium2/Trainium3 chips and Google’s TPU v6/v7 lines are also worth serious evaluation here, since both offer meaningfully lower cost-per-FLOP than Nvidia GPUs for teams willing to adapt their training code.
Regulated Industry Needing Data Residency and Compliance
A healthcare or financial services team constrained by data residency rules often has less pricing flexibility, since only one or two clouds may have compliant regions or certifications for their specific use case. In this scenario, GPU price becomes a secondary factor behind compliance scope, and all three clouds now offer HIPAA-eligible and SOC 2-audited AI services, so the decision usually comes down to which provider already holds the compliance paperwork for the rest of the company’s infrastructure.
Pricing Comparison Table: Committed-Use Discounts and Entry Costs
| Commitment Type | AWS Discount | Azure Discount | Google Cloud Discount |
|---|---|---|---|
| On-demand (no commitment) | Baseline | Baseline | Baseline |
| Spot / preemptible | Up to ~75% off (H100: $1.95-$2.50/hr) | Low-priority VM, variable by region | Up to ~75% off (H100: $2.10-$2.80/hr) |
| 1-year commitment | ~30% | ~38% | Committed use discount, varies by resource |
| 3-year commitment | ~50% | ~62% | Committed use discount, varies by resource |
| Free trial credit | 750 free-tier hours (12 mo) | $200 credit (30 days) + 12-mo free tier | $300 credit (90 days) |
| Minimum billing increment | Per-second | Per-second | Per-second (1-min minimum) |
Azure’s committed-use discount curve is the steepest of the three, which matters if your organization can accurately forecast GPU demand three years out. Few AI teams can, given how fast model architectures and hardware generations turn over, which is part of why spot and preemptible pricing gets so much attention in 2026 FinOps circles despite the availability tradeoffs.
Migration Guide: Moving a GPU Workload Between Clouds
Switching cloud providers for GPU-heavy AI workloads is more involved than a typical lift-and-shift, mainly because of dataset egress costs, driver/CUDA version drift, and orchestration differences. Here’s the general path teams have followed in 2026 migrations.
- Audit current GPU utilization. Pull 90 days of instance-hour data to separate steady-state (reservation-eligible) load from bursty (spot-eligible) load before pricing out the new cloud.
- Price the target cloud on the same shape. Match instance-for-instance (e.g., AWS p5.48xlarge to Azure ND H100 v5 8-GPU to GCP a3-highgpu-8g) rather than comparing list per-GPU rates in isolation, since node-level bundling changes the effective price.
- Calculate egress cost for the dataset move. Training datasets in the multi-terabyte range can cost thousands of dollars to transfer out of the source cloud; get a quote before committing to a switch.
- Containerize the training/inference stack. Docker images with pinned CUDA, cuDNN, and framework versions reduce the risk of subtle numerical drift when moving between AWS’s, Azure’s, and Google’s slightly different GPU driver stacks.
- Re-provision via managed Kubernetes. Whether that’s AKS, GKE, or EKS, standardizing on Kubernetes manifests up front makes the actual cutover closer to a config change than a rewrite.
- Run a parallel inference shadow test. Route a small percentage of production traffic to the new cloud’s endpoint for 1-2 weeks before fully cutting over, comparing latency and output quality against the baseline.
- Lock in committed-use pricing only after the shadow test passes. Committing to a 1- or 3-year discount before validating performance on the new provider risks paying for a downgrade.
- Decommission the old environment on a delay. Keep the source cloud’s resources live at minimum scale for 30 days post-cutover as a rollback path.
Teams already running vLLM for self-hosted inference have an easier migration path than teams tied to a single cloud’s managed model API, since vLLM’s OpenAI-compatible serving layer can point at any cloud’s GPU instances without an application-layer rewrite.
Pros and Cons: AWS vs Azure vs Google Cloud for GPU and AI Workloads
AWS
- Largest global infrastructure footprint and market share (~28-31%)
- Mid-tier H100 pricing with the widest range of instance shapes
- Claude models via Bedrock offer competitive quality-to-price ratio
- Mature Trainium/Inferentia alternative to Nvidia GPUs
- Deepest existing enterprise footprint (RDS, S3, IAM) reduces integration cost
- Not the cheapest on raw per-GPU rate for smaller instance shapes
- Savings Plans require accurate multi-year forecasting to pay off
- Bedrock model selection narrower than Vertex AI’s catalog
Microsoft Azure
- Exclusive-tier access to OpenAI’s GPT-5.x model family
- Steepest committed-use discounts (up to ~62% at 3 years)
- Best fit for existing Microsoft 365/Entra ID enterprise shops
- $678B earnings backlog signals long-term platform stability
- Most expensive H100 on-demand pricing of the three clouds
- GPT-5.5 Pro tier priced far above comparable Claude/Gemini tiers
- Maia accelerator still limited-availability, less mature than TPU or Trainium
Google Cloud
- Cheapest flagship model pricing across the board (Gemini tiers)
- Lowest normalized H100 per-GPU rate in most 2026 indexes
- TPU v6/v7 lineup is the most mature non-Nvidia AI silicon option
- Named AI/ML category winner in at least one 2026 same-app cost study
- Smallest market share, smaller partner/ISV ecosystem than AWS or Azure
- B200 pricing is the most expensive of the three clouds (~$16.11/hr median)
- Fewer regions than AWS, which matters for data residency requirements
Real-World Examples From 2026 Deployments
Publicly documented cost comparisons and vendor case studies from 2026 illustrate how these pricing gaps play out in practice.
- Mid-size SaaS company (per OPTI Software’s August 2026 analysis): Running an equivalent 1M-input/200K-output token workload across all three clouds, Gemini 2.5 Flash-Lite cost roughly $0.18, versus a few dollars for comparable Claude and GPT-5.x tiers, a gap that becomes six figures a year at production scale.
- Enterprise AI agent deployment (per OpenHelm’s 2026 comparison): A Vertex AI Gemini 1.5 Pro workload ran about $120/month, versus roughly $360/month for a comparable AWS Bedrock Claude 3.5 deployment before reserved discounts, and around $450/month on Azure AI Studio with GPT-4 Turbo before credits.
- Large-scale AI training buyer (per Zarif Automates’ June 2026 benchmark): At a workload of 1B input and 200M output tokens per month, AWS Bedrock’s Claude Sonnet 4.6 stack landed near $72,000/month, compared to roughly $3.25 million a year on both Azure OpenAI with GPT-5 and Google Cloud with Gemini 2.5 Pro at list pricing for the same volume.
- GPU marketplace arbitrage (per AIMultiple’s GPU index, July 2026): Teams willing to use non-hyperscaler marketplaces alongside their primary cloud found B200 rates as low as $3.75/hr on Packet.ai versus $16.11/hr on Google Cloud on-demand, an over 4x spread for the same chip generation.
- FinOps-driven H100 cost reduction (per nOps’ August 2026 report): Teams combining GCP preemptible instances with committed-use baseline capacity reported effective blended H100 rates well under half of pure on-demand Azure pricing for similar total compute-hours.
FinOps Strategies for Controlling GPU Cloud Costs
Regardless of which cloud you land on, the FinOps playbook for 2026 GPU spend looks broadly similar across teams that have gotten costs under control.
Blend spot/preemptible capacity with a smaller reserved baseline. Since spot H100 pricing runs roughly 75-80% below on-demand across all three clouds, teams that can tolerate interruption (checkpointed training jobs, batch inference) should default to spot first and reserve capacity only for latency-sensitive production inference. Model-tier routing is the second-biggest lever: sending simple classification or extraction tasks to a cheap tier like Gemini 2.5 Flash-Lite or Claude Haiku 4.5, while reserving GPT-5.5 Pro or Claude Opus for tasks that actually need frontier reasoning, can cut blended token costs by an order of magnitude without a quality hit on the bulk of traffic.
Multi-cloud GPU brokering is gaining traction too, where teams route training jobs to whichever cloud has the cheapest available capacity that week using tools that abstract over AWS, Azure, and GCP APIs. This adds operational complexity but can meaningfully undercut single-cloud commitment pricing for teams with flexible workload scheduling. Finally, egress costs deserve their own line item: moving large training datasets or model checkpoints between clouds isn’t free, and teams that don’t budget for it are frequently surprised by a five-figure transfer bill mid-migration.
The Verdict: Which Cloud Wins for GPU and AI Compute
There’s no single winner, but the data points to clear defaults by scenario. For raw cost efficiency on both GPU rental and token pricing, Google Cloud comes out ahead in the majority of 2026 comparisons, with the lowest normalized H100 rates in several indexes and the cheapest flagship model tier (Gemini 2.5 Flash-Lite at $0.10/$0.40 per million tokens) of any provider tracked. That lead evaporates on B200 pricing, where Google is currently the most expensive of the three.
AWS is the safest default for teams that value breadth over the lowest possible price: the largest market share, the broadest instance catalog, and a Claude-based Bedrock stack that undercut both rivals by nearly 40x in at least one large-scale cost comparison ($72,000/month versus roughly $270,000/month equivalent for the Azure and GCP flagship-model comparisons in the same study). Azure remains the only rational choice for organizations whose AI roadmap is built specifically around GPT-5.x, and its steep three-year commitment discounts (up to 62%) can close much of the on-demand price gap for teams with predictable, long-horizon GPU demand. But on both H100 on-demand pricing and flagship model token costs, Azure was the most expensive option in nearly every 2026 report we reviewed.
The practical takeaway for August 2026: price your actual workload shape (node size, commitment horizon, spot tolerance, and model tier) on all three clouds before committing, because the “which cloud is cheapest” answer changes depending on whether you’re comparing full 8-GPU nodes (near-parity, all three within 14 cents an hour) or normalized per-GPU spot pricing (up to 4x apart). Don’t assume last year’s market-share leader is this year’s price leader for GPU compute. It isn’t.
Frequently Asked Questions
Which cloud has the cheapest H100 GPU pricing in 2026?
Google Cloud generally shows the lowest normalized per-GPU H100 on-demand rate in 2026 indexes, as low as $3.00-$3.50/hr on its A3 series in some regions, though at the full 8-GPU node level AWS, Azure, and GCP are within about 14 cents an hour of each other, roughly $98.32-$98.46/hr.
Is Azure always the most expensive cloud for GPU compute?
On normalized per-GPU on-demand H100 pricing, yes, Azure’s ND H100 v5 series consistently priced highest ($12.25-$14.50/hr) in 2026 pricing surveys. Azure partially offsets this with the steepest committed-use discounts, up to roughly 62% at a 3-year term.
What’s the difference between H100 and B200 pricing?
B200 is Nvidia’s newer Blackwell-generation GPU and generally costs more to rent than H100. A July 2026 index put the B200 median rental price around $6.25/hr with a range of $3.75-$16.11/hr across providers, versus H100’s roughly $1.65-$14.19/hr range in the same period. For teams weighing rental against outright purchase, Mercatus AI’s June 2026 pricing sheet put a new H100 SXM5 GPU at $25,000-$30,000, so even the higher end of on-demand rental only catches up to the cost of buying after many months of continuous use.
Should I use spot/preemptible GPU instances for AI training?
For checkpointed, interruption-tolerant workloads like most large-scale training runs, spot and preemptible instances cut H100 costs by roughly 75-80% versus on-demand across AWS, Azure, and Google Cloud, and EMMA’s May 2026 index found spot H100 capacity as low as $1.25 an hour against a $12.29-an-hour Azure on-demand ceiling in the same dataset, nearly a 10x spread. They’re a poor fit for latency-sensitive production inference that can’t tolerate interruption.
Which cloud offers the cheapest AI model API pricing?
Google’s Vertex AI consistently priced lowest across model tiers in 2026 comparisons, with Gemini 2.5 Flash-Lite at $0.10 input / $0.40 output per million tokens. AWS Bedrock’s Claude models sit in the middle, and Azure OpenAI’s GPT-5.5 tiers, especially the Pro tier at $30/$180 per million tokens, priced highest.
Do AWS, Azure, and Google Cloud all offer non-Nvidia AI chip alternatives?
Yes. AWS offers Trainium2/Trainium3, Google Cloud offers its TPU v6/v7 lineup (the most mature of the three), and Azure offers the Maia 200 accelerator, though Maia has more limited availability than the other two options as of August 2026.
How much does it cost to migrate a GPU training workload between clouds?
Costs vary by dataset size, but egress fees for multi-terabyte training datasets alone can run into the thousands of dollars, before accounting for engineering time to re-containerize and re-validate on the new cloud’s GPU driver stack. Budgeting for a parallel shadow-test period before full cutover is standard practice in 2026 migrations.
Is cloud market share still a useful signal for choosing a provider?
Less than it used to be for AI-specific workloads. AWS’s 28-31% share, Azure’s 21-25%, and Google Cloud’s 12-14% (2026 estimates) reflect overall IaaS revenue, not GPU pricing or AI model quality, which is why buyers increasingly compare workload-specific costs rather than defaulting to the market leader.
Related Coverage
- AWS Graviton vs Intel/AMD: 45% Cheaper, 25% Faster [2026]
- AWS Hikes EC2 GPU Pricing 20%, Second Time in 2026
- Nvidia Vera CPU Takes On Intel, AMD With $20B Bet [2026]
- Microsoft Azure Earnings: $678B Backlog, Up 84% [2026]
- Backblaze B2 vs Google Cloud vs Azure: $6.95 vs $23/TB [2026]
- Nvidia, SK Hynix Ink $500B AI Memory Deal [2026]
- How to Set Up vLLM: 12 Steps, 90 Min [2026]


