Nvidia’s newest AI accelerator is no longer a slide in a keynote deck. The GB300 “Blackwell Ultra” platform is shipping in volume to Microsoft, Oracle Cloud, CoreWeave and other hyperscalers, and the numbers behind that ramp landed on August 27, 2026, when Nvidia reported fiscal second-quarter revenue of $96.2 billion, according to the company’s own financial results release. Data center revenue hit a record $89.0 billion, up 117% from a year earlier, and Nvidia said the jump was driven by the ramp of its Blackwell Ultra infrastructure.
That single line explains a lot of what has happened to the AI hardware market since mid-August. CEO Jensen Huang has described the GB300 platform as being in full-scale mass production, and the rack-scale systems built around it, badged GB300 NVL72, are now the default configuration hyperscalers order when they want the fastest inference and training throughput Nvidia sells. For engineers and infrastructure buyers trying to plan the next 12 months of AI capacity, GB300 availability, pricing and how it stacks up against Google’s new TPUs, Cerebras’ wafer-scale systems and Microsoft’s homegrown chips have become the questions that matter most.
Don't miss new tech stories on Google
Add Tech Insider once in the Google app and our stories appear in your news suggestions.
What the GB300 Blackwell Ultra actually is
The GB300, sold under the Blackwell Ultra name, is Nvidia’s mid-cycle refresh of the Blackwell architecture that succeeded Hopper (H100/H200) in 2025. According to specs compiled by GPU marketplace Spheron Network, each B300 GPU carries 288GB of HBM3e memory across eight 12-Hi stacks, up from 192GB on the original B200, with 8 TB/s of memory bandwidth, and Nyaws Daily reported in May 2026 that the added memory bandwidth translates into roughly 40% higher FP8 inference throughput than B200 on comparable workloads. Nvidia’s own GB300 NVL72 product page describes the full rack as 72 Blackwell Ultra GPUs paired with 36 Grace CPUs in a single liquid-cooled cabinet, pooling roughly 20TB of fast memory as one coherent system, and the fifth-generation NVLink Switch fabric tying those racks together can now scale a single cluster to as many as 576 GB300 GPUs, the same Nyaws Daily report noted.
One correction worth flagging up front: several early reports conflated GB300 with HBM4 memory. It doesn’t use it. GB300 runs on HBM3e, the same memory family as B200, just with more stacks per package. HBM4 is reserved for Nvidia’s next architecture, code-named Rubin, which the company unveiled at CES 2026 in January and is targeting for broader availability later this year into 2027, reportedly built on TSMC’s 3nm process. GB300 is the bridge product carrying Nvidia’s data center business through that transition, and going by the earnings numbers, it’s carrying a lot of revenue with it.
Power draw is the other number worth sitting with. A single GB300 GPU has a 1,400W TDP, up from 1,000W on B200. Multiply that across 72 GPUs per rack and a fully loaded GB300 NVL72 cabinet pulls somewhere between 132kW and 155kW, per figures compiled from supply-chain and system-integrator reporting; a June 2026 teardown from VRLA Tech pegged the figure at 132kW to 140kW while measuring 1.1 EFLOPS of FP4 compute and 130 TB/s of aggregate NVLink bandwidth moving between the rack’s 72 GPUs. That’s a small power plant’s worth of draw for a single server rack, and it’s a big reason data center operators have spent the past year racing to secure liquid cooling capacity and grid interconnects.
Nvidia GB300 vs B200 vs H200 vs H100: the spec jump
The generational gap between GB300 and the Hopper-era chips still running in most data centers today is the clearest way to understand why hyperscalers are paying up to swap fleets. Per Spheron’s compiled benchmarks, Blackwell Ultra’s FP4 throughput runs roughly 11 to 15 times the per-GPU inference throughput of an H100 on large language model workloads, with the B300 estimated to hit more than 100,000 tokens per second at FP8 and above 150,000 tokens per second at FP4 on Llama-class 70B inference runs. On agentic workloads specifically, FutureTimeline reported in February 2026 that the Blackwell Ultra platform can deliver up to 50 times higher inference performance per watt than the prior generation, underscoring that power efficiency, not just raw throughput, has become as much a selling point for the chip as speed.
| Spec | H100 | H200 | B200 | B300 (Blackwell Ultra) |
|---|---|---|---|---|
| Memory | 80GB HBM3 | 141GB HBM3e | 192GB HBM3e | 288GB HBM3e |
| Memory bandwidth | 3.35 TB/s | 4.8 TB/s | 8 TB/s | 8 TB/s |
| FP8 compute | ~1,979 TFLOPS | ~1,979 TFLOPS | 4,500 TFLOPS | 7,000 TFLOPS |
| FP4 compute | N/A | N/A | 9,000 TFLOPS | 15,000 TFLOPS |
| TDP (per GPU) | 700W | 700W | 1,000W | 1,400W |
| NVLink generation | NVLink 4 (900 GB/s) | NVLink 4 (900 GB/s) | NVLink 5 (1.8 TB/s) | NVLink 5 (1.8 TB/s) |
| Networking | ConnectX-7 (800G) | ConnectX-7 (800G) | ConnectX-7 (800G) | ConnectX-8 (1.6T) |
Cloud pricing has moved along with the specs. GPU rental marketplaces were quoting roughly $9.08 an hour for on-demand B300 access as of July 2026, compared with around $3.70 an hour for H200 and a range of $2.01 to $3.92 an hour for H100, depending on PCIe versus SXM5 configuration, per Spheron’s pricing tracker. That’s a real premium, and it’s one enterprise buyers are apparently willing to pay: Nvidia’s data center segment grew 18% sequentially in the same quarter GB300 shipments ramped, on top of the 117% year-over-year jump.
Who’s actually deploying GB300 right now
Microsoft Azure, Oracle Cloud and CoreWeave are the three hyperscale and neocloud names most frequently cited as early GB300 NVL72 deployers, with TechGadgetOrbit reporting as early as February 2026 that all three already had Blackwell Ultra racks live, building out the racks in existing and new data center campuses to support both first-party AI products and rented GPU capacity for enterprise and AI-lab customers. Nvidia’s earnings commentary pointed to broad-based demand: hyperscale revenue more than doubled year over year, and a category Nvidia calls ACIE (AI cloud, enterprise and other) revenue grew 138% year over year and 25% sequentially, which the company attributed to AI-native startups, enterprises and sovereign AI programs alongside the hyperscalers.
The financing side of this buildout has been just as active as the hardware itself. Neocloud provider Lambda, which is backed by Nvidia, has reportedly raised roughly $1 billion in private debt specifically to buy Nvidia AI chips it then leases to Microsoft, according to TechCrunch’s AI coverage. Separately, Anthropic is reported to have signed a cloud computing agreement worth roughly $35 billion with Lambda to expand the GPU infrastructure behind Claude’s training and serving, per financial reporting from TradingKey. Whether every dollar figure in these deals is final or still being negotiated, the pattern is consistent: capital is flowing toward whoever can get GB300-class capacity online fastest.
The competition isn’t waiting around
Nvidia’s dominance in AI training and inference silicon is real, but 2026 has been the year every major cloud provider and several AI labs made clear they don’t want to depend on a single vendor for their most important workloads. Three moves stand out.
Google introduced its eighth-generation TPU lineup at Cloud Next ’26, splitting the chip into two specialized designs instead of one general-purpose part. TPU 8t is built for large-scale training: a full superpod links 9,600 chips with 2 petabytes of shared HBM and 121 FP4 exaFLOPS of pod-level compute, which Google says delivers up to 2.7x better performance-per-dollar than the prior Ironwood generation, according to coverage of the Cloud Next announcements. TPU 8i is tuned for low-latency inference instead, with 10.1 FP4 petaflops per chip, 288GB of HBM and 8.6 TB/s of memory bandwidth, and Google claims an 80% performance-per-dollar gain over Ironwood on large mixture-of-experts models.
Cerebras took a different route entirely. Its CS-4 system, unveiled August 19, 2026 at a launch event in San Francisco, stacks three of the company’s dinner-plate-sized WSE-3 Turbo wafer-scale chips into a single rack connected by a new interconnect scheme called Nexus, claiming around 750 petaFLOPs of AI compute and inference speeds the company says can run up to 30 times faster than traditional GPU alternatives. Notably, OpenAI is reportedly using Cerebras hardware for the fastest serving tier of its GPT-5.6 Sol model, with early reports putting output speed at roughly 1,300 tokens per second on CS-4, nearly double what the same tier delivered at its mid-August launch.
Microsoft, meanwhile, is preparing to unveil its own Maia 300 chip this fall, potentially as soon as September 2026, according to reporting from The Information cited by Reuters and multiple other outlets. Microsoft is reportedly negotiating with TSMC for capacity to produce upward of 300,000 Maia 300 units for delivery in 2027, with an eventual ambition north of a million units, though supply constraints make that stretch goal uncertain. The move builds on Maia 200, Microsoft’s second-generation in-house AI chip introduced in January 2026 on TSMC’s 3nm process, and reflects a broader push to cut the company’s dependence on Nvidia silicon for its own Azure and Copilot workloads.
Competitive landscape: GB300 vs the alternatives
| Platform | Maker | Launch status (as of Sept 2026) | Headline compute claim | Primary use case |
|---|---|---|---|---|
| GB300 Blackwell Ultra (NVL72) | Nvidia | Mass production, shipping | 15,000 FP4 TFLOPS per GPU; 20TB pooled memory per rack | Training + inference, general purpose |
| TPU 8t | Announced, cloud rollout underway | 121 FP4 exaFLOPS per 9,600-chip superpod | Large-scale pre-training | |
| TPU 8i | Announced, cloud rollout underway | 10.1 FP4 petaflops per chip | Low-latency inference | |
| CS-4 | Cerebras | Launched Aug 19, 2026 | ~750 petaFLOPs per rack; up to 30x GPU inference speed (vendor claim) | Ultra-fast single-stream inference |
| Maia 300 | Microsoft | Unveiling expected fall 2026 | Specs not yet disclosed | Internal Azure/Copilot workloads |
| Rubin | Nvidia | Announced CES 2026, GA later 2026 into 2027 | HBM4-based successor to Blackwell Ultra | Next-gen training + inference |
Historical context: how we got to a 1,400-watt GPU
It’s worth remembering that this level of per-chip power draw would have been unthinkable in data centers a decade ago, when a high-end server GPU pulled 250W to 300W. The shift began with Nvidia’s Ampere generation (A100) in 2020, accelerated through Hopper (H100/H200) in 2022 and 2024, and has now compounded through two Blackwell revisions inside 18 months. Each generation has traded higher power and cooling complexity for a steeper drop in cost-per-token, and cloud providers have absorbed that trade because the alternative, running more of the older, less efficient chips, costs more in aggregate power and rack space for the same throughput.
The pace is also notable. Nvidia has now shipped a new flagship data center architecture or major revision roughly once a year since 2022, compressing what used to be a two-year GPU cadence. That acceleration is itself a competitive response: Google’s TPU program, AWS’s Trainium and Inferentia chips, and now Microsoft’s Maia line all shortened Nvidia’s runway to coast on any single generation, pushing the faster Blackwell-to-Blackwell-Ultra-to-Rubin cadence buyers are seeing now.
Market impact: what $89 billion in data center revenue signals
Nvidia’s Blackwell platform, which includes both B200 and B300 variants, now accounts for close to 70% of the company’s total data center compute revenue, based on figures the company has disclosed around its recent quarters. That concentration means GB300’s success over the next two quarters will move Nvidia’s overall results more than almost any other single product decision the company has made. Nvidia guided third-quarter revenue to $108 billion, above the $104.2 billion Wall Street had modeled, and did so while explicitly assuming zero revenue from China, where export restrictions have kept Nvidia’s most advanced chips out of the market for over a year.
That guidance matters for competitors too. Every dollar of GB300 backlog Nvidia converts into shipped racks is a dollar of budget hyperscalers aren’t spending on TPUs, Trainium, or eventually Maia chips, at least for workloads where Nvidia’s CUDA software ecosystem remains the path of least resistance. But the inverse is also true: Google’s willingness to sell TPU 8t and 8i as customer-facing cloud products rather than internal-only tools, and Cerebras’ willingness to sell CS-4 capacity to a marquee customer like OpenAI for a specific workload, both signal that the “GPU shortage forces everyone to Nvidia” dynamic of the last few years is giving way to a market where workload-specific silicon gets a real look.
What GB300 means for AI training costs
For teams actually training or serving large models, the practical question isn’t which chip has the biggest exaFLOP number on a keynote slide, it’s cost per useful token. SemiAnalysis benchmarks that Nvidia itself has cited put GB300 NVL72 inference cost at roughly $0.123 per million tokens while serving 116 tokens per second per user as of June 2026, down from the $0.24 per million tokens at 102 TPS per user that SemiAnalysis measured on DGX B300 running DeepSeek-R1 just two months earlier, in April 2026. At roughly $9 an hour for on-demand B300 access versus under $4 for H200, GB300 only pays for itself if the 11x-to-15x inference throughput gain Spheron’s benchmarks describe actually materializes on real production workloads rather than idealized benchmark runs. Early adopters, largely the hyperscalers with the volume to negotiate better-than-list pricing and the engineering staff to optimize kernels for the new architecture, are best positioned to capture that gain first. Smaller AI labs and enterprises renting capacity secondhand through neoclouds may see a longer lag before GB300 pricing drops enough to make the switch a clear win over sticking with H200 fleets a while longer.
That dynamic is also why financing structures like Lambda’s roughly $1 billion debt raise to buy Nvidia chips for lease to Microsoft matter as much as the chip specs themselves. GB300 hardware is expensive enough, with rack-level estimates ranging from $3.7 million to as high as $6.5 million depending on the source and configuration, per third-party analyst notes, and with per-GPU list estimates that Tech-Insider put at $40,000 to $50,000 for a standalone B300 and around $300,000 for a full DGX B300 system as of March 2026, that most of the market is accessing it through rental agreements rather than outright purchase. Spot pricing has offered a cheaper on-ramp for those rentals: Spheron Network data cited by NeuralCoreTech showed B300 spot rates around $2.45 an hour per GPU as of April 2026, well below the on-demand rate the market has since settled into. Nvidia does not publish an official list price for the NVL72 system, so every dollar figure circulating is a supply-chain estimate, not a confirmed number.
TSMC and the packaging bottleneck behind GB300 supply
None of Nvidia’s roadmap matters if the chips can’t actually be built at volume, and GB300 depends on a packaging process almost as constrained as the wafers themselves. Every Blackwell Ultra GPU relies on TSMC’s CoWoS-L advanced packaging to bond the compute die to its HBM3e stacks, a process with far less spare capacity than standard wafer fabrication. TSMC has spent much of 2026 expanding CoWoS output specifically to keep pace with Nvidia, Google, Microsoft and AMD orders simultaneously, and that shared bottleneck is a big reason Microsoft’s own Maia 300 chip, also fabricated at TSMC, is queuing for 2027 delivery rather than shipping this year. Any hiccup in CoWoS-L capacity, or in HBM3e supply from Samsung, SK hynix and Micron, would ripple through GB300 shipment schedules well before Nvidia’s design or Nvidia’s demand becomes the limiting factor.
What GB300 means if you’re planning AI infrastructure spend
For engineering and infrastructure leaders weighing GB300 against sticking with existing H100 or H200 fleets, the calculus comes down to workload shape more than raw specs. Teams running latency-sensitive inference at scale, where the 11x-to-15x throughput gain Spheron documented actually shows up in production, have the clearest case for paying the roughly $9-an-hour premium over H200. Teams running smaller fine-tuning jobs or intermittent workloads are less likely to see that gain translate into lower total cost, since GB300’s advantage compounds with scale and utilization. Given how tight allocation remains, the more immediate decision most teams face isn’t GB300 versus H200, it’s whether to lock in capacity now through a neocloud or hyperscaler reservation, or wait for Rubin and risk another allocation scramble when that generation ships.
Predictions: where the AI chip race goes next
- GB300 supply stays tight through Q4 2026. With Nvidia guiding to $108 billion in next-quarter revenue and hyperscale demand still outpacing shipments, expect continued allocation battles among Microsoft, Oracle, CoreWeave and newer neoclouds rather than open-market availability.
- Rubin becomes the next flashpoint by early 2027. With HBM4 samples already reaching select hyperscaler labs, pricing and performance leaks on Nvidia’s Rubin platform will likely dominate AI hardware coverage once GB300 output stabilizes.
- Microsoft’s Maia 300 unveiling narrows the internal-chip gap, but doesn’t close it. Even if Microsoft hits its 300,000-unit target for 2027, that’s a fraction of Nvidia’s shipment volume, meaning Maia will likely remain a cost-control tool for specific internal workloads rather than a full Nvidia replacement in the near term.
- Google’s TPU 8t and 8i push more workloads toward specialized silicon. Expect more AI labs to follow Google’s lead and split training and inference onto different hardware, rather than running both on the same general-purpose GPU fleet.
- Power and cooling, not chip supply, becomes 2027’s real bottleneck. At 132kW to 155kW per GB300 rack, with Rubin expected to push power draw higher still, data center operators’ ability to secure grid interconnects and liquid cooling capacity may constrain AI buildouts more than silicon availability does.
Frequently asked questions
What is the Nvidia GB300 Blackwell Ultra?
GB300 is Nvidia’s Blackwell Ultra data center GPU, an upgraded version of the B200 chip with 288GB of HBM3e memory per GPU instead of 192GB, higher FP4 and FP8 compute, and a 1,400W power envelope. It ships as part of the GB300 NVL72 rack, which combines 72 GPUs with 36 Grace CPUs in one liquid-cooled system.
Does GB300 use HBM4 memory?
No. GB300 uses HBM3e memory, the same family as the prior B200 chip, just with more memory stacks per package. HBM4 is associated with Nvidia’s next architecture, Rubin, which the company has not yet fully launched.
How much does a GB300 NVL72 rack cost?
Nvidia does not publish an official price. Third-party analyst estimates for a fully configured GB300 NVL72 rack range from roughly $3.7 million to $6.5 million, depending on the source and system configuration. Cloud rental pricing for individual B300 GPU access was around $9.08 an hour on demand as of July 2026.
Which companies are deploying GB300 first?
Microsoft, Oracle Cloud and CoreWeave are among the earliest named deployers of GB300 NVL72 systems, based on industry reporting. Nvidia’s own earnings commentary also cited broad hyperscale and enterprise demand, without naming every customer.
How does GB300 compare to Google’s TPU 8t and TPU 8i?
GB300 is a general-purpose GPU aimed at both training and inference. Google split its eighth-generation TPU into two specialized chips instead: TPU 8t for large-scale training, with 121 FP4 exaFLOPS per superpod, and TPU 8i for low-latency inference, with 10.1 FP4 petaflops per chip. Google claims meaningful performance-per-dollar gains over its prior Ironwood generation for each use case.
Is Nvidia’s Rubin platform out yet?
Not for general availability. Nvidia unveiled Rubin at CES 2026 in January, and hyperscaler labs reportedly have early samples as of August 2026, but broad availability is expected later in 2026 into 2027, with GB300 Blackwell Ultra serving as the current-generation flagship in the meantime.
Why did Nvidia’s data center revenue jump 117% year over year?
Nvidia attributed the record $89.0 billion in quarterly data center revenue directly to the ramp of Blackwell Ultra infrastructure, meaning GB300 shipments to hyperscalers and cloud providers, alongside continued demand for the earlier B200 chips still in the Blackwell family.


