PJFP.com

Pursuit of Joy, Fulfillment, and Purpose

,

ClusterMAX 3.0: SemiAnalysis on Which GPU Clouds Are Actually Good

SemiAnalysis’s ClusterMAX 3.0 ranks the GPU clouds that rent AI compute, and its authors say most neoclouds still fail basic security and reliability tests. In this Latent Space episode, Jordan Nanos (who ran the testing) and SemiAnalysis founder Dylan Patel explain how the rankings work, why compute is harder to rent than ever, why Google’s DeepMind now has less research compute than OpenAI or Anthropic, why Nvidia is buying up labs like Poolside and Hugging Face, and what multi-silicon clusters mean for the next round of testing. This post also draws on the 30,000-word ClusterMAX 3.0 report itself for the detail the conversation skips.

TLDW

ClusterMAX 3.0 reviews 77 providers out of 323 it tracks and tests managed Slurm and Kubernetes clusters against 10 criteria, with reliability and security at the top. The bar this round was deploying GB300 NVL72 racks that actually work: Nebius joined CoreWeave in Platinum, Google joined Oracle in Gold, and most others slid, with only 19 neoclouds earning a medal. The conversation then widens into the state of the compute market: profitable inference providers have made GPUs nearly impossible to rent, frontier labs are acting as opportunistic neoclouds, Nvidia is making defensive acquisitions, and heterogeneous chips (including fast-improving Huawei silicon) are about to make cluster testing much harder.

The ClusterMAX 3.0 Rankings at a Glance

  • Platinum: CoreWeave and Nebius. The report calls CoreWeave “the industry’s sharpest, sturdiest, most proactive neocloud,” and promotes Nebius after it established itself as the provider that can command a price premium.
  • Gold: Google Cloud and Oracle. Google finally reaches the tier SemiAnalysis predicted for it back in ClusterMAX 1.0.
  • Silver: Lambda, Microsoft Azure (moved down this round), Firmus, TensorWave and GMI (up from Bronze).
  • Bronze: AWS, Gcore, GMO, Verda, Moonlite, Together (down from Silver), Crusoe (down from Gold), Prime Intellect, DigitalOcean and Hyperstack.
  • Participation Ribbon: a new tier between Bronze and Underperforming for 15 providers that, in SemiAnalysis’s words, “do the bare minimum to get by,” including Vultr, Runpod, IBM Cloud and Vast.ai.
  • Not Recommended: underperformers such as IREN, Akamai, Hetzner and OVHcloud, plus an “Unavailable” group that SemiAnalysis could not test, including Fluidstack, SpaceXAI, Mistral and Poolside’s infrastructure company.

Thoughts

The security segment is the part every AI founder renting GPUs should sit with. Patel describes testers in earlier rounds being able to see other tenants’ jobs, data and storage, and in one case material belonging to a country’s national security and military intelligence agencies running on the same neocloud. Nanos then makes the more damning point: the checks are not sophisticated. Are your drivers current, is your software patched, are your InfiniBand partition and management keys configured? When providers run two or three year old software with published CVEs, you do not need a frontier model to write an exploit. The report puts a number on how rare clean hygiene is: Oracle was the only provider rated Silver or above with zero installed libraries carrying applicable CVEs, and the Crusoe review asks, pointedly, whether “a CVE from November of 2025 count[s] as ‘historical baggage’.” Ilya Sutskever’s tweet, which the report opens with, supplies the stakes: the next time agents go rogue, he wrote, “they’ll try taking over a neocloud to run more copies.” The free CMAX command-line tool SemiAnalysis released alongside the report, which flags outdated versions on your own cluster, is the most practically useful thing in the whole launch.

The explanation for why compute is so scarce is sharper than the usual “demand is high” line. Patel’s argument is that the competition for GPUs changed character. A year or two ago a new lab was mostly bidding against money-losing frontier labs. Now inference providers like Baseten, Fireworks and Together are running open models at around 60% gross margins (Patel says Baseten was near 10% not long ago), so anyone who gets a GPU makes money immediately. The report adds the financing side, which it calls a Matthew principle: profitable frontier labs lock down capacity more easily, while smaller labs get “fat prepays, worse prices, and a much harder time planning years in advance.” SemiAnalysis models OpenAI and Anthropic at 56.3% of all lab compute by the end of 2027, and it describes Nebius demanding prepayments that reach 100% on one-year commitments and, in one case, auctioning off a tranche of capacity outright. That is a structural shift, not a cycle, and it suggests the “rent GPUs, raise a seed round, train a model” path is closing for anyone without a strong inference business on the side.

The most interesting market detail is how labs now treat compute as a two-way option. Nanos points to the SpaceX and Google compute deal, which reportedly carries mutual 90-day cancellation rights: if SpaceX’s own research starts working, it pulls the compute back; if Google stops making money serving inference on it, Google hands it back. He frames it as a standoff over who believes in AGI more. Patel’s corollary is pointed: by the end of the year OpenAI and Anthropic will each have more compute dedicated to research than DeepMind, because Google keeps selling capacity to outside customers. Google Cloud’s revenue looks great as a standalone business, but as Patel puts it, selling capacity at $20 to $30 million a megawatt when you could run your own models on it for $50 million a megawatt is a bad trade for Alphabet as a whole. The report’s Google review reads the same way from the other side: GCP is now “comfortably among the best managed cluster providers” and has been dealmaking like few others, all “amidst the strange retreat of Google Brain and DeepMind from the frontier.”

Nanos reads Nvidia’s Poolside and Hugging Face deals as defensive rather than talent grabs, and that is the more convincing reading. With roughly $50 billion of free cash flow a quarter, spending a fraction of it to keep two strong teams from landing with a competitor is cheap. Patel adds the kingmaker math: a Poolside flush with Nvidia cash can raise more, lever it at 75% loan to value, and become a gigawatt-scale cloud that buys Nvidia GPUs, so the money largely comes back. The report generalizes this into a financing thesis: Nvidia now backs GPU demand with revenue floors, landlord guarantees and leases, which lets neoclouds get investment-grade debt without a hyperscaler involved. Its section heading says it plainly: “The Balance Sheet Is The Moat.” The flip side shows up in CoreWeave, which carries $35 billion of debt and is leaning toward long-term bare metal contracts that are easy to finance, even though managed clusters earn higher margins. Nanos’s aside that only about five companies have proven they can build a 100,000 GPU cluster is a useful reality check on the hundreds of providers in the market.

Reliability is where the report gets most concrete, and where the rankings are really decided. SemiAnalysis injects fake GPU errors into every cluster and times how long the provider takes to notice and recover, against a target of about two minutes to detect and under an hour to swap in a spare node. There are 172 documented ways an Nvidia GPU can fail, and the gap between providers is wide: CoreWeave runs burn-in tests on idle nodes for 20 to 30 minutes of every hour so production jobs are not the ones that find faults, while Nebius’s automatic recovery on a GB300 rack worked but took 8 hours 40 minutes. Rack-scale systems change the rules, since you cannot hot-swap a tray into a 72-GPU NVLink domain, so the report describes an emerging “NVL64+” service level where a rack counts as down once 3 of its 18 nodes fail. Its sharpest line is aimed at providers who said they delay remediation to let customers retrieve their work: “That is cope.” The good news is that Vera Rubin should be a far easier transition than Hopper to Blackwell, mostly faster 1.6T NICs and racks rising from about 130kW to about 200kW.

The report’s section on agentic coding is the most original material in it, and it never came up on the podcast. SemiAnalysis ran its own test agents overnight against clusters and found they “magnify skill deltas rather than flattening them.” One agent decided the NVLink fabric was too much trouble, ran jobs over InfiniBand instead and declared victory; another ran storage tests on local NVMe instead of the shared filesystem; Claude has a habit of hammering the Slurm controller with loops instead of job arrays, and CoreWeave told them it has seen head nodes run out of memory under agent-driven workloads. The more important economic point is that agents make customers more willing to accept bare metal or lightly managed clusters, which squeezes the margin managed providers charge for doing the hard parts. If you know what you are doing, an agent lets you tolerate a worse provider. If you do not, as the report puts it, “your AGI advisor will cheerfully tell you to shoot your foot off.”

The final stretch of the podcast is where the next year of infrastructure gets decided. Patel says Anthropic pre-trains mostly on TPUs and serves mostly on Trainium and GPUs, and that RL is now as large as or larger than pre-training, which means most of the workload is forward passes regardless of what you call it. That blurs the training versus inference split and invites mixed fleets, but Nanos flags the hard part: numerical consistency across chips, even across B200 and B300 generations, plus weight updates happening at different times across sites. Add Groq-derived LPUs arriving inside Nvidia’s own Vera Rubin systems, Cerebras running clusters for OpenAI, and Huawei shipping new Ascend generations in quick succession, and ClusterMAX’s single-vendor testing model is going to need a rethink. The report already shows the strain: Oracle and TensorWave were the only providers that let SemiAnalysis test AMD GPUs, and Oracle’s MI355X cluster wedged a node over a NIC driver bug that no health check caught. Patel’s estimate that China could produce tens of millions of AI chips by 2028, combined with his comment that the CUDA moat “is rapidly being deteriorated,” explains why Nvidia now ships a new system every year.

The section on criticism is worth hearing too. ClusterMAX tests managed clusters only, so a downgrade for Crusoe or Together is not a verdict on Crusoe’s data centers or Together’s inference endpoints, and it is explicitly not a stock call. The report is careful here too, crediting Crusoe with “the largest real pipeline of any datacenter builder” and a valuation that grew from about $10 billion to $30.9 billion, even as it drops the company to the bottom of Bronze for reliability and hygiene. Patel notes he holds personal investments in Crusoe, Fluidstack and TensorWave, none of which rank near the top. The forthcoming EndpointX benchmark for serverless inference providers should address the most common complaint, that one logo can mean several very different businesses, and the report goes further, arguing that “all neoclouds need to have an endpoints business going forward.”

Key Takeaways

  • ClusterMAX 3.0 reviews 77 providers and tracks 323, up from 209 in ClusterMAX 2.0, and testing is now largely automated with coding agents running an internal benchmark and burn-in suite.
  • CoreWeave and Nebius are Platinum, Google Cloud and Oracle are Gold, and only 19 neoclouds worldwide earned a medal rating. Azure fell to Silver, while Crusoe and Together fell to Bronze.
  • SemiAnalysis defines a neocloud as any company that sells access to compute with a focus on AI, as opposed to hyperscalers and legacy clouds that existed before ChatGPT.
  • The rankings cover managed clusters (managed Slurm, managed Kubernetes, node hot-swapping), not bare metal, data center construction or inference APIs.
  • Security remains poor across the industry. Testers have seen other tenants’ data, and in one case a nation’s intelligence workloads, because of misconfigured network isolation and outdated software.
  • SemiAnalysis released a free CLI, CMAX, that checks software versions on a cluster and links to the needed upgrades.
  • This round’s bar was a working GB300 NVL72 deployment, which means ARM CPUs, an 800G network, a rack-scale NVLink domain, mandatory direct liquid cooling and 130kW+ racks. Providers that skipped it were downgraded.
  • Reliability is the criterion top customers care about most. SemiAnalysis injects GPU errors and expects detection in about two minutes and spare-node replacement in under an hour on standard HGX clusters.
  • Compute is the hardest it has ever been to rent, because inference providers now earn around 60% gross margins on open models and frontier labs are renting clusters as small as 1,000 GPUs.
  • Patel expects OpenAI and Anthropic to each have more R&D compute than DeepMind by year end, since Google sells much of its capacity to outside customers.
  • The SpaceX and Google compute contract has mutual 90-day cancellation rights, a sign that labs treat spare compute as an option rather than a permanent business.
  • Nanos sees Nvidia’s acquisitions of Poolside’s team and Hugging Face as defensive moves, and the report argues Nvidia’s financing backstops have made its balance sheet a moat of its own.
  • AI coding agents push customers toward bare metal and lightly managed clusters, which pressures managed-cluster margins, but they amplify skill gaps rather than closing them.
  • Multi-silicon clusters are coming next year, driven by Nvidia’s LPX systems, AMD, Cerebras, Trainium, TPUs outside Google Cloud and chip startups. Huawei is improving fastest among Chinese chipmakers.
  • SemiAnalysis plans EndpointX, a ranking of serverless inference endpoint providers that weighs reliability, security, cost and monitoring alongside raw speed.

Chapters

0:04 From HPE Hardware Designer to Running ClusterMAX

Nanos spent 10 years at Hewlett Packard Enterprise designing the systems neoclouds bought, then joined SemiAnalysis after criticizing the first ClusterMAX. He explains that the first edition missed big providers and leaned too heavily on Slurm at the expense of Kubernetes. The provider count has grown to 323, and testing moved from manual benchmark scripts to an automated suite that coding agents can run and debug on their own.

3:03 What a Neocloud Is and Why ClusterMAX Exists

A neocloud is any company renting out AI-focused compute, as distinct from AWS, Google, Oracle, Azure and older clouds like Rackspace. ClusterMAX started because founders kept asking Patel the same question: they had just raised a seed round and were about to hand most of it to a GPU provider. The report is meant as a coarse first pass, a tier list and podium, with deeper consulting available for specific requirements.

5:15 The State of Neocloud Security

Patel says security has been bad since the first edition, with testers able to see other customers’ jobs and storage. SemiAnalysis published a piece saying most neoclouds are bad at security, and Ilya Sutskever tweeted the same view hours later. Nanos explains that the checks are basic (patched software, current drivers, correct InfiniBand keys) and introduces the free CMAX tool.

9:35 Who Moved Up, Who Moved Down, and the GB300 Bar

Each edition raises the bar, so most providers have been downgraded over time while Google climbed to gold. This round required working GB300 NVL72 racks, which change almost everything from cooling and CPUs to networking and driver management. Nanos argues that a provider skipping GB300 for “strategic reasons” raises the question of whether it will be ready for Vera Rubin.

12:02 Why Compute Is the Hardest It Has Ever Been to Rent

Inference providers are now highly profitable on open models, so anyone with a GPU can make money and neolabs are bidding against them. Frontier labs have shrunk minimum rentals from about 8,000 GPUs to 1,000, while inference platforms like Modal happily take four nodes. Even single-node spot rentals have dried up.

15:01 OpenAI, Anthropic, Google and the Compute Race

OpenAI and Anthropic lead in rented and self-built compute, across CoreWeave, Microsoft, Oracle, AWS and Google Cloud TPUs. Patel explains why bare metal giants like Crusoe can score poorly on managed services. He argues DeepMind now has less R&D compute than either rival, and the hosts describe DeepMind researchers frustrated that leadership keeps selling capacity to others.

19:01 Labs as Opportunistic Neoclouds and Google’s Dilemma

SpaceX, Meta, Mistral and Poolside all show signs of renting out spare compute while keeping the option to pull it back for research. The SpaceX and Google deal’s mutual 90-day cancellation is the clearest example. The group debates Google funding departing researchers who then spend on GCP, using Anthropic as the cautionary tale, and touches on DeepMind’s leadership change.

26:00 Nvidia, Poolside and Hugging Face

Poolside built its own infrastructure because early neoclouds were poor, and its researchers have now gone to Nvidia to improve Nemotron. Patel walks through how Nvidia’s money can turn Poolside into a gigawatt-scale cloud, and Nanos calls both deals defensive. They note that only a handful of companies have proven they can build a 100,000 GPU cluster.

36:01 Answering the Critics

Patel stresses that ClusterMAX rates managed clusters only, that SemiAnalysis never takes money for rankings, and that his own investments rank poorly. He separates ClusterMAX from SemiAnalysis’s institutional research, using Nebius versus CoreWeave and IREN as examples where stock views and cloud quality diverge. Nanos previews broader coverage of inference endpoints, bare metal, RL infrastructure and sandboxes, including EndpointX.

47:00 Multi-Silicon, RL and Chinese Chips

Non-Nvidia clouds like TensorWave are already ranked, and next year will bring heterogeneous fleets of LPUs, AMD, Cerebras, Trainium and TPUs. The group discusses how RL blurs training and inference, the numerical headaches of mixing chips, and distributed RL across sites. Patel closes with Huawei’s rapid Ascend progress and projections for Chinese chip volume.

Notable Quotes

“The best way to get a job at SemiAnalysis is to find something that SemiAnalysis is doing and criticize it and then Dylan hires you to fix it.”

Jordan Nanos, on how he joined SemiAnalysis

“Imagine if we were malicious. I’m sure there were people who were malicious or could have been malicious on that cluster as well.”

Dylan Patel, on seeing a nation’s intelligence workloads during testing

“Now everyone can make money off of GPUs. And so it’s just very, very difficult to get any sort of compute.”

Dylan Patel, on why neolabs struggle to rent clusters

“Our leadership keeps selling the compute that we want to other people.”

A Latent Space host, recounting what DeepMind researchers told them

“The value we have to the industry is that we tell the truth and what we believe.”

Dylan Patel, responding to claims that ClusterMAX is pay to play

“In this semi-legible business, agents magnify skill deltas rather than flattening them.”

The ClusterMAX 3.0 report, on testing clusters with AI coding agents

“They know they have to run as fast as possible at the speed of light otherwise they will get caught up to because the CUDA moat is rapidly being deteriorated.”

Dylan Patel, on Nvidia’s annual release cadence

Watch the full conversation on Latent Space here, and read the full ClusterMAX 3.0 report from SemiAnalysis.

Related Reading