Key Moments
Which GPU Clouds Are Actually Good? | ClusterMAX 3.0
Key Moments
GPU clouds often have critical cybersecurity flaws, with one instance exposing national intelligence data, yet the market for AI compute is becoming incredibly difficult to access and more expensive.
Key Insights
ClusterMAX 3.0 now tracks 323 providers, a significant increase from earlier versions.
Major cybersecurity issues persist in new GPU clouds, including improperly implemented network controls and outdated software with documented vulnerabilities.
Google has shown consistent improvement across ClusterMAX generations, moving from bronze to gold tiers.
The market for AI compute is significantly tougher, with inference providers operating at up to 60% gross margins, making it difficult for new labs to secure resources.
NVIDIA's recent acquisitions of Poolside and Hugging Face are seen as defensive moves to maintain their ecosystem dominance and prevent competitors from gaining strategic assets.
Chinese GPU manufacturers, particularly Huawei, are rapidly improving but still lag behind NVIDIA in production volume and overall performance.
Evolution of ClusterMAX and its rigorous evaluation criteria
The ClusterMAX project, spearheaded by SemiAnalysis, has evolved significantly since its inception. Initially, the evaluation of GPU clouds was largely manual, with a focus on Slurm. The subsequent versions, particularly ClusterMAX 3.0, have introduced extensive automation, a broader range of providers (now tracking 323), and a more comprehensive testing methodology. This includes not only configuration, workload performance, and reliability but also response to failures and crucial cybersecurity assessments. The research aims to provide a holistic view rather than relying on isolated benchmarks, helping organizations make informed decisions about where to deploy their AI workloads.
Pervasive cybersecurity vulnerabilities in new GPU clouds
A critical finding from ClusterMAX 3.0 is the alarming prevalence of cybersecurity flaws in many new GPU cloud providers. These issues range from basic misconfigurations like outdated software and improperly secured networks to more severe vulnerabilities. In one striking example, sensitive national intelligence data was exposed due to lax security controls. This lack of basic security hygiene is surprising, as exploiting these known vulnerabilities does not require advanced AI techniques, merely basic reconnaissance. The research highlights that even simple checks like software version updates and proper network segmentation are often overlooked, posing significant risks to users and their data. The ClusterMAX CLI tool (cmax) was developed to help users identify and address these basic security misconfigurations.
The escalating difficulty and cost of accessing AI compute
Securing AI compute resources has become considerably more challenging and expensive compared to previous years. The rise of highly profitable inference providers, operating with gross margins up to 60%, has intensified competition for GPUs. This means that new AI labs, often funded by early-stage venture capital, are now competing not only against other research labs but also against profitable commercial entities for limited resources. This has led to a contraction in the available compute for training, with labs that once rented thousands of GPUs now seeking allocations of only around 1,000. The market is consolidating, with smaller allocations becoming the norm for emerging labs, further complicating access and increasing overall costs.
Google's sustained improvement and the changing AI compute hierarchy
Within the evolving GPU cloud landscape, Google has demonstrated consistent progress, moving up the rankings across successive ClusterMAX iterations. Starting at a lower tier, Google has advanced to the gold category, indicating significant improvements in their offerings. This upward trajectory is contrasted with many other providers whose rankings have stagnated or declined due to a lack of substantial upgrades. The conversation also touches upon the broader compute hierarchy, noting that while Google possesses vast compute resources, companies like OpenAI and Anthropic are increasingly commanding larger portions of the most advanced hardware through dedicated partnerships and large-scale leasing. This shift suggests that sheer scale alone is no longer a guarantee of dominance; strategic partnerships and focused R&D compute are becoming paramount.
NVIDIA's strategic acquisitions of Poolside and Hugging Face
NVIDIA's recent acquisitions of Poolside, a company known for its work on open-source models and infrastructure, and Hugging Face, a vital hub for the AI ecosystem, are viewed as strategic defensive moves. These acquisitions aim to consolidate NVIDIA's control over its ecosystem and prevent key players from falling into the hands of competitors. By acquiring talent and critical platforms, NVIDIA ensures that these assets remain within its sphere of influence, potentially bolstering its own offerings like the Neuron compiler and preventing rivals from leveraging them against NVIDIA. The move also signals NVIDIA's potential waning faith in open-source contributions from US entities and a desire to maintain a competitive edge by controlling key parts of the AI development pipeline.
The emergence of China's domestic GPU capabilities
Chinese GPU manufacturers, particularly Huawei, are showing rapid advancements. Despite past US export restrictions, companies like Huawei have quickly released more powerful chips, such as the Ascend 940, 950, and 960. While still trailing NVIDIA in production volume and overall performance, these domestic alternatives are gaining traction, especially for inference workloads. Several Chinese labs are reporting significant usage of non-NVIDIA accelerators, indicating a growing diversification of the hardware landscape. As China invests heavily in its national AI strategy, production volumes are expected to increase substantially, potentially forcing NVIDIA to accelerate its own innovation cycles to maintain its leading position.
Critiques and clarifications regarding ClusterMAX's methodology
The ClusterMAX report has faced several criticisms, primarily centered around its scope and perceived biases. One major point of confusion is that ClusterMAX focuses specifically on 'managed clusters'—services offering ease of use, automated failure recovery, and managed orchestration. This excludes other areas like bare-metal provisioning, data center construction, or pure inference endpoints, where some providers excel but may not perform as well in the managed cluster category. Another common critique questions whether the rankings are paid for. SemiAnalysis clarifies that they do not charge for rankings, as it would be illegal and compromise their integrity. Their revenue comes from selling research and data to buyers within the infrastructure supply chain, and maintaining their reputation for truthfulness is paramount. The challenge of balancing detailed reporting with reader accessibility is also acknowledged, with plans for more granular data dashboards and specialized research products.
Mentioned in This Episode
●Products
●Software & Apps
●Companies
●Organizations
●Studies Cited
●Concepts
●People Referenced
Common Questions
Cluster Max is a platform developed by SemiAnalysis that tests and ranks cloud providers based on 10 criteria, including performance, security, reliability, ease of use, pricing, and availability. It aims to provide comprehensive evaluations beyond individual benchmark tests.
Topics
Mentioned in this video
A platform that tests and ranks cloud providers based on various criteria. The discussion covers its evolution from version 1.0 to 3.0.
A research product offered by SemiAnalysis for licensing, providing insights into cloud costs.
A model trained by Anthropic, reportedly on Google's infrastructure.
A workload manager for cluster computing that was a significant focus in the first version of Cluster Max, but Kubernetes became more important for later versions.
A future product from SemiAnalysis that will test serverless inference endpoint providers.
Its release in 2022/2023 marked a significant boom in the cloud market, leading to the emergence of many new cloud companies.
A container orchestration system that became a key focus for testing in Cluster Max versions after the first.
Mentioned as one of the traditional cloud providers.
A Chinese company that has used 50,000 non-NVIDIA AI chips for inference.
One of the major cloud providers considered a 'traditional' cloud, contrasted with the 'new clouds' focused on AI.
Discussed in detail regarding its AI strategy, its role in training Anthropic's models, its cloud infrastructure, and its potential competitive challenges.
NVIDIA's parallel computing platform and API, whose dominance is being challenged by the diversification of the chip market.
Mentioned as a company still attempting to work with open-source models, in contrast to NVIDIA's perceived loss of hope in American open-source.
A product from SemiAnalysis focused on testing inference performance of chips and systems.
A major AI lab that is a significant renter of cloud compute, with partnerships with AWS, Google Cloud, and CoreWeave. Also building its own data centers.
A company where the speaker had an early investment, but it was ranked as underperforming.
Jordan's previous employer for 10 years, where he designed hardware systems for emerging cloud companies.
Mentioned as a company that, like Poolside and Meta, might consider offering cloud services opportunistically.
A cloud provider discussed for its bare metal infrastructure and short-term contracts, making it agile in a rising GPU price market.
A company with a large data center potentially going online, discussed in the context of its bare metal infrastructure versus managed services.
Mentioned as one of the world's largest companies, likely a client of SemiAnalysis's research, though not directly related to Cluster Max's ranked clouds.
A cloud provider that rents out resources, but is heavily booked by Anthropic, limiting availability for others. The speaker has an investment in the company.
A company acquired by NVIDIA, recognized for its work on building and training AI models and its infrastructure expertise.
A competitor to NVIDIA in the chip market, mentioned as having its own offerings and potential for multi-chip systems.
A traditional cloud provider and a subject of discussion regarding its strategy, investments in AI, and its own cloud services (Google Cloud).
An older cloud provider mentioned in the context of traditional clouds.
Google's AI research lab, discussed in relation to its compute needs and the leadership changes within Google.
Mentioned as one of the traditional cloud providers.
Mentioned as a potential 'new cloud' provider and in relation to its partnership with Google.
An older cloud provider mentioned in the context of traditional clouds.
A company mentioned for its bare metal infrastructure and data center building capabilities, but ranked lower for its managed cluster offerings.
One of several Chinese companies advancing in the AI chip space.
Mentioned as an older cloud provider with AI capabilities.
Mentioned as a company involved in open-sourcing some of its work.
Mentioned as a type of company that appreciates Cluster Max's methodology.
A startup in the chip space mentioned as entering the market.
A chip startup mentioned as entering the market.
A leading AI lab that rents significant cloud compute and is discussed in the context of its resource needs and strategic decisions.
Separating its chip business into a standalone company, indicating growth in China's chip ecosystem.
A key player in the GPU market, mentioned in relation to cybersecurity issues and its strategic acquisitions of Poolside and Hugging Face.
Acquired by NVIDIA, it's seen as a defensive move to keep a vital AI ecosystem entity within NVIDIA's control.
Mentioned as a buyer of Iris Energy's bare metal infrastructure and a partner of OpenAI.
Mentioned as one of the world's largest companies, likely a client of SemiAnalysis's research, though not directly related to Cluster Max's ranked clouds.
One of several Chinese companies advancing in the AI chip space.
A cloud provider that has partnered with Anthropic and is also a customer of SemiAnalysis's research. Mentioned in the context of managed clusters and bare metal infrastructure.
Mentioned as a partner of Anthropic, alongside AWS.
A chip startup mentioned as entering the market.
Elon Musk's AI company, mentioned in the context of its potential cloud needs and availability.
Mentioned as a type of company that appreciates Cluster Max's methodology.
A Chinese company whose AI chips are rapidly improving, particularly in inference. They are seen as NVIDIA's main competitor in China.
Mentioned as a company that might explore offering cloud services opportunistically.
A company developing AI chips and infrastructure, mentioned in relation to OpenAI and potential heterogeneity in cloud offerings.
A major cloud provider and chip manufacturer (Trainium), discussed in its relationship with Anthropic and potential future chip production.
A chip startup mentioned as entering the market.
One of several Chinese companies advancing in the AI chip space.
Mentioned as an AMD cloud provider that SemiAnalysis has invested in but ranks poorly.
CEO of Google Cloud, tasked with 'fixing' the company's cloud business.
A former Google researcher who left to found Discovery Loop, discussed in relation to Google's retention strategies.
Mentioned as the leader of a business model that others should consider.
Founder of DeepMind, whose departure from leadership at Google was seen as a 'black swan' event.
Mentioned as the CEO of Cisco, used as an example for leadership analogy.
Quoted as stating that many new clouds have 'terrible cybersecurity', a sentiment echoed by the speakers.
Mentioned as another former Google researcher who left to found a company, likely Discovery Loop.
Mentioned in the context of leadership departures, specifically comparing his situation to others.
Used as an analogy for a significant leadership departure.
Mentioned in the context of SpaceX and xAI, and his broader ventures.
A recent Huawei AI chip release.
A generation of NVIDIA GPUs discussed in the context of memory capacities and batch sizes affecting training.
Mentioned alongside InfiniBand as a networking technology with different security aspects.
Mentioned alongside NVIDIA's LPU as part of future chip offerings.
A networking technology with specific key configurations (M-keys and P-keys) that are part of network security checks.
Google's Tensor Processing Unit, used by Anthropic and discussed in the context of Google's internal hardware strategy.
A Huawei AI chip mentioned as part of their improved offerings.
A recent Huawei AI chip release.
A high-end NVIDIA GPU, mentioned in comparison to production volumes of other chip manufacturers.
The NVIDIA product that Poolside's acquired talent is intended to improve.
Language Processing Unit, a type of chip mentioned as part of NVIDIA's future offerings.
A Huawei AI chip mentioned as part of their improved offerings.
An AI accelerator chip from Amazon that can be used with Hugging Face models.
A high-end NVIDIA GPU, mentioned in comparison to production volumes of other chip manufacturers.
A generation of NVIDIA GPUs discussed in the context of memory capacities and batch sizes affecting training.
A recent Huawei AI chip release.
More from Latent Space
View all 260 summaries
104 minRecursive Language Models — Alex Zhang, MIT PhD
41 minInside OpenAI DevDay: Superhuman Computer Use, Decisions API, and the AI Cloud — Ari & Nikunj
95 minThe Future of Claude Code: Mods, Mutable Software, & Multiplayer Agents — Thariq Shihipar, Anthropic
99 minRunway’s Bet Beyond Video: World Models, Robotics, and the Neural OS — Anastasis Germanidis
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free