ClusterX™GPU Cloud Rating & Ranking System
ClusterX rates managed GPU clusters on what customers actually receive: configuration, workload performance, reliability, and how the provider responds to failures. Hands-on cluster testing, documentation review, and customer feedback feed ten criteria, from security and networking to pricing and availability.
ClusterX 3.0 rankings
September 202677 providers on the board
A ranking of managed clusters. Bare metal and inference endpoints are not rated here.
Platinum
2 providers
- CoreWeave
- Nebius
Gold
2 providers
- Oracle Cloud
- Google Cloud
Silver
5 providers
- Azure
- Firmus
- Lambda
- GMI Cloud
- TensorWave
Bronze
10 providers
- AWS
- Gcore
- Verda
- Moonlite
- Prime Intellect
- Together AI
- Crusoe
- DigitalOcean
- GMO GPU Cloud
- Hyperstack
Participation Ribbon
15 providers
- Vultr
- Neysa
- VESSL AI
- Shadeform
- Runpod
- Radiant
- FPT AI Factory
- Core42
- Latitude.sh
- IBM Cloud
- BuzzHPC
- Vast.ai
- Bitdeer
- Hyperbolic
- STN
Not recommended
43 providers
Underperforming11
- Sharon AI
- IREN
- Hydra Host
- FarmGPU
- WhiteFiber
- PaleBlueDot
- Akamai
- Hetzner
- Mithril
- OVHcloud
- Massed Compute
Unavailable32
- FluidStack
- Cirrascale
- Lightning AI
- Scaleway
- CUDO Compute
- Denvr Dataworks
- Atlas Cloud
- SpaceXAI
- Mistral AI
- Poolside Infrastructure Company
- Nscale
- Highrise
- Corvex
- Andromeda
- Volta
- Firebird
- Tatra SuperCompute
- Sesterce
- Groq
- Yotta
- Darya AI
- Boostrun
- Global AI
- Argentum AI
- QumulusAI
- Alibaba Cloud
- MegaSpeed
- BytePlus
- Humain
- SK Telecom
- Naver Cloud
- Indosat
A low ClusterX rank is not a verdict on bare metal. Several providers ranked low here are strong bare metal providers, and the reverse is also true.
What the tiers mean
A relative rating. Each provider is assessed against its peers.
After evaluating a provider against the ten criteria (Security, Lifecycle, Orchestration, Storage, Networking, Reliability, Monitoring, Pricing, Partnerships, Availability) we assign one of these ratings. Gold and Platinum providers rise to the top by introducing features and functionality that others do not have.
Six of these definitions were written for the 2.x releases; the Participation Ribbon was added in 3.0. The board above shows current standings.
- Platinum
- The best GPU cloud providers in the industry. They consistently excel across evaluation criteria, are proactive and innovative, and maintain an active feedback loop with their users. In practice they command a pricing premium because their total cost of ownership is better even when the raw $/GPU-hr is higher.
- Gold
- Strong performance across all evaluation categories with some opportunities for improvement. Gold-tier providers may have small gaps or inconsistencies but are responsive to feedback. A great choice that generally wins deals, especially with the best $/GPU-hr and availability on the customer timeline.
- Silver
- An adequate offering with noticeable gaps compared to Gold or Platinum. Some users will not consider a Silver provider despite attractive pricing and availability. There is clear room for improvement, and we encourage adopting industry best practices to catch up to peers.
- Bronze
- Fulfills our minimum criteria, and the last tier we directly recommend. Common issues include inconsistent support, subpar networking, unclear SLAs, limited Kubernetes/Slurm integration, or less competitive pricing. Several providers here are making real effort to catch up.
- Participation Ribbon
- Introduced in 3.0 as a tier between Bronze and Underperforming. It more accurately describes our opinion that these providers do the bare minimum to get by.
- Underperforming
- Based on hands-on testing, these providers can quickly rise to Bronze or Silver by fixing one or more critical issues, such as offering only older GPUs, missing basic security attestation (SOC 2, ISO 27001), misconfiguring key server features (leaving PCIe ACS enabled, or failing to enable GPUDirect RDMA), or charging for GPU hours during cluster creation or hardware downtime.
- Unavailable
- An interesting service we are excited about but cannot verify yet: not launched publicly, sold out with no plans to add capacity, government-only, or otherwise untestable. We keep a “trust but verify” approach until we can complete hands-on testing.
Ten criteria
ClusterX evaluates GPU cloud providers across ten dimensions. Each one expands into requirements that differ by deployment scenario.
Security
Data protection, compliance, and infrastructure security measures including encryption, access controls, and audit capabilities.
Lifecycle
Resource provisioning, scaling, and management capabilities including automation and self-service features.
Orchestration
Container and workload management features including Slurm scheduling and Kubernetes support.
Storage
High-performance, scalable shared storage that mounts reliably and protects data, including throughput, durability, and backup guarantees.
Networking
Network performance, latency, and connectivity options including RDMA and high-speed interconnects.
Reliability
Uptime guarantees, fault tolerance mechanisms, and disaster recovery capabilities.
Monitoring
Observability, alerting, and diagnostic capabilities for infrastructure and workloads.
Pricing
Cost transparency, value for money, and flexible pricing models including spot/spot-like instances.
Partnerships
Ecosystem and integration support with major AI/ML frameworks and tools.
Availability
Geographic reach, service accessibility, and capacity planning capabilities.
Requirements by scenario
Each criterion is checked against a list of requirements. Most apply to every deployment scenario; a few are specific to Slurm, Kubernetes, or standalone machines, so the columns do not add up to the total.
| Criterion | Slurm | Kubernetes | Standalone | Total |
|---|---|---|---|---|
| Security | 20 | 21 | 20 | 21 |
| Lifecycle | 15 | 14 | 14 | 15 |
| Orchestration | 18 | 18 | 18 | 18 |
| Storage | 15 | 16 | 14 | 16 |
| Networking | 10 | 9 | 9 | 10 |
| Reliability | 24 | 25 | 24 | 25 |
| Monitoring | 16 | 14 | 14 | 16 |
| Pricing | 6 | 6 | 6 | 6 |
| Partnerships | 6 | 6 | 6 | 6 |
| Availability | 9 | 9 | 9 | 9 |
| All scenarios | 139 | 138 | 134 | 142 |
How ClusterX rates
Hands-on testing, documentation review, and customer feedback, combined into one picture of each provider.
What ClusterX is
The ClusterX™ Rating System provides a comprehensive framework for evaluating GPU cloud providers based on critical dimensions of cloud infrastructure quality and service delivery. Our independent analysis helps you make informed decisions when selecting a GPU cloud provider.
Unlike simple performance benchmarks or cost comparisons, ClusterX™ evaluates the complete ecosystem of GPU cloud services, from security and compliance to orchestration capabilities and long-term reliability.
Why it matters
The GPU cloud market has exploded with options, making it increasingly difficult for organizations to choose the right provider. Traditional evaluation methods often focus on narrow metrics like price per GPU hour or peak performance, missing critical factors that determine real-world success.
Three inputs to every rating
Hands-on testing
We deploy real workloads on each provider's infrastructure, testing everything from basic GPU performance to complex multi-node training jobs, including network performance, storage throughput, and reliability under load.
Documentation review
We analyze each provider's technical documentation, security certifications, SLA commitments, and compliance frameworks to understand their capabilities and limitations.
Customer feedback
We collect and analyze feedback from actual users across different industries and use cases, providing insight into real-world performance and support quality.
Ratings are updated as the market moves
The GPU cloud landscape evolves rapidly, with new providers entering the market and existing ones continuously improving their offerings. ClusterX™ ratings are updated regularly to reflect these changes, ensuring our evaluations remain current and relevant.
Expectations
Detailed requirements for each deployment model and operational domain.
Slurm
Slurm is an open-source job scheduler and the de facto standard for HPC for over 20 years. Used on over 60% of TOP500 supercomputers and over 50% of AI training clusters.
Kubernetes
Kubernetes is the industry-standard container orchestration software. It is the de facto standard for inference and growing in popularity for training.
Standalone
Individual GPU machines provide direct hardware access for training, inference, data processing, and research workloads. Standalone systems offer maximum flexibility and performance without orchestration overhead.
Monitoring
A complete monitoring dashboard includes high-level and low-level views of the cluster, providing comprehensive visibility into system performance, resource utilization, and potential issues before they impact workloads.
Health checks
Proactive health monitoring identifies issues before they impact workloads through both active diagnostic testing and passive continuous monitoring that automatically remediates common problems.
Releases
Previous releases remain available as historical records.
ClusterX 3.0
September 2026CurrentManaged GPU clusters evaluated through audit, performance, reliability, and fault-tolerance testing, covering 77 providers, with the market view widened to 323 providers and well over 200 end users interviewed. Only 19 neoclouds achieve a medallion rating this round, and a Participation Ribbon tier is introduced between Bronze and Underperforming.
ClusterX 2.1
April 2026Added a small set of newly tested providers rather than a full re-test. Four moved into Bronze on initial evaluation, several serious providers moved into Unavailable ahead of launch, and the GPU cluster TCO and goodput expense calculators were published.
ClusterX 2.0
November 2025Our second major evaluation, conducted from August to September 2025. Rated 84 neoclouds, up from 26, and tracked 209 providers. Interviewed over 140 end users, clarified and expanded all ten criteria with itemized requirements, and extended testing to Kubernetes clusters and standalone machines.
ClusterX 1.0
March 2025The inaugural comprehensive evaluation of the GPU cloud market, conducted in Q1 2025. It established baseline ratings for major providers and defined the evaluation methodology.
Licensing and partnerships
ClusterX is developed and published by SemiAnalysis.
Write to the team
Licensing inquiries and partnership proposals go to the ClusterX mailbox.
clusterx@semianalysis.com