Which GPU cloud platform fits your need?
Request a demo“I need to send my customer a correct invoice for what they used — this month, not next year.”
Most control planes stop at metering. You get usage records or a CSV, then you are expected to build rating, tax, credit notes, dunning and a payment gateway yourself.
What it costs you2–3 quarters of consumed GPU hours before the first correct invoice leaves the building.
Who can serve this — ranked
- #1Partial
Rafay
Real usage metering APIs plus a utility that emits a billing-ready CSV.
Where it stops: Invoicing is explicitly out of scope — their own engineering blog says so.
Sources - #2Does not cover it
NVIDIA (Run:ai / BCM)
Scheduling and cluster management.
Where it stops: Licenses per server to you. No tenant-facing invoicing at all.
www.nvidia.comwww.nvidia.com/en-us/software/run-aiwww.nvidia.com/en-us/data-center/base-command/managerSourcesdocs.nvidia.com/ai-enterprise/latest/product-support-matrix/index.htmlchecked 2026-08-23 - #3Does not cover it
Canonical
MAAS, Juju and Charmed Kubernetes with a support subscription.
Where it stops: No in-product customer invoicing; billing is deferred to you or a partner.
Sourcesubuntu.com/kubernetes/charmed-k8schecked 2026-08-23 - #4Does not cover it
Platform9
Managed private cloud and Kubernetes with a support subscription.
Where it stops: No tenant-facing invoicing; monetisation is left to your own billing stack.
Sourcesplatform9.comchecked 2026-08-23 - #5Does not cover it
Cast AI · Exostellar
Cost visibility and optimisation.
Where it stops: Shows spend, never produces an invoice — and meters you per CPU while doing it.
- #6Partial
OpenNebula
Built-in Showback: monthly per-VM, per-group cost reports with configurable CPU/memory/disk cost, exportable to a billing system.
Where it stops: Showback is a cost report, not an invoice — its own docs describe it as integrating with chargeback and billing platforms. No rating engine, tax, credit notes, dunning or payment gateway.
Sources - #7Does not cover it
k0rdent
Multi-tenant cluster provisioning with GPU metering in the k0rdent AI Starter Pack — 'provision, manage and meter GPU resources'.
Where it stops: Metering only. There is no customer-facing invoicing, no price list, no markup chain; monetisation is left to your own billing stack.
Sourceswww.mirantis.com/k0rdent-ai-pricingchecked 2026-08-23www.mirantis.com/software/mirantis-k0rdent-enterprisechecked 2026-08-23
And where Cognition-AI sits on this need
- #8Strong fit
Cognition-AIus
Hard cost → margin → unit price → downline markup, inside the product that provisions the resource.
Where it stops: Billable on day one; no integration project. Younger product than the US incumbents.
Sourcescognition-ai.comchecked 2026-08-23
“My partners must resell my cloud under their own brand, at their own price, with my margin preserved.”
White-label in this market usually means logo, colour and domain. It rarely means a parent→child pricing relationship the platform can actually enforce.
What it costs youPartner deals get priced in spreadsheets, so operators cap partners at what finance can reconcile by hand.
Who can serve this — ranked
- #1Does not cover it
OpenNebula · k0rdent
Groups, VDCs and multi-tenant clusters you can brand and hand to a partner.
Where it stops: Tenancy is an isolation boundary, not a commercial one. Neither carries a parent→child price list, so reseller markup is reconciled outside the platform.
Sourcesopennebula.iochecked 2026-08-23www.mirantis.com/software/mirantis-k0rdent-enterprisechecked 2026-08-23 - #2Partial
Rafay
OEM white-label: brand, colour, domain, tenant language. Three tiers: org → team → user.
Where it stops: No parent→child markup chain, so reseller margin lives outside the platform.
- #3Does not cover it
NVIDIA
Per-server licensing.
Where it stops: The licence model does not contemplate you reselling margin on top.
Sourcesdocs.nvidia.com/ai-enterprise/latest/product-support-matrix/index.htmlchecked 2026-08-23 - #4Does not cover it
Core42
A sovereign cloud you can resell capacity from.
Where it stops: Its billing faces its customers. You are in their channel, not running your own.
Sourcescore42.aichecked 2026-08-23
And where Cognition-AI sits on this need
- #5Strong fit
Cognition-AIus
Six tenancy tiers — super-admin → billing → sub-org → org → project → practitioner — with pricing at every level.
Where it stops: Requires you to actually define a channel price list; the platform will not invent one.
Sourcescognition-ai.comchecked 2026-08-23
“I have 8–128 GPUs, not 8,000. I need a real multi-tenant cloud at my size, without hyperscaler assumptions.”
Enterprise control planes are priced, staffed and scoped for large fleets. Below a certain size the licence and the platform team cost more than the GPUs earn.
What it costs youSmall fleets get pushed into DIY OpenStack — or out of the market entirely.
Who can serve this — ranked
- #1Partial
OpenNebula
Genuinely lightweight: a single control node can run a small estate, and the full platform is Apache-licensed with no per-GPU fee.
Where it stops: Cheap to license, expensive to staff. Enterprise/AI Factory subscriptions add support, but you still build the tenant-facing commercial layer yourself, and delivery is EU-centred.
- #2Partial
k0rdent
The k0rdent AI Starter Pack is explicitly a 'GPU Cloud in a Box' for first-time operators, supporting up to 144 Hopper/Blackwell GPUs.
Where it stops: Sized for NVIDIA Hopper/Blackwell estates and sold with enterprise subscription and services; below a rack the platform and support cost dominates the GPU revenue.
Sourceswww.mirantis.com/k0rdent-ai-pricingchecked 2026-08-23 - #3Partial
Rafay
Works technically at any size; sold to enterprises and neoclouds.
Where it stops: Reviewers flag licensing cost as high and a learning curve — both bite hardest on small fleets.
Sourceswww.peerspot.com/products/rafay-reviewschecked 2026-08-23 - #4Does not cover it
NVIDIA AI Enterprise
A complete AI stack on certified systems.
Where it stops: $2,500–$5,000 per GPU per year scales linearly downward too — it never gets cheap.
Sourcesblog.spheron.networkchecked 2026-08-23www.nvidia.com/en-us/data-center/products/ai-enterprisechecked 2026-08-23 - #5Does not cover it
HUMAIN · Core42 · Khazna
Hyperscale sovereign capacity.
Where it stops: They are the operator. A few racks of your own is not the business they are in.
And where Cognition-AI sits on this need
- #6Strong fit
Cognition-AIus
Designed for quarter-rack to few-rack estates: installed, metered and billable at that scale.
Where it stops: Not competing for 100,000-GPU national build-outs.
Sourcescognition-ai.comchecked 2026-08-23
“I want to buy whatever accelerator my procurement cycle and my tender allow — and switch next cycle.”
Vendor orchestration only manages the vendor's own silicon, and licences are counted per certified server. Hardware neutrality clauses in public tenders fail on the spot.
What it costs you$320K–$640K per year on a 128-GPU fleet, in a currency and a licence you do not control.
Who can serve this — ranked
- #1Strong fit
OpenNebula
Runs on commodity x86 and ARM, KVM/LXC/Firecracker, with NVIDIA PCI passthrough plus vGPU and MIG partitioning.
Where it stops: GPU documentation is NVIDIA-first; AMD and alternative accelerators work as generic PCI passthrough with no partitioning story.
Sourcesopennebula.iochecked 2026-08-23 - #2Partial
k0rdent
Infrastructure-agnostic cluster templates across public cloud, bare metal and edge.
Where it stops: The AI/GPU packaging is built around NVIDIA Hopper and Blackwell; non-NVIDIA silicon is not part of the productised path.
Sourceswww.mirantis.com/software/mirantis-k0rdent-enterprisechecked 2026-08-23www.mirantis.com/k0rdent-ai-pricingchecked 2026-08-23 - #3Strong fit
Rafay · Exostellar
Genuinely vendor-agnostic GPU and Kubernetes orchestration.
Where it stops: Neutral on hardware — but they leave IaaS, storage and invoicing to you.
- #4Strong fit
Canonical
Open-source stacks that run on commodity hardware.
Where it stops: No commercial layer on top.
Sourcesubuntu.com/kubernetes/charmed-k8schecked 2026-08-23 - #5Does not cover it
NVIDIA AI Enterprise / Base Command
The deepest NVIDIA feature set available.
Where it stops: NVIDIA-Certified Systems only. AMD, Intel and alternative accelerators are simply not covered.
Sourceswww.nvidia.com/en-us/data-center/products/ai-enterprisechecked 2026-08-23docs.nvidia.com/ai-enterprise/latest/product-support-matrix/index.htmlchecked 2026-08-23
And where Cognition-AI sits on this need
- #6Strong fit
Cognition-AIus
Hardware-agnostic by architecture; MIG-aware GPU slicing over open components.
Where it stops: You still buy your own vendor support contracts for the metal.
Sourcescognition-ai.comchecked 2026-08-23
“The control plane runs on my metal, in my jurisdiction, and the people who can escalate are in my time zone.”
Foreign SaaS in a local data centre passes the location test and fails the jurisdiction test. Escalation across nine time zones is a 10–14 hour round trip on a P1.
What it costs youClassified, defence and ministry workloads — the highest-margin tier in the region.
Who can serve this — ranked
- #1Strong fit
Core42
Air-gapped / restricted / connected tiers, cleared local personnel, 16 UAE data centres behind it.
Where it stops: Best answer if you want to rent sovereignty. You become their tenant, on their economics.
- #2Partial
Rafay via Moro Hub
Real UAE presence through the Oct 2025 Digital DEWA partnership.
Where it stops: Sunnyvale software, Sunnyvale roadmap and escalation. Channel-shaped sovereignty.
- #3Partial
OpenNebula
Fully open source under Apache 2.0, installable air-gapped, vendor based in Madrid — a strong European digital-sovereignty story with no US SaaS dependency.
Where it stops: Sovereignty is European. No MENA-resident engineering, no local install team, no Arabic-language operations, and support hours follow EU business time.
- #4Partial
k0rdent
Open-source core, deployable on-prem and at the edge for sovereign AI-factory builds.
Where it stops: US-headquartered vendor, now being acquired by IREN — a roadmap and ownership question you inherit, with escalation outside your time zone.
Sourceswww.mirantis.com/software/mirantis-k0rdent-enterprisechecked 2026-08-23www.mirantis.com/company/press-centerchecked 2026-08-23 - #5Does not cover it
Cast AI
Public-cloud-oriented SaaS optimisation.
Where it stops: Wrong shape for an air-gapped estate.
Sourcescast.aichecked 2026-08-23
And where Cognition-AI sits on this need
- #6Strong fit
Cognition-AIus
On-prem, air-gap ready, UAE-resident engineering, installed by the team that built it.
Where it stops: On-prem means you own the hardware lifecycle.
Sourcescognition-ai.comchecked 2026-08-23
“One platform, one identity model, one metering pipeline, one SLA — not a CMP plus a GPU orchestrator.”
The common architecture bolts Run:ai or Rafay onto VMware, OpenShift or Huawei Cloud Stack. Two control planes, two upgrade calendars, two support contracts.
What it costs youDuplicate licences, a second ops team, and hours of vendor finger-pointing inside a minutes-long SLA.
Who can serve this — ranked
- #1Partial
Canonical
Charmed OpenStack + Kubernetes + Ceph via Juju — the closest architectural match.
Where it stops: Sells charms and support. Your team becomes the platform team.
- #2Does not cover it
Rafay · Run:ai
Excellent GPU and Kubernetes scheduling.
Where it stops: No IaaS, no bare metal service, no storage. A layer, not a platform.
- #3Partial
OpenNebula
VMs, containers, bare-metal provisioning, storage and edge under one control plane and one identity model — architecturally the closest European equivalent.
Where it stops: Kubernetes arrives as an appliance add-on (OneKE / Rancher Prime), and there is no tenant invoicing — Showback stops at cost visibility.
- #4Partial
Mirantis (MKE · OpenStack)
OpenStack-on-Kubernetes and MKE from one vendor: containers, AI workloads and VMs with a service catalogue, GitOps and a Run:ai integration across clouds, edge and bare metal.
Where it stops: Everything is Kubernetes-native — VMs run as a workload on K8s rather than as first-class IaaS — and there is no tenant invoicing or regional install team.
Sourceswww.mirantis.comchecked 2026-08-23 - #5Partial
k0rdent
Open-source multi-cluster Kubernetes and AI platform with GPU provisioning, metering and a service catalogue.
Where it stops: Kubernetes-only: no first-class VM or bare-metal IaaS layer, no tenant invoicing, and no local install or support team.
Sourceswww.mirantis.com/software/mirantis-k0rdent-enterprisechecked 2026-08-23k0rdent.iochecked 2026-08-23
And where Cognition-AI sits on this need
- #6Strong fit
Cognition-AIus
IaaS, Kubernetes, storage and GPU from one control plane, one SLA, one team.
Where it stops: Phase-based rollout: IaaS → containers → GPU-PaaS.
Sourcescognition-ai.comchecked 2026-08-23
“First billable tenant this quarter. The hardware started losing value the day it landed.”
Run:ai reviewers report “very complex configuration and setup ... for running basic things.” OpenStack operators name upgrades as their primary pain. Platform9 buyers report long implementations.
What it costs youEvery month in integration is depreciation, power and cooling against zero revenue.
Who can serve this — ranked
- #1Partial
k0rdent
The AI Starter Pack is sold on exactly this promise: operationalise GPUs the same day the hardware arrives, with tenant isolation and GPU slicing pre-integrated.
Where it stops: Fast to schedule, not fast to invoice — metering ships, billing does not, so first tenant and first correct invoice are still two different dates.
Sourceswww.mirantis.com/k0rdent-ai-pricingchecked 2026-08-23 - #2Partial
OpenNebula
Small footprint and a genuinely quick install for a first cluster.
Where it stops: Day-one install is fast; the commercial layer, GPU partitioning policy and Kubernetes appliance integration are still your project.
Sourcesopennebula.iochecked 2026-08-23 - #3Partial
Platform9
Managed private cloud that emulates familiar VMware behaviours like HA and DRS.
Where it stops: Buyers report implementation takes time and integrations need ongoing maintenance.
- #4Partial
Rafay
Self-service delivery of GPU environments, including packaging Run:ai for consumption.
Where it stops: Fast to schedule, slow to monetise — the billing project still comes after.
Sourcesrafay.cochecked 2026-08-23 - #5Does not cover it
Raw OpenStack · Run:ai DIY
Maximum control, zero licence cost.
Where it stops: Complexity and upgrade burden are the documented reason projects slip past a quarter.
And where Cognition-AI sits on this need
- #6Strong fit
Cognition-AIus
Scoped install by the vendor's own engineers, with billing configured as part of go-live.
Where it stops: Deployment capacity is regional — MENA first.
Sourcescognition-ai.comchecked 2026-08-23
“We bought the DGX boxes as one company, but we are twenty divisions and a hundred projects. Give each one its own share — and a number finance can charge back.”
Direct-purchase clusters arrive as one flat pool. Allocation says the cluster is full while DCGM says the GPUs are at 15%; the loudest team wins the queue and the platform team eats the whole bill. Cast AI's 2026 report puts average Kubernetes GPU utilisation at ~5%, and 80%+ of enterprises surveyed by VentureBeat say their GPUs run at half capacity or less.
What it costs youAn idle H100 is roughly $8,850 a month of capacity nobody is charged for — and no division has any reason to release it.
Who can serve this — ranked
- #1Partial
NVIDIA (Run:ai / BCM)
Genuinely strong at the scheduling half: projects, departments, fair-share quotas, guaranteed and over-quota GPU allocation, MIG and fractional GPUs on certified NVIDIA systems.
Where it stops: Quota is not accounting. There is no internal price list, no per-division statement, no cost model finance can reconcile — and it only governs NVIDIA-certified hardware.
www.nvidia.comwww.nvidia.com/en-us/software/run-aiwww.nvidia.com/en-us/data-center/base-command/managerSourceswww.nvidia.com/en-us/software/run-aichecked 2026-08-23www.nvidia.com/en-us/data-center/base-command/managerchecked 2026-08-23www.nvidia.com/en-us/data-center/products/ai-enterprisechecked 2026-08-23 - #2Partial
Cast AI · Saturn Cloud · FinOps tooling
DCGM-based per-pod, per-namespace attribution, idle detection, and reclamation of GPUs held by dormant notebooks — good visibility into who is wasting what.
Where it stops: It reports on a cluster somebody else provisions. Cost attribution is not a chargeback statement, and these tools do not own the allocation, the tenancy boundary or the invoice.
Sourcescast.ai/blog/gpu-cost-monitoring-kuberneteschecked 2026-08-23saturncloud.io/services/gpu-cost-and-chargebackchecked 2026-08-23cast.ai/reports/kubernetes-optimization-reportchecked 2026-08-23 - #3Partial
OpenNebula
Native user groups and VDCs, plus Showback: monthly per-VM, per-group cost reports with configurable rates.
Where it stops: Showback is a report, not a chargeback engine — its own docs describe integrating with a billing platform. GPU-hour and MIG-slice rating is not modelled.
Sources - #4Partial
Rafay · k0rdent
Multi-tenant environments and namespaces per team, with usage metering behind them.
Where it stops: Metering stops at a CSV. Turning that CSV into a defensible internal chargeback across twenty cost centres is your project.
Sourceswww.mirantis.com/k0rdent-ai-pricingchecked 2026-08-23 - #5Does not cover it
Spreadsheets and goodwill
What most direct-purchase clusters actually run on today.
Where it stops: When scarce GPU time appears free, demand grows faster than governance — documented as the standard failure mode of shared on-prem AI platforms.
Sourcessysart.consulting/insights/gpu-chargeback-quotas-on-prem-ai-platformschecked 2026-08-23sysart.consulting/insights/qos-fairness-shared-gpu-inference-on-premiseschecked 2026-08-23
And where Cognition-AI sits on this need
- #6Strong fit
Cognition-AIus
Divisions, projects and subsidiaries as real tenants with their own quotas, MIG/vGPU slices, and an internal price list — the same rating engine that bills an external customer produces the internal chargeback or showback statement per division.
Where it stops: You still have to agree the policy: who pays for reserved-but-idle capacity. We will help you set it, we cannot decide it for you.
Sourcescognition-ai.comchecked 2026-08-23
“I want to sell my GPUs plus somebody else's inference API, storage tier or software — one catalogue, one markup, one invoice to my downline.”
Buyers in 2026 are not comparing H100 hours, they are comparing tokens, latency and platform primitives. Operators are told to move up the stack — but the control planes they run only know how to provision their own resources. Anything a partner supplies lives outside the system, so it gets priced in a spreadsheet and invoiced separately.
What it costs youEvery partner service you cannot rate and mark up is margin that stays with the partner instead of your P&L.
Who can serve this — ranked
- #1Partial
hosted.ai
Turnkey neocloud stack with customisable add-ons, white-label UI, automated billing and self-service provisioning across bare metal, VMs and Kubernetes.
Where it stops: The add-on model is built around their own service flavours; it is a monetisation stack for your infrastructure rather than an open marketplace of external providers you rate and resell.
Sourceshosted.ai/platformchecked 2026-08-23 - #2Partial
ARK Labs · Saturn Cloud
Move the operator up the stack: managed multi-tenant inference, token endpoints and a customer-facing portal on top of a mixed fleet.
Where it stops: One layer, sold per GPU. It adds a product to your catalogue — it does not become the catalogue, and it does not carry your partners' services or your reseller price chain.
Sourcesark-labs.cloud/solutions/for-neocloudschecked 2026-08-23saturncloud.io/docs/operatorschecked 2026-08-23 - #3Partial
Rafay
A genuine service catalogue: self-service environments and packaged third-party software (including Run:ai) delivered to tenants, plus an AI Token Factory for inference endpoints.
Where it stops: The catalogue delivers software; it does not price it. No cost basis, no markup chain, no invoice — monetisation is left to a billing system you supply.
- #4Does not cover it
NVIDIA stack
A deep first-party catalogue — NIM, NeMo, AI Enterprise — on certified NVIDIA systems.
Where it stops: It is their catalogue, licensed per GPU. You resell NVIDIA's list, you do not compose your own with your own margin on top.
Sourceswww.nvidia.com/en-us/data-center/products/ai-enterprisechecked 2026-08-23blog.spheron.networkchecked 2026-08-23 - #5Does not cover it
k0rdent
Cluster templates and a catalogue of services for what runs on your own infrastructure.
Where it stops: A technical catalogue with no commercial layer. Nothing rates a partner's service, applies a markup, or produces a downline price.
Sourceswww.mirantis.com/software/mirantis-k0rdent-enterprisechecked 2026-08-23 - #6Does not cover it
OpenNebula
Appliance marketplace for VM and container images on your own cloud.
Where it stops: Images, not offers. No third-party service onboarding, markup rules, or reseller price list.
Sourcesopennebula.iochecked 2026-08-23 - #7Does not cover it
Platform9
Managed private cloud with a curated add-on set on top of your hardware.
Where it stops: Add-ons are operational, not commercial. Nothing turns a partner service into a rated, marked-up SKU.
Sourcesplatform9.comchecked 2026-08-23
And where Cognition-AI sits on this need
- #8Strong fit
Cognition-AIus
Third-party and partner services are onboarded as catalogue items alongside your own compute: we load them, set the cost basis, apply your margin and the downline reseller markup, and they bill on the same invoice as GPU hours.
Where it stops: Each new provider is an onboarding exercise — connectors and metering hooks are configured with you, not self-serve on day one.
Sourcescognition-ai.comchecked 2026-08-23
“I put capital into this facility and this hardware. Tell me the highest-return way to run it — and show me the numbers monthly.”
Capital arrived before an operating model. The facility was underwritten as real estate; the GPUs inside it depreciate like semiconductors. Colocation owners moving into GPUaaS are taking on technology-obsolescence risk their business model never carried, and most cannot yet see revenue per GPU, per rack or per megawatt in one place.
What it costs youThe asset only pays back if it is sold as a service. Racks that sit as unsold capacity are financed depreciation with no offsetting revenue.
Who can serve this — ranked
- #1Partial
Advisory and capital firms
Frameworks for monetising stranded capacity, underwriting GPU density, and structuring residual-value cover on GPUs as a financeable asset class.
Where it stops: Strategy and structuring. Nobody hands you a running platform at the end of it.
Sourcesankuracapitaladvisors.com/insights/from-cost-center-to-cash-machinechecked 2026-08-23www.globaldatacenterhub.com/p/how-to-underwrite-gpu-density-inchecked 2026-08-23www.datacenterdynamics.com/en/opinions/colos-are-buying-into-gpuschecked 2026-08-23 - #2Partial
DCIM (Modius and peers)
Unifies power, cooling and IT telemetry to prove where stranded capacity hides and turn recovered megawatts into sellable inventory.
Where it stops: Facility-side truth only. It never sees GPU-hours, tenants or revenue.
Sources - #3Partial
Rafay
A mature control plane that makes the hardware consumable quickly.
Where it stops: Answers 'can workloads run?' — not 'what is this asset earning?'. Utilisation dashboards, not a P&L per rack.
Sourcesrafay.cochecked 2026-08-23 - #4Partial
k0rdent
Fast multi-cluster standup across the fleet.
Where it stops: Cluster-centric telemetry. No revenue, cost or payback view tied to the physical asset.
Sourceswww.mirantis.com/k0rdent-ai-pricingchecked 2026-08-23 - #5Partial
Platform9
Managed operations that keep the estate running without a large platform team.
Where it stops: Ops SLAs, not asset economics. Nothing reports yield per rack or per invested dollar.
Sourcesplatform9.comchecked 2026-08-23 - #6Partial
Sell or lease the capacity wholesale
Fastest path to predictable cash: hand the racks to a single large tenant or an operator.
Where it stops: You cap the return at wholesale rates and give the retail margin, the customer relationship and the upside to someone else.
And where Cognition-AI sits on this need
- #7Strong fit
Cognition-AIus
Turns the asset into a sellable product: tenants, price list, margin per SKU and revenue per GPU and per rack reported from the same system that provisions it — so the operating model and the investor reporting are one thing, not two.
Where it stops: We supply the platform and the commercial model, not the demand. Filling the racks is still a go-to-market job.
Sourcescognition-ai.comchecked 2026-08-23
“My power is fixed and expensive. I want more billable output from the same megawatts — and I will pay for whatever gets me there.”
Power, not floor space, is the ceiling. LBNL-based analysis puts 30–50% of installed capacity in US facilities as unused, while inside the racks GPUs sit at a fraction of their capability: ~5% average utilisation in Kubernetes clusters, 30–40% in typical enterprise fleets. You pay for the watts either way; only sold, busy GPUs earn against them.
What it costs youTwo operators on the same megawatts can differ by multiples in revenue. The difference is scheduling, slicing, reclamation and pricing — not more hardware.
Who can serve this — ranked
- #1Partial
NVIDIA (DSX MaxLPS, Dynamo, Run:ai)
Serious performance-per-watt engineering: dynamic power redistribution across racks, disaggregated inference to recover stranded compute under thermal limits, and fractional GPU scheduling.
Where it stops: Tied to NVIDIA reference systems and certified silicon, and it optimises throughput — not the price you charge for it. Efficiency without a billing layer is cost saving, not revenue.
Sourcesdeveloper.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlpschecked 2026-08-23perspectives.nvidia.com/ai-infrastructure/total-cost-of-ownership/task/faq/recover-stranded-gpu-capacity-thermal-power-constraintschecked 2026-08-23www.nvidia.com/en-us/software/run-aichecked 2026-08-23 - #2Partial
Cast AI · Exostellar · HAMi
Bin-packing, GPU virtualisation and idle reclamation that measurably raise utilisation on clusters you already run.
Where it stops: Kubernetes-scoped optimisation layers. They lift utilisation, then hand the commercial question — who is charged for the recovered capacity — back to you.
Sourcescast.aichecked 2026-08-23www.businesswire.comchecked 2026-08-23project-hami.iochecked 2026-08-23 - #3Partial
DCIM and power-side tooling
Finds stranded power across the facility and converts recovered megawatts into recurring revenue per megawatt.
Where it stops: Stops at the rack PDU. It cannot see whether the GPUs drawing that power are earning.
Sources - #4Does not cover it
Buy more GPUs
The reflex answer when capacity feels short.
Where it stops: In a power-constrained site there is nothing to plug in. Under strict thermal and power limits, recovery is a software problem — NVIDIA says so itself.
And where Cognition-AI sits on this need
- #5Strong fit
Cognition-AIus
Packs paying tenants onto the same silicon with MIG and vGPU slicing, reclaims idle allocations back into sellable inventory, and prices each slice — so higher utilisation shows up as invoiced revenue against the same power draw, not just a greener dashboard.
Where it stops: We optimise how the GPUs are sold and shared. Facility power, cooling design and PUE remain the data centre's own engineering.
Sourcescognition-ai.comchecked 2026-08-23
“I just want sovereign GPU capacity on a contract. I have no interest in owning a platform.”
Nothing is broken here — this buyer should not be sold an orchestration platform at all.
What it costs youBuying a platform you will not operate is the most expensive mistake on this page.
Who can serve this — ranked
- #1Strong fit
Core42 · HUMAIN · Ezditek × Gcore
National-scale sovereign capacity you rent, with cleared staff and regional data centres.
Where it stops: You rent margin as well as compute.
And where Cognition-AI sits on this need
- #2Does not cover it
Cognition-AIus
A platform for operators who own their metal and bill their own tenants.
Where it stops: Wrong fit if you never intend to run infrastructure. We will say so on the call.
Sourcescognition-ai.comchecked 2026-08-23
Six things the bigger platforms will not do for you.
On market cap and years-in-market, several vendors on this page are ahead of us. None of them will do the six things below — not because they cannot, but because their model does not allow it at your size.
- 01
Data sovereignty, with local experts on the ground
By default, Cognition-AI is the platform for organizations where data sovereignty is the top priority. We run entirely inside your perimeter, with no external control plane, no telemetry dependency, and no foreign SaaS for your compliance team to approve. And because our engineers are physically present, you get on-site human support that understands your rack, your policy, and your regulator.
You keep data inside the organization, and you have a local team standing next to the hardware when it matters.
- 02
Turnkey installation, with people on the ground
We have a physical presence in the region and we install it ourselves: racking, MAAS/Juju bring-up, Ceph, network, GPU enablement, first tenant live. You are not handed a licence, a PDF and a Slack channel in another timezone.
No integrator contract, no 3–6 month deployment project billed by the hour.
- 03
Human-led support, not ticket triage
When a client gets stuck, our engineers go into their environment, reproduce the issue and fix it. Support is done by the team that built the product, not by a first-line queue reading a runbook.
Hours of downtime instead of days — on hardware that costs you money every idle hour.
- 04
Customisation, because you are not too small to matter
Odd accelerator mix, a tenant model nobody else has, an internal portal to integrate, a report your CFO insists on — we build it. The largest platforms are too big to care about a request from a quarter-rack or few-rack operator.
You stop paying for workarounds and shadow tooling around a product that will not bend.
- 05
GPU and CPU in one platform
Our depth in any single niche may not beat a specialist, but we cover the whole surface an operator actually runs. GPU compute and CPU compute, bare metal and VMs, tenants and quotas — one control plane.
One platform instead of two or three licences, two or three teams and the glue between them.
- 06
Pricing and policy engineering for real organisations
This is where most of our engineering has gone: price lists, markups, reseller tiers, internal chargeback rates, approval workflows, quota and policy rules that map to how your organisation is actually governed.
Your finance and compliance teams sign off on day one instead of blocking launch.
Four kinds of vendor. Each solves a different slice.
If you would rather scan the market by category than by need: here is who is out there, what they sell, and where each category stops.
Hardware-vendor orchestration
Software that only manages the vendor's own silicon.
WhoNVIDIA Run:ai, Base Command / DGX, AMD ROCm, Intel
Where it stopsCertified hardware only. You buy their GPUs or you get nothing.
GPU / Kubernetes control plane
GPU + K8s scheduling on any hardware, sold to enterprises and neoclouds.
WhoRafay, Mirantis k0rdent, Exostellar, Cast AI, Saturn Cloud, ARK Labs, WhaleFlux, LayerOps
Where it stopsNo IaaS, no storage, no invoicing. A layer, not a business.
Multi-stack private cloud platform
IaaS + containers + storage unified for sovereign and private use.
WhoOpenNebula, Mirantis, Canonical, Platform9, hosted.ai, Huawei Cloud Stack, Red Hat
Where it stopsStrong stacks, zero customer-facing billing, no regional install team.
Regional sovereign operators
They run the cloud. You rent from them.
WhoCore42 / Khazna, e& OneCloud, Telnyx Dubai, HUMAIN, Ezditek × Gcore
Where it stopsYou become their tenant — not the operator of your own margin.
Tell us your need. We will tell you honestly if it is us.
Ninety minutes with our engineers: your rack design, your accelerator mix, your tenant model and your margin structure — mapped against what you are being quoted today.
Every source on this page.
Vendor documentation, press releases, analyst reports and public review sites, as retrieved in August 2026. Pricing and partnership details change — verify before quoting.
- 01cognition-ai.comhttps://cognition-ai.com · checked 2026-08-23
- 02rafay.cohttps://rafay.co · checked 2026-08-23
- 03rafay.co — Rafay for AIhttps://rafay.co/rafay-for-ai/ · checked 2026-08-23
- 04Rafay blog — GPU cloud billing: from usage metering to billinghttps://rafay.co/the-kubernetes-current/gpu-cloud-billing-from-usage-metering-to-billing/ · checked 2026-08-23
- 05Rafay — AI Token Factoryhttps://rafay.co/platform/ai-token-factory/ · checked 2026-08-23
- 06Rafay press — Moro Hub (Digital DEWA) partnership, Oct 2025https://rafay.co/company/press-releases/ · checked 2026-08-23
- 07Moro Hub (Digital DEWA)https://www.morohub.com/ · checked 2026-08-23
- 08Rafay Partner Elevate programhttps://rafay.co/partners/ · checked 2026-08-23
- 09PeerSpot — Rafay reviewshttps://www.peerspot.com/products/rafay-reviews · checked 2026-08-23
- 10G2 — Run:ai reviewshttps://www.g2.com/products/run-ai/reviews · checked 2026-08-23
- 11NVIDIA — Run:aihttps://www.nvidia.com/en-us/software/run-ai/ · checked 2026-08-23
- 12NVIDIA — NVIDIA AI Enterprise licensing / certified systemshttps://www.nvidia.com/en-us/data-center/products/ai-enterprise/ · checked 2026-08-23
- 13NVIDIA AI Enterprise documentation — licensing prerequisiteshttps://docs.nvidia.com/ai-enterprise/latest/product-support-matrix/index.html · checked 2026-08-23
- 14NVIDIA Base Command Managerhttps://www.nvidia.com/en-us/data-center/base-command/manager/ · checked 2026-08-23
- 15Spheron — NVIDIA AI Enterprise pricing analysis ($2,500–$5,000/GPU/yr)https://blog.spheron.network/ · checked 2026-08-23
- 16r/HPC — Bright Cluster Manager repricing ($260 → $4,500 per node)https://www.reddit.com/r/HPC/ · checked 2026-08-23
- 17The Next Platform — NVIDIA, Bright and cluster managementhttps://www.nextplatform.com/ · checked 2026-08-23
- 18NVIDIA acquires SchedMD (Slurm), Dec 2025https://www.linkedin.com/company/nvidia/ · checked 2026-08-23
- 19Reuters — NVIDIA completes Run:ai acquisition (~$700M)https://www.reuters.com/technology/nvidia-completes-acquisition-israeli-ai-firm-runai-2024-12-30/ · checked 2026-08-23
- 20Tom's Hardware — NVIDIA to open-source Run:aihttps://www.tomshardware.com/tech-industry/artificial-intelligence · checked 2026-08-23
- 21Mirantishttps://www.mirantis.com/ · checked 2026-08-23
- 22Mirantis blog — k0rdent Enterprisehttps://www.mirantis.com/blog/ · checked 2026-08-23
- 23k0rdent (open source)https://k0rdent.io/ · checked 2026-08-23
- 24Mirantis k0rdent Enterprise — Kubernetes platformhttps://www.mirantis.com/software/mirantis-k0rdent-enterprise/ · checked 2026-08-23
- 25Mirantis k0rdent AI Starter Pack — 'GPU Cloud in a Box', up to 144 GPUshttps://www.mirantis.com/k0rdent-ai-pricing/ · checked 2026-08-23
- 26Mirantis — IREN acquisition noticehttps://www.mirantis.com/company/press-center/ · checked 2026-08-23
- 27OpenNebula — open source cloud & AI platformhttps://opennebula.io/ · checked 2026-08-23
- 28OpenNebula — subscriptions & Enterprise/AI Factory editionshttps://opennebula.io/subscriptions/ · checked 2026-08-23
- 29OpenNebula docs — Showback (usage cost reports, integrates with billing platforms)https://docs.opennebula.io/7.4/product/cloud_system_administration/multitenancy/showback/ · checked 2026-08-23
- 30OpenNebula docs — NVIDIA vGPU & MIG, GPU passthroughhttps://docs.opennebula.io/7.4/product/cluster_configuration/hosts_and_clusters/vgpu/ · checked 2026-08-23
- 31Canonical — Charmed Kuberneteshttps://ubuntu.com/kubernetes/charmed-k8s · checked 2026-08-23
- 32Canonical — Cephhttps://canonical.com/ceph · checked 2026-08-23
- 33Platform9https://platform9.com/ · checked 2026-08-23
- 34Platform9 blog — VMware alternativeshttps://platform9.com/blog/ · checked 2026-08-23
- 35Independent vendor analysis — Platform9https://getbreakout.ai/ · checked 2026-08-23
- 36Core42 — Signature Private Cloudhttps://core42.ai/ · checked 2026-08-23
- 37Core42 — $550M HSBC trade finance (Feb + May 2026)https://core42.ai/newsroom · checked 2026-08-23
- 38G42https://g42.ai/ · checked 2026-08-23
- 39Khazna Data Centershttps://khaznadatacenters.com/ · checked 2026-08-23
- 40DataCenterDynamics — Khazna UAE footprinthttps://www.datacenterdynamics.com/ · checked 2026-08-23
- 41e& enterprise — OneCloud with Oracle Alloyhttps://www.eand.com/ · checked 2026-08-23
- 42Telnyx — Dubai GPU infrastructurehttps://telnyx.com/ · checked 2026-08-23
- 43Go Data — UAE GPU cloudhttps://godataglobal.com/ · checked 2026-08-23
- 44HUMAIN (PIF, Saudi Arabia)https://humain.ai/ · checked 2026-08-23
- 45NVIDIA newsroom — HUMAIN 18,000 GB300 deploymenthttps://nvidianews.nvidia.com/ · checked 2026-08-23
- 46Vision 2030 — Saudi AI programmehttps://www.vision2030.gov.sa/ · checked 2026-08-23
- 47Public Investment Fund — HUMAINhttps://www.pif.gov.sa/ · checked 2026-08-23
- 48Ezditek × Gcore — nine Saudi AI data centershttps://ezditek.com/ · checked 2026-08-23
- 49DCNN Magazine — Ezditek and Gcorehttps://dcnnmagazine.com/ · checked 2026-08-23
- 50S&P Global — Saudi data center market CAGRhttps://www.spglobal.com/ · checked 2026-08-23
- 51Cast AIhttps://cast.ai/ · checked 2026-08-23
- 52nOps — Cast AI cost analysishttps://www.nops.io/blog/ · checked 2026-08-23
- 53SoftwareFinder — Cast AI pricinghttps://softwarefinder.com/ · checked 2026-08-23
- 54BusinessWire — Exostellar Software Defined GPUhttps://www.businesswire.com/ · checked 2026-08-23
- 55Startup Stash — Exostellar listinghttps://startupstash.com/ · checked 2026-08-23
- 56HAMi — CNCF incubating GPU virtualizationhttps://project-hami.io/ · checked 2026-08-23
- 57Grand View Research — sovereign cloud market ($117.5B 2025 → $648.9B 2033)https://www.grandviewresearch.com/ · checked 2026-08-23
- 58Fortune Business Insights — sovereign cloud forecast (27.0% CAGR)https://www.fortunebusinessinsights.com/ · checked 2026-08-23
- 59IDC — sovereign AI stack split forecasthttps://blogs.idc.com/ · checked 2026-08-23
- 60r/vmware — Broadcom price-increase threadshttps://www.reddit.com/r/vmware/ · checked 2026-08-23
- 61OpenStack — operator upgrade pain pointshttps://www.openstack.org/blog/ · checked 2026-08-23
- 62OpenMetal — OpenStack operational complexityhttps://openmetal.io/ · checked 2026-08-23
- 63OpenStack Foundation — c12n reference stack (OpenStack + K8s + Ceph)https://www.openstack.org/ · checked 2026-08-23
- 64Cast AI — 2026 State of Kubernetes Optimization Report (23,000+ clusters)https://cast.ai/reports/kubernetes-optimization-report/ · checked 2026-08-23
- 65Cast AI — GPU cost monitoring: average GPU utilisation ~5%, idle H100 ≈ $8,850/GPU/monthhttps://cast.ai/blog/gpu-cost-monitoring-kubernetes/ · checked 2026-08-23
- 66VentureBeat Research, July 2026 — 80%+ of enterprises say GPUs run at half capacity or lesshttps://venturebeat.com/orchestration/wall-street-is-debating-the-ai-buildout-enterprises-just-answered-86-say-their-gpus-run-at-half-capacity-or-less · checked 2026-08-23
- 67INFINITIX — ROI of GPU idle time (most enterprise clusters average 30–40% utilisation)https://ai-stack.ai/en/gpu-roi · checked 2026-08-23
- 68Cloud Cost Room — allocating shared GPU cluster costs across teamshttps://cloudcostroom.com/blog/how-to-allocate-shared-gpu-cluster-costs-across-teams · checked 2026-08-23
- 69SysArt — GPU chargeback and quotas for shared on-prem AI platformshttps://sysart.consulting/insights/gpu-chargeback-quotas-on-prem-ai-platforms/ · checked 2026-08-23
- 70SysArt — QoS and fairness for shared on-prem GPU inference clustershttps://sysart.consulting/insights/qos-fairness-shared-gpu-inference-on-premises/ · checked 2026-08-23
- 71Saturn Cloud — GPU chargeback: 'allocation says full, DCGM says 15%'https://saturncloud.io/services/gpu-cost-and-chargeback/ · checked 2026-08-23
- 72Saturn Cloud — platform layer for GPU operators and neocloudshttps://saturncloud.io/docs/operators/ · checked 2026-08-23
- 73hosted.ai — turnkey neocloud stack: overcommit, white-label UI, billinghttps://hosted.ai/platform/ · checked 2026-08-23
- 74ARK Labs — 'stop selling GPUs, start selling inference' (neocloud inference layer)https://ark-labs.cloud/solutions/for-neoclouds/ · checked 2026-08-23
- 75Ankura Capital Advisors — monetising stranded enterprise data center capacityhttps://ankuracapitaladvisors.com/insights/from-cost-center-to-cash-machine/ · checked 2026-08-23
- 76Modius — how colocation providers find and reclaim stranded power capacityhttps://modius.com/blog/how-colocation-providers-find-and-reclaim-stranded-power-capacity/ · checked 2026-08-23
- 77DataCenterDynamics — colos buying into GPUs take on technology obsolescence riskhttps://www.datacenterdynamics.com/en/opinions/colos-are-buying-into-gpus/ · checked 2026-08-23
- 78Global Data Center Hub — how to underwrite GPU density in AI data centershttps://www.globaldatacenterhub.com/p/how-to-underwrite-gpu-density-in · checked 2026-08-23
- 79NVIDIA — recovering stranded GPU capacity under thermal and power constraintshttps://perspectives.nvidia.com/ai-infrastructure/total-cost-of-ownership/task/faq/recover-stranded-gpu-capacity-thermal-power-constraints/ · checked 2026-08-23
- 80NVIDIA — maximising AI factory performance per watt with DSX MaxLPShttps://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/ · checked 2026-08-23
- 81Hammerhead / LBNL — 30–50% of installed data center power capacity sits unusedhttps://hammerheadco.ai/the-sleeping-giant-tapping-into-the-hidden-power-of-ai-data-centers/ · checked 2026-08-23
- 82Cognizant — GPU segmentation and multi-tenancy for cost efficiency (MIG, fractional GPU)https://www.cognizant.com/en_us/services/documents/optimizing-resource-utilization-and-reducing-costs-through-gpu-segmentation.pdf · checked 2026-08-23