← Insights
GPU Economics10 min

Sharon AI’s $356m facility puts billable GPU-hours first

GPU-backed finance does not make idle capacity productive. A worked operating worksheet shows how smaller estates can connect reservations, metering, power and tenancy to collected revenue.

GPU server racks beside a reservation ledger and electricity meter, illustrating the conversion of installed capacity into billable usage.

A financing commitment is not a utilisation result

Sharon AI’s October 1, 2026 announcement describes a committed $356 million GPU-backed debt facility carrying fixed interest of 9.95%, excluding fees. The change is access to a financing framework for GPU infrastructure deployments, not evidence that every financed accelerator is installed, serving customers or generating cash. Operators should keep those milestones separate.

The announcement appears through Sharon AI’s press release and an SEC-hosted exhibit. Those are publication routes for the company’s disclosure, not two independent investigations. Dealroom is an aggregation signal rather than independent confirmation. Repetition across these channels should not increase confidence in assumptions about drawdowns, customer commitments or achieved utilisation that the disclosure does not establish.

A committed facility is also different from drawn debt. Conditions, collateral eligibility, deployment schedules and other contractual provisions can govern when funds become available. The headline interest rate excludes fees, so it is not an all-in borrowing cost. Applying 9.95% to the entire $356 million would imply $35.422 million in annual interest only if that whole amount were outstanding for a full year on those terms.

For a quarter-rack operator, the transferable lesson is not the facility’s size. It is the need to connect financeable assets to enforceable customer obligations and reliable service delivery. Installed GPU count measures inventory. Billable GPU-hours, collection timing and contribution after operating costs determine whether that inventory supports its financing.

The bottleneck is the revenue conversion chain

A small estate encounters the same conversion problem without a large finance team. Hardware must move through acceptance, provisioning, available capacity, customer allocation, successful service delivery, metering, invoicing and collection. Every transition can strand value. A healthy GPU with no contracted tenant produces no revenue; a busy GPU with disputed usage records may produce an unpaid invoice.

Start by maintaining three separate utilisation measures. Physical utilisation describes activity on the accelerator. Allocation utilisation describes capacity assigned to tenants. Billable utilisation describes capacity that the contract permits you to invoice. None automatically equals the others, especially when reservations, minimum commitments and fractional GPUs coexist.

A reservation can be billable while the customer is not executing work, provided the agreement charges for exclusive availability. Conversely, high device activity can come from internal testing, failed workloads, free trials or unmetered sessions. Reporting either condition as simply “80% utilised” hides the information needed for pricing and debt service.

The operating worksheet should therefore join tenant identity, contract, allocation, service interval, meter and invoice line. This is the practical role for orchestration: make commercial boundaries correspond to enforceable infrastructure boundaries, then retain enough evidence to explain each charge.

Build a capacity ledger before a revenue forecast

Define the saleable unit first. It might be an exclusive whole GPU-hour, a named partition-hour, or a VM-hour containing a specified accelerator allocation. Do not assume that fractional resources have interchangeable performance. Memory capacity, bandwidth, isolation and workload compatibility can make apparently similar slices commercially different.

Next, declare the reporting period and service calendar. Deduct planned maintenance, acceptance testing and known unavailable capacity before making commitments. Keep unplanned downtime visible rather than retrospectively deleting it from the denominator. Otherwise, reliability deterioration can misleadingly improve your reported utilisation percentage.

Use a ledger such as this one, reconciling each row monthly:

Ledger itemDefinitionEvidence
Installed hoursAccepted GPUs × calendar hoursAsset inventory
Available hoursInstalled hours less planned exclusionsMaintenance calendar
Reserved hoursCapacity promised under contractReservation schedule
Delivered hoursTenant allocation actually servedAllocation events
Billable hoursContract-qualified charging unitsMeter and tariff rules
Collected revenueCash received against invoicesReceipts ledger

Specify whether “delivered” means allocation time or measured execution time. Allocation-time billing is often easier to audit, but customers must understand that an idle allocated device can remain chargeable. Execution telemetry is useful for efficiency analysis; it should not silently replace the agreed billing basis.

A worked worksheet for 32 GPUs

Consider an illustrative estate of 32 identical whole-GPU billing units over a 30-day month. These are planning assumptions, not Cognition-AI or Clastiq Orchestrator benchmarks. There are 720 calendar hours per GPU, giving 23,040 installed GPU-hours. Planned maintenance removes 24 hours per device, or 768 GPU-hours, leaving 22,272 available GPU-hours.

Customers reserve 12,000 of those available hours under availability-based contracts and actually allocate 9,000. Other customers purchase 4,000 on-demand allocation-hours outside the reserved blocks. Assume no service failures, credits or overlaps. Delivered allocation totals 13,000 GPU-hours, while billable capacity totals 16,000 GPU-hours because the unused 3,000 reserved hours remain chargeable.

Billable utilisation against installed capacity is 16,000 ÷ 23,040, or 69.4%. Delivered allocation utilisation against installed capacity is 13,000 ÷ 23,040, or 56.4%. Against available capacity, delivered allocation utilisation is 58.4%. These percentages answer different questions; the dashboard should display their labels and denominators.

The reservation schedule matters as much as the monthly total. A customer cannot reserve more GPUs at a particular moment than the estate can supply. Nor can the operator sell unused reserved capacity twice unless the contract explicitly allows reclaimable service and the scheduler can honour the reservation when needed.

Assume the following monthly economic costs:

  • Hardware depreciation: $24,000.
  • Electricity: $7,200.
  • Operations and support allocation: $10,000.
  • Network and facility charges, excluding electricity: $4,000.
  • Storage depreciation and operating allocation: $4,000.

Total cost is $49,200. The electricity assumption is a constant average facility-attributable draw of 50 kW across the month: 50 × 720 = 36,000 kWh, priced at $0.20/kWh. This boundary includes the estate’s allocated cooling overhead; adding cooling again would double-count it.

Cost per available GPU-hour is $2.21, rounded from $49,200 ÷ 22,272. Cost per billable GPU-hour is $3.075. Dividing by installed hours instead produces $2.14, a reassuring but commercially misleading figure if used as the minimum sustainable selling price.

At $3.20 per reserved GPU-hour, reservations generate $38,400. At $4.00 per on-demand GPU-hour, on-demand sales generate $16,000. Total revenue is $54,400, leaving $5,200 before financing costs and tax. The blended realised price is $3.40 per billable GPU-hour.

This model is sensitive to lost demand. Losing 2,000 on-demand hours removes $8,000 of revenue. Even if electricity falls somewhat, depreciation, staffing and facility commitments do not disappear. With costs unchanged for a conservative first pass, the $5,200 surplus becomes a $2,800 deficit. More installed GPUs would not fix that demand gap.

Separate economic margin from debt-service cash

Now add a hypothetical $1 million outstanding loan at 9.95%, using simple annual interest divided by twelve for illustration. Monthly interest is approximately $8,292, before fees. The worksheet’s $5,200 surplus becomes a roughly $3,092 loss before tax after interest. This is a sensitivity test using the announced rate, not a representation of Sharon AI’s borrowing schedule.

Do not then subtract loan principal from that accounting result and call it profit. Depreciation is non-cash; principal repayment is a financing cash outflow. Maintain a separate cash bridge that starts with receipts, subtracts cash operating expenses, interest, fees, principal and capital expenditure, and reflects customer payment delays.

Likewise, do not count the same GPU through both depreciation and a full equipment lease payment without checking the accounting treatment. For operators, the immediate objective is simpler than producing statutory accounts: know the economic replacement cost and the actual payment calendar, without mixing them into an unusable margin number.

Reservations help only when the customer obligation is dependable. Examine cancellation rights, acceptance conditions, service credits, payment terms and counterparty exposure. A non-binding forecast is not contracted revenue. A signed commitment is not collected cash. A reservation from one customer can improve utilisation while making the estate more vulnerable to a single default.

Turn reservations into enforceable tenancy

A system such as Clastiq Orchestrator addresses the infrastructure side of this chain through bare-metal provisioning, Kubernetes and VM tenancy, GPU partitioning, quotas, policy engineering and per-tenant metering and chargeback. It does not manufacture customer demand or make a weak contract bankable. Its role is to make sold capacity governable on the operator’s own hardware.

Provisioning should produce a known node state before capacity becomes saleable. Record hardware identity, accelerator inventory, firmware and driver baselines, storage attachments and acceptance results. A device still undergoing burn-in belongs in inventory, not in the reservation pool used to support service commitments.

Tenancy then binds capacity to an accountable customer. Kubernetes namespaces alone are not a complete isolation policy. Define which workloads require dedicated nodes, VMs, network separation or supported GPU partitions. Hardware-agnostic operation means selecting supported controls for each device and workload, not pretending every accelerator offers identical partitioning or isolation.

Quotas should cover concurrent accelerators, partition profiles, CPU and memory, storage capacity and relevant network limits. Reservations require time-aware admission decisions in addition to static quotas. A tenant allowed eight GPUs must not consume another tenant’s guaranteed window simply because eight devices happen to appear idle.

For partitioned service, meter the sold profile directly. Four small partition-hours should not automatically become one whole-GPU hour in the financial ledger. Keep an explicit conversion policy for internal capacity analysis, while preserving the customer-facing unit, price and profile history for chargeback.

Meter the boundary the contract actually sells

A useful usage record includes tenant, contract identifier, resource identifier, partition profile, allocation start and end, tariff version and billable status. Add adjustment references for outages and credits. Prefer reproducible calculations over dashboard screenshots, and preserve raw events so corrected invoices can be explained.

Clock synchronisation and duplicate handling are billing controls. A repeated allocation event must not create a second charge; a missing release event must not create an indefinite one. Reconcile metered duration against scheduler allocations and node availability, then quarantine exceptions rather than automatically invoicing every anomaly.

A short reconciliation check can catch a fundamental error before billing. This illustrative Python block is not a Clastiq CLI:

installed = 32 * 30 * 24
planned_maintenance = 32 * 24
available = installed - planned_maintenance
reserved = 12000
on_demand = 4000
billable = reserved + on_demand
assert billable <= available
print(f"Billable utilisation: {billable / installed:.1%}")

The assertion is necessary but insufficient: it validates monthly totals, not simultaneous reservations. Run interval-level checks as well. Keep the finance export versioned, with explicit currency, time zone, rounding policy and treatment of partial hours. Otherwise, two teams can calculate different invoices from apparently identical records.

Storage, power and sovereignty constrain saleable hours

GPU availability does not guarantee workload readiness. Slow dataset staging, insufficient scratch space or a saturated storage network can leave allocated accelerators waiting. Ceph storage should enter the operating model through capacity, performance requirements, replication overhead and recovery behaviour, not just a cost-per-terabyte line.

Separate persistent tenant data, images, checkpoints and temporary scratch where their service requirements differ. Test recovery under representative load before promising availability. A storage rebuild that degrades every tenant can turn one infrastructure event into several service-credit claims. Include storage maintenance in reservation planning rather than treating it as unrelated background work.

Power imposes a second admission limit. Nameplate GPU capacity can exceed what the rack’s electrical or cooling envelope can sustain concurrently. Measure whole-system draw, including hosts and networking, and account for cooling at a clearly stated boundary. Use that information when setting concurrency limits and maintenance windows.

Sovereignty adds operational obligations rather than removing economics. Identify where tenant data, usage records, administrative access and backups reside. For air-gap deployments, plan local image distribution, approved update transfer, clock discipline and offline meter retention. Any exported billing summary needs an explicit approval path consistent with the operator’s policy.

Cognition-AI, based in Muscat, Oman, describes Clastiq Orchestrator as supporting air-gap deployments and human-led local support. Those capabilities matter when an operator must maintain local control, but each deployment still needs documented access procedures, recovery ownership and contractual support arrangements. They should not be interpreted as an unstated certification or availability guarantee.

What to do this week

  1. Publish three utilisation figures. Calculate installed, available, delivered and billable hours for the last complete month. Label every denominator and investigate the gap between busy devices and invoiced service.
  2. Audit reservation promises. Map each contract to capacity windows, cancellation terms, downtime treatment and payment dates. Check concurrent commitments against both accelerator inventory and the power envelope.
  3. Rebuild the cost worksheet. Separate depreciation, electricity, storage, support and facility costs. Add financing sensitivity and a distinct cash bridge; test a 20% demand reduction and a delayed major payment.
  4. Reconcile one tenant end to end. Follow provisioning, allocation, meter events, tariff application, credits and invoice approval. Correct duplicate events, ambiguous partition units and missing release records before expanding automation.
  5. Stress-test concentration and recovery. Model losing the largest customer, then rehearse a storage or node outage during a reserved window. Assign owners for capacity restoration, customer communication and billing adjustments.

Clastiq Orchestrator runs this operating model on the operator’s own hardware: request a demo.

Sources

Clastiq Orchestrator runs this on your own infrastructure — from a quarter rack to a few racks.

Request a demo