Enterprise · GPU

Your GPUs are the most contended resource you own and the least governed

Realm9 does not replace your scheduler. It sits above it — deciding who is entitled to what, for how long, at what cost, and who has to approve it. Placement stays with Kubernetes, Slurm or your hypervisor.

Fair-shareDeserved quota per project, with controlled over-quota borrowing
FractionalMIG partitioning and time-slicing exposed as bookable units
Multi-clusterKubernetes, Slurm, bare metal and neocloud in one entitlement view
ChargebackEvery GPU-hour attributed to a principal, a project and an approval

We deliberately do not build a scheduler

Kubernetes now handles GPUs properly. Dynamic Resource Allocation is GA, NVIDIA donated its DRA driver to the CNCF, and the KAI Scheduler is a CNCF sandbox project. Placement, bin-packing, gang scheduling and preemption are solved.

What is not solved is entitlement: which team is allowed how much, for how long, at what cost, and who signs off when they want more. Realm9 owns that layer and delegates placement downward.

  • Delegates to KAI, Kueue and Volcano on Kubernetes; to Slurm on HPC estates
  • Delegates to vCenter and Proxmox for VM-attached accelerators
  • Adds the quota, approval, budget and ledger those schedulers do not have
  • Works identically across owned clusters and rented neocloud capacity
GPU lease request · data-platform, 8×H100, 6h
Deserved quota16 GPU (project)
In use12 GPU
Requested8 GPU (+4 over quota)
BorrowableYes — shared pool idle 21 GPU
Est. cost$811.20 @ $16.90/GPU-hr
Budget remaining$14,180 of $40,000
Approval
Over-quota — routed to @ml-platform-leads
Placement delegated to kai-scheduler

The problems that actually cost you money

Idle at night, queued at noon

Utilisation is a scheduling problem only after it is an entitlement problem. TTLs, idle detection and preemptible over-quota borrowing recover capacity nobody is using.

The team that took everything

Deserved quota per project with controlled borrowing means one team's experiment cannot starve another team's deadline — without an administrator refereeing it manually.

Fragmented fleets

Owned clusters, HPC partitions and rented neocloud capacity appear as one entitlement surface, so demand routes to whatever is free and cheapest.

Whole GPUs for small jobs

MIG partitions and time-sliced shares are bookable units in their own right, so an inference job does not consume an H100 it cannot use.

Agents with no ceiling

An autonomous training pipeline is a principal like any other: named, quota-bound, budget-capped and revocable in one call.

No answer for finance

Every GPU-hour attributed to a project, a principal and an approval, on the same ledger as cloud and token spend.

Turn your GPU fleet into governed, accountable capacity

Works with the schedulers you already run. No migration, no replacement.