GPU as a Service  ·  GPUaaS

One card or five hundred

Containerised pods for iteration, bare metal for production, and interconnected clusters when a single node stops being enough. Same catalogue, same console, same bill.

Catalogue

Available today

AcceleratorMemoryInterconnectBest forOn-demandRegion
NVIDIA H200 SXM141 GB HBM3eNVLink + 400G IBFrontier-scale training₹319 /hrBOM1 · PNQ1
NVIDIA H100 SXM80 GB HBM3NVLink + 400G IBMulti-node training₹239 /hrBOM1 · PNQ1 · MAA1
NVIDIA H100 PCIe80 GB HBM3PCIe Gen5Single-node training₹199 /hrAll regions
NVIDIA A100 SXM80 GB HBM2eNVLink + 200G IBFine-tuning, HPC₹149 /hrAll regions
NVIDIA L40S48 GB GDDR6PCIe Gen4Inference, render₹99 /hrAll regions
NVIDIA RTX 6000 Ada48 GB GDDR6PCIe Gen4Visualisation, CAD₹79 /hrPNQ1 · DEL1
NVIDIA RTX 409024 GB GDDR6XPCIe Gen4Development, LoRA₹44 /hrPNQ1 · DEL1
AMD Instinct MI300X192 GB HBM3Infinity FabricLarge-context inference₹269 /hrBOM1

Rates are per GPU-hour, billed per second, exclusive of GST. Reserved and committed terms are lower; see pricing.

U38

Three shapes

Pick how close to the metal

Pods

Containerised

Your image, our scheduler. Starts in about 40 seconds, stops when idle, and bills only while running. This is where most teams begin.

Bare metal

Whole machines

Root on the host, no hypervisor tax, your own drivers and kernel. Monthly or longer, with the NICs and NVMe wired the way you specify.

Clusters

Interconnected

16 to 512 GPUs on non-blocking 400G InfiniBand, delivered as a named environment with Slurm or Kubernetes already running.

Cluster engineering

Fabric that actually scales

A rack of H100s is not a cluster. Getting linear scaling out of sixty-four nodes is a topology and tuning problem, and it is the part we do for you before handover.

  • Rail-optimised fat tree: non-blocking, with SHARP in-network reduction on supported fabrics
  • NCCL tuned per topology: we hand you the benchmark numbers, not a brochure
  • GPUDirect Storage: checkpoints stream to NVMe without a bounce through host memory
  • Health-checked nodes: every node passes a burn-in and DCGM suite before it enters your pool
  • Slurm or Kubernetes: your choice, configured, with the operator stack in place
nccl-tests: 64× H100 SXM, BOM1
# all_reduce_perf -b 8 -e 8G -f 2 -g 8

size(B)      time(us)   busbw(GB/s)
1073741824   4812.3      371.4
2147483648   9503.7      376.1
4294967296  18871.2      378.9
8589934592  37622.8      380.2

# scaling efficiency 8 → 64 nodes: 94.1%

Representative figures from an acceptance run. We publish yours before you sign anything.

Need more than a single node?

Tell us the model size and the deadline. We'll come back with a topology, a price and a date.