GPU as a Service · GPUaaS
One card or
five hundred
Containerised pods for iteration, bare metal for production, and interconnected clusters when a single node stops being enough. Same catalogue, same console, same bill.
Catalogue
Available today
| Accelerator | Memory | Interconnect | Best for | On-demand | Region |
|---|---|---|---|---|---|
| NVIDIA H200 SXM | 141 GB HBM3e | NVLink + 400G IB | Frontier-scale training | ₹319 /hr | BOM1 · PNQ1 |
| NVIDIA H100 SXM | 80 GB HBM3 | NVLink + 400G IB | Multi-node training | ₹239 /hr | BOM1 · PNQ1 · MAA1 |
| NVIDIA H100 PCIe | 80 GB HBM3 | PCIe Gen5 | Single-node training | ₹199 /hr | All regions |
| NVIDIA A100 SXM | 80 GB HBM2e | NVLink + 200G IB | Fine-tuning, HPC | ₹149 /hr | All regions |
| NVIDIA L40S | 48 GB GDDR6 | PCIe Gen4 | Inference, render | ₹99 /hr | All regions |
| NVIDIA RTX 6000 Ada | 48 GB GDDR6 | PCIe Gen4 | Visualisation, CAD | ₹79 /hr | PNQ1 · DEL1 |
| NVIDIA RTX 4090 | 24 GB GDDR6X | PCIe Gen4 | Development, LoRA | ₹44 /hr | PNQ1 · DEL1 |
| AMD Instinct MI300X | 192 GB HBM3 | Infinity Fabric | Large-context inference | ₹269 /hr | BOM1 |
Rates are per GPU-hour, billed per second, exclusive of GST. Reserved and committed terms are lower; see pricing.
Three shapes
Pick how close
to the metal
Containerised
Your image, our scheduler. Starts in about 40 seconds, stops when idle, and bills only while running. This is where most teams begin.
Whole machines
Root on the host, no hypervisor tax, your own drivers and kernel. Monthly or longer, with the NICs and NVMe wired the way you specify.
Interconnected
16 to 512 GPUs on non-blocking 400G InfiniBand, delivered as a named environment with Slurm or Kubernetes already running.
Cluster engineering
Fabric that
actually scales
A rack of H100s is not a cluster. Getting linear scaling out of sixty-four nodes is a topology and tuning problem, and it is the part we do for you before handover.
- Rail-optimised fat tree: non-blocking, with SHARP in-network reduction on supported fabrics
- NCCL tuned per topology: we hand you the benchmark numbers, not a brochure
- GPUDirect Storage: checkpoints stream to NVMe without a bounce through host memory
- Health-checked nodes: every node passes a burn-in and DCGM suite before it enters your pool
- Slurm or Kubernetes: your choice, configured, with the operator stack in place
size(B) time(us) busbw(GB/s)
1073741824 4812.3 371.4
2147483648 9503.7 376.1
4294967296 18871.2 378.9
8589934592 37622.8 380.2
# scaling efficiency 8 → 64 nodes: 94.1%
Representative figures from an acceptance run. We publish yours before you sign anything.
Need more than
a single node?
Tell us the model size and the deadline. We'll come back with a topology, a price and a date.