NVIDIA
Popular

NVIDIA RTX 5090

32GB GDDR7 Blackwell — high-throughput inference, AIGC and fine-tuning at the lowest entry point

Reserved and on-demand, deployed in Thailand

$0.75/hr per GPU

All prices in USD, ex-tax.

Compare GPU models

TATC-deployed RTX 5090 bare-metal servers feature 32GB GDDR7 memory and 1.79 TB/s bandwidth for high-density AI compute. Engineered for 7B–70B model inference, AIGC generation, and single-node fine-tuning with full root access and single-tenant isolation.

Memory
32GB GDDR7
1.79 TB/s bandwidth
Architecture
Blackwell
Family · Enterprise Blackwell Series
Compute (FP16)
~3.35 PFLOPS (8-GPU, FP16 tensor, est.)
BF16: 1.68 PFLOPS (8-GPU, BF16 tensor) · FP8: 6.70 PFLOPS (8-GPU, FP8 tensor)
Interconnect
PCIe Gen 5 x16 Interconnect
TDP · 575W

Highlights

  • 32GB GDDR7 with 1.79 TB/s — 70B-class LoRA inference on a single card
  • Blackwell 5th-gen Tensor Cores with FP8 tensor acceleration (6.70 PFLOPS, 8-GPU)
  • Best price-per-GPU entry into the TATC bare-metal catalogue

Software stack

CUDA 12.8+
PyTorch 2.8+
Triton
vLLM
TensorRT-LLM

Best-fit scenarios

  • High-throughput LLM inference (7B–70B with offloading)
  • AIGC image / video generation
  • Single-node fine-tuning and LoRA training

Ready to deploy NVIDIA RTX 5090?

Spin up a dedicated server in days, or talk to our sales engineers about custom configurations.

Contact Us