NVIDIA

NVIDIA RTX PRO 6000

96GB GDDR7 Blackwell-architecture GPU — high-density inference and AIGC acceleration

Reserved and on-demand, deployed in Thailand

$1.45/hr per GPU

All prices in USD, ex-tax.

Compare GPU models

Deployed as high-density physical bare-metal infrastructure by TATC, the RTX PRO 6000 features 96GB of high-capacity GDDR7 memory and 1.8 TB/s bandwidth, optimized for large-scale LLM inference, AIGC workflows, and model fine-tuning where NVLink scale-out is not required.

Memory
96GB GDDR7
1.8 TB/s bandwidth
Architecture
Blackwell
Family · NVIDIA RTX PRO
Compute (FP16)
~4.03 PFLOPS (8-GPU, FP16 tensor, est.)
BF16: 2.01 PFLOPS (8-GPU, BF16 tensor) · FP8: 8.06 PFLOPS (8-GPU, FP8 tensor)
Interconnect
100G RoCE High-Speed Fabric
TDP · 600W

Highlights

  • Cost-effective entry point for high-VRAM CUDA inference
  • Holds 70B–200B quantised models entirely in VRAM
  • Seamless integration with Hugging Face transformers ecosystem

Software stack

CUDA 12.x
PyTorch 2.x
Triton
vLLM

Best-fit scenarios

  • LLM inference (70B-class models fully in VRAM)
  • AIGC image / video generation
  • Small-model fine-tuning

Ready to deploy NVIDIA RTX PRO 6000?

Spin up a dedicated server in days, or talk to our sales engineers about custom configurations.

Contact Us