NVIDIA
NVIDIA RTX PRO 6000
96GB GDDR7 Blackwell-architecture GPU — high-density inference and AIGC acceleration
Reserved and on-demand, deployed in Thailand
$1.45/hr per GPU
All prices in USD, ex-tax.
Deployed as high-density physical bare-metal infrastructure by TATC, the RTX PRO 6000 features 96GB of high-capacity GDDR7 memory and 1.8 TB/s bandwidth, optimized for large-scale LLM inference, AIGC workflows, and model fine-tuning where NVLink scale-out is not required.
Memory
96GB GDDR7
1.8 TB/s bandwidth
Architecture
Blackwell
Family · NVIDIA RTX PRO
Compute (FP16)
~4.03 PFLOPS (8-GPU, FP16 tensor, est.)
BF16: 2.01 PFLOPS (8-GPU, BF16 tensor) · FP8: 8.06 PFLOPS (8-GPU, FP8 tensor)
Interconnect
100G RoCE High-Speed Fabric
TDP · 600W
Highlights
- Cost-effective entry point for high-VRAM CUDA inference
- Holds 70B–200B quantised models entirely in VRAM
- Seamless integration with Hugging Face transformers ecosystem
Software stack
CUDA 12.x
PyTorch 2.x
Triton
vLLM
Best-fit scenarios
- LLM inference (70B-class models fully in VRAM)
- AIGC image / video generation
- Small-model fine-tuning
Ready to deploy NVIDIA RTX PRO 6000?
Spin up a dedicated server in days, or talk to our sales engineers about custom configurations.