GPU explorer
LLM inference planning — compare GPU generations by memory, bandwidth, and cost efficiency
Preset:
Vendor:
X Axis:
Y Axis:
Loading GPU catalog…
💡 Top-right = high VRAM and throughput. These GPUs handle larger models and longer contexts.
Bubble size represents BF16 TFLOPs. Cost efficiency axes and tokens-per-dollar will be restored once the Costings API is available.
Throughput Index is a planning metric derived from memory bandwidth, VRAM, and architecture generation. It enables relative GPU comparison — not exact model throughput.
Inference performance depends on model architecture (GQA vs MHA), sequence length, batching, and inference backend (vLLM, TensorRT-LLM, etc.).
Hardware cost and cloud pricing axes are pending the Costings REST API.