v0.2.2 · c1d32fd
2026-08-20 06:42:17 UTC

GPU explorer

LLM inference planning — compare GPU generations by memory, bandwidth, and cost efficiency

Preset:

Vendor:

X Axis:

Y Axis:

Loading GPU catalog…

💡 Top-right = high VRAM and throughput. These GPUs handle larger models and longer contexts.

Bubble size represents BF16 TFLOPs. Cost efficiency axes and tokens-per-dollar will be restored once the Costings API is available.

Throughput Index is a planning metric derived from memory bandwidth, VRAM, and architecture generation. It enables relative GPU comparison — not exact model throughput.

Inference performance depends on model architecture (GQA vs MHA), sequence length, batching, and inference backend (vLLM, TensorRT-LLM, etc.).

Hardware cost and cloud pricing axes are pending the Costings REST API.