Performance estimate
See time to first token, throughput, and memory estimates in seconds — refine as needed.
Checking...
Tested: , ... — type to autocomplete
Gated model? Add your HF token in Settings →
Based on your configuration — ISL 2048, OSL 128, FP16 KV cache, 32 concurrent users.
Select a model and GPU system above, then press Calculate to see results.
Want to change assumptions?