v0.2.9 · 6f58929
2026-10-04 23:59:18 UTC

Predict performance

See time to first token, throughput, and memory estimates in seconds — refine as needed.

Tested: Nemotron, DeepSeek V4, Gemma 4, Kimi, ... — search by size (“70b”), type (“vision”, “moe”) or precision (“fp8”)
Serving mode:
Based on your configuration — ISL 2048, OSL 128, FP16 KV cache, 1 concurrent users.
Want to change assumptions?

Every number above comes from these. Open a section to tune it — closed sections show their current values.