Decode performance of vLLM, SGLang, and TensorRT-LLM on H100 SXM across Qwen 3 0.6B-32B models for 1-256 batch sizes analysed with Nsight Systems.
Somewhat surprisingly, the benchmarks didn’t record a host-bound regime even for the smallest model, meaning the GPU was always busy and the CPU never became the bottleneck (at least with CUDA graphs turned on). However, being busy does not always mean being efficient or fast, as the GPU can spend time on kernels, copies, or CUDA graph launches without fully utilising its memory bandwidth. Each step below is split into the bandwidth floor (reading weights and the KV cache from HBM at the card’s measured memory bandwidth), kernel excess (kernel time over the bandwidth floor), and kernel launch cost, with host stalling and copies as the remainder.
Less surprisingly, with smaller models in general and smaller batch sizes for smaller models, kernels’ fixed cost couldn’t be amortised by the bandwidth floor. Even with CUDA graphs, the inter-kernel launch overhead still exists, now within the card itself.
With the same cuBLAS kernels, the biggest difference between engines was when serving smaller models, with TensorRT-LLM underperforming because of its custom attention kernel (assessed via kernel name-matching of Nsight traces, kernel counts per step, and time per step). Hollow markers below are within the noise.
For engine optimisations, CUDA graphs made the biggest difference overall and are disproportionately more important for smaller models and smaller batch sizes in general. Scheduler overlap also reduces step time, rescuing performance for smaller models, although its relationship with batch size varied by engine. Model Runner V2 for vLLM made no meaningful difference in these runs.
This benchmark ran on a host with one H100 SXM GPU and 8 CPU cores, on dense bf16 models, and is tied to specific releases of the serving engines (vLLM 0.30.0, SGLang 0.5.20, TensorRT-LLM 1.2.1). Prompts of 896 tokens; 288 tokens generated. Three timed generations per cell after two warm-up generations. Noise measured separately with representative runs. Code is available at zhebrak/decode_decoded.
Related work and further reading
- Spector et al., Look Ma, No Bubbles!, 2025
- Li et al., ForgeMegakernel, 2026
- Vellaisamy et al., TaxBreak, ISPASS 2026
- Recasens et al., Mind the Memory Gap, 2025
- Chen, Memory-Bound but Not Bandwidth-Limited, 2026
- Siavashi et al., Dissecting GPU Utilization for LLM Inference on Nvidia Hopper, 2026
- Siavashi et al., Blink, 2026
- Ye et al., FlashInfer, MLSys 2025
- vLLM: Model Runner V2, 2026