Skip to content
Alex Zhebrak
Go back

Decode Decoded

· blog

Decode performance of vLLM, SGLang, and TensorRT-LLM on H100 SXM across Qwen 3 0.6B-32B models for 1-256 batch sizes analysed with Nsight Systems.

Somewhat surprisingly, the benchmarks didn’t record a host-bound regime even for the smallest model, meaning the GPU was always busy and the CPU never became the bottleneck (at least with CUDA graphs turned on). However, being busy does not always mean being efficient or fast, as the GPU can spend time on kernels, copies, or CUDA graph launches without fully utilising its memory bandwidth. Each step below is split into the bandwidth floor (reading weights and the KV cache from HBM at the card’s measured memory bandwidth), kernel excess (kernel time over the bandwidth floor), and kernel launch cost, with host stalling and copies as the remainder.

Decode step share for Qwen3-0.6B by batch size, one stacked bar per engine: bandwidth floor, copies, kernel excess, launch cost and host stall Decode step share for Qwen3-0.6B by batch size, one stacked bar per engine: bandwidth floor, copies, kernel excess, launch cost and host stall

Less surprisingly, with smaller models in general and smaller batch sizes for smaller models, kernels’ fixed cost couldn’t be amortised by the bandwidth floor. Even with CUDA graphs, the inter-kernel launch overhead still exists, now within the card itself.

Decode step share by model size at batch 1 and 32, two stacked bars per engine in each model group Decode step share by model size at batch 1 and 32, two stacked bars per engine in each model group

With the same cuBLAS kernels, the biggest difference between engines was when serving smaller models, with TensorRT-LLM underperforming because of its custom attention kernel (assessed via kernel name-matching of Nsight traces, kernel counts per step, and time per step). Hollow markers below are within the noise.

Each engine's decode step over the fastest engine's, by model size, at batch 1 and 16 Each engine's decode step over the fastest engine's, by model size, at batch 1 and 16

For engine optimisations, CUDA graphs made the biggest difference overall and are disproportionately more important for smaller models and smaller batch sizes in general. Scheduler overlap also reduces step time, rescuing performance for smaller models, although its relationship with batch size varied by engine. Model Runner V2 for vLLM made no meaningful difference in these runs.

Engine optimisations switched off one at a time: step time and GPU busy share with CUDA graphs off, and step time with the scheduler off, by model and batch Engine optimisations switched off one at a time: step time and GPU busy share with CUDA graphs off, and step time with the scheduler off, by model and batch

This benchmark ran on a host with one H100 SXM GPU and 8 CPU cores, on dense bf16 models, and is tied to specific releases of the serving engines (vLLM 0.30.0, SGLang 0.5.20, TensorRT-LLM 1.2.1). Prompts of 896 tokens; 288 tokens generated. Three timed generations per cell after two warm-up generations. Noise measured separately with representative runs. Code is available at zhebrak/decode_decoded.


Share this post on:

Next Post
Specialisation & Co-design