← Founder Notes
Archive

Serving got 2.8x faster with zero new hardware. vllm's september 13 optimization of kimi k3 cut…

Yethikrishna ROriginal on Threads

serving got 2.8x faster with zero new hardware. vllm's september 13 optimization of kimi k3 cut latency 56-60% and first-token time up to 85% just by tuning the engine.

the model didn't change, the serving code did.

Context

vLLM's blog of 13 September 2026 compares v0.27.1 against a main-branch commit on Kimi K3, with an 8K input and 1K output workload on a B300 node, TP8 and eight-token DSpark speculation, at concurrency 1, 4 and 16. Average latency fell from 12.37 to 5.30 seconds (minus 57.2 percent), 23.67 to 10.50 (minus 55.6) and 55.90 to 22.17 (minus 60.3); average first-token time fell from 2262.9 to 376.3 milliseconds (minus 83.4 percent), 2314.9 to 640.5 (minus 72.3) and 7601.1 to 1121.0 (minus 85.3); throughput rose 120, 150 and 180.6 percent.

How it compares

It is vendor-reported, on one workload and one hardware tier, and not independent. The 2.8x is throughput and not latency, and the latency gain is about 2.3 to 2.5 times by my arithmetic from the table. The comparison is against a main-branch commit and not a tagged release. Zero new hardware holds as the same B300 node on both sides, but just by tuning the engine simplifies many software changes such as scheduler, prefix cache and kernels. Other figures in the blog use different configurations and are not merged here.

Related work

Watch next

  • The release tag containing the commit and independent reruns.

Sources

  1. The road to 2.8x throughput (vLLM, 13 Sep 2026)vllm.ai

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 16:22 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/serving-got-2-8x-faster-with-zero-new-Ddi_T7sgMHU" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Serving got 2.8x faster with zero new hardware. vllm's september 13 optimization of kimi k3 cut…"></iframe>

More notes