← Founder Notes
Archive

The open source serving engine just made the same model 2.8 times faster. vllm's kimi k3…

Yethikrishna ROriginal on Threads

the open source serving engine just made the same model 2.8 times faster. vllm's kimi k3 optimizations, out september 13, combine scheduling, prefix caching and pd disaggregation.

nobody touched the weights.

Context

The vLLM blog of September 13, 2026 reports merged work for Kimi K3, comparing v0.27.1 against main at commit 82a85dc1. It reports 56 to 60 percent lower latency, 2.2 to 2.8x throughput and 72 to 85 percent lower time to first token across concurrency 1, 4 and 16, on an 8K input, 1K output workload with TP8, eight-token DSpark speculation and a B300 node on CUDA 13.3. The listed changes include adaptive speculative-token budgets, internal KDA prefix checkpoints, zero-copy mixed batches and deferred MXFP4 finalization. The post describes prefill-decode disaggregation in a separate section.

How it compares

This is a blog about merged work and not a versioned release, and the project reports its own numbers on one hardware and configuration, so 2.8x is the upper end at one concurrency and not universal. Prefix caching was disabled in both benchmark runs because v0.27.1 had a known Kimi K3 prefix-caching issue fixed later, so it did not contribute to the headline, and no measured effect of prefill-decode disaggregation on the 2.8x was stated in the text read. The note's combination of scheduling, prefix caching and pd disaggregation therefore overstates what the headline benchmark measured. Nobody touched the weights is the author's take.

Watch next

  • A tagged vLLM release containing these changes and independent replications.

Sources

  1. vLLM blog: Kimi K3 performance optimization (September 13, 2026)vllm.ai

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 22 September 2026 at 17:46 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-open-source-serving-engine-just-made-the-DdltuKFDUwf" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The open source serving engine just made the same model 2.8 times faster. vllm's kimi k3…"></iframe>

More notes