← Founder Notes
Archive

Speculative decoding, the trick most inference stacks lean on, is 2.9x faster at batch 1 and loses…

Yethikrishna ROriginal on Threads

speculative decoding, the trick most inference stacks lean on, is 2.9x faster at batch 1 and loses past batch 32, and a drafter with 19 percent acceptance actively slows you down 12 percent, per the new speed-bench. the fastest path depends on your batch size, not your model.

most teams are optimizing the wrong regime.

Context

SPEED-Bench (arXiv 2604.09557), listed by NVIDIA research on 23 February 2026, is a unified and diverse benchmark for speculative decoding. Its qualitative split runs at batch size 32 on TensorRT-LLM and SGLang with Qwen3 models, and reports speedups across categories up to about 3 in its tables. A small simulation repository by JohnScheuer, created 19 July 2026, reports a gamma crossover for batched speculative decoding.

How it compares

The note does not say which speed-bench it means, and SPEED-Bench is the established one, so it is not new in September. The note's 2.9x at batch 1, the loss past batch 32, the 19 percent acceptance and the 12 percent slowdown were not located in the SPEED-Bench text read, nor in the small repository, and the table cells seen were not batch-1 results, so all four figures are unverified. The repository is not shown to be the note's source. The paper says acceptance is data-dependent, and speedup depends on draft model, acceptance rate, draft length, engine and workload. Most teams are optimizing the wrong regime is the author's opinion.

Related work

Watch next

  • Which benchmark the note's numbers come from, and the per-batch-size tables.

Sources

  1. SPEED-Bench (arXiv 2604.09557)arxiv.org
  2. batched-speculative-decoding-bench (GitHub)github.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 18:48 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/speculative-decoding-the-trick-most-inference-stacks-lean-DdbhnfXiCFB" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Speculative decoding, the trick most inference stacks lean on, is 2.9x faster at batch 1 and loses…"></iframe>

More notes