Speculative decoding, the trick most inference stacks lean on, is 2.9x faster at batch 1 and loses…
speculative decoding, the trick most inference stacks lean on, is 2.9x faster at batch 1 and loses past batch 32, and a drafter with 19 percent acceptance actively slows you down 12 percent, per the new speed-bench. the fastest path depends on your batch size, not your model.
most teams are optimizing the wrong regime.
Context
SPEED-Bench (arXiv 2604.09557), listed by NVIDIA research on 23 February 2026, is a unified and diverse benchmark for speculative decoding. Its qualitative split runs at batch size 32 on TensorRT-LLM and SGLang with Qwen3 models, and reports speedups across categories up to about 3 in its tables. A small simulation repository by JohnScheuer, created 19 July 2026, reports a gamma crossover for batched speculative decoding.
The note does not say which speed-bench it means, and SPEED-Bench is the established one, so it is not new in September. The note's 2.9x at batch 1, the loss past batch 32, the 19 percent acceptance and the 12 percent slowdown were not located in the SPEED-Bench text read, nor in the small repository, and the table cells seen were not batch-1 results, so all four figures are unverified. The repository is not shown to be the note's source. The paper says acceptance is data-dependent, and speedup depends on draft model, acceptance rate, draft length, engine and workload. Most teams are optimizing the wrong regime is the author's opinion.
Related work
- SPEED-Bench (NVIDIA Research, Feb 2026) ↗Snippet only.
- SPEED-Bench (Hugging Face blog, 19 Mar 2026) ↗Snippet only.
- ICML 2026 slides ↗Snippet only.
Watch next
- Which benchmark the note's numbers come from, and the per-batch-size tables.
Sources
- SPEED-Bench (arXiv 2604.09557)arxiv.org
- batched-speculative-decoding-bench (GitHub)github.com
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 18:48 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →