← Founder Notes
Archive

Alibaba shipped qwen3.8-flash-next with a 51b parameter component designed to live in system ram…

Yethikrishna ROriginal on Threads

alibaba shipped qwen3.8-flash-next with a 51b parameter component designed to live in system ram instead of gpu memory. the interesting part is not the count. it is that the inference graph now treats cpu as a tier of vram.

the line between model and context just got blurry.

Context

The Hugging Face model card for Qwen3.8-Flash-Next lists 125B parameters with 6B activated, plus a 51B n-gram embedding and 4B MTP. The card says the n-gram embedding requires less computation and is more amenable to offloading than Mixture-of-Experts, and suits memory-constrained accelerators.

How it compares

The card supports a 51B n-gram embedding component and the offloading language. It does not state that the component is designed to live in system RAM, so that wording is the author's reading. The card date seen in search metadata is 26 August 2026, before the note, so the model was not new on 5 September. Treating CPU as a tier of VRAM and the model and context line blurring are the author's opinion.

Related work

Watch next

  • Qwen documentation on where the n-gram embedding is meant to run.

Sources

  1. Qwen3.8-Flash-Next model card (Hugging Face)huggingface.co
  2. Qwen3.8-Flash-Next technical reportraw.githubusercontent.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 5 September 2026 at 07:03 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/alibaba-shipped-qwen3-8-flash-next-with-a-Dc4yi_RiOmU" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Alibaba shipped qwen3.8-flash-next with a 51b parameter component designed to live in system ram…"></iframe>

More notes