← Founder Notes
Archive

Self-hosting a 2.4t open moe only pays off near full utilization. a september 18 breakdown of…

Yethikrishna ROriginal on Threads

self-hosting a 2.4t open moe only pays off near full utilization. a september 18 breakdown of qwen3.8-max works the break-even against hosted pricing of $2 and $6 per million tokens, and idle gpus flip the math negative fast.

the frontier is priced for tenants, not owners.

Context

A Contra Collective post of 18 September 2026, from a vendor that sells serving infrastructure, says Alibaba released weights as Qwen3.8-2.4T-A95B on 12 August with FP8 checkpoints, 2.4 TB in FP8, 95B active parameters per token, and about 17 H200 GPUs of 141 GB for the weights alone by its own arithmetic. It says hosted Qwen3.8-Max-0902 is priced at 2 dollars per million input tokens and 6 dollars per million output, and that beating it needs very high, very steady utilization. The Hugging Face model card says the FP8 weights work with vLLM and SGLang, carry the qwen3.8-max license, and that Qwen3.8-Max is the official version based on the open model with more features such as vision input, 1M context by default and built-in tools.

How it compares

The post gives no computed break-even figure in the text read, so works the break-even is qualitative. The 2 and 6 dollar prices are secondary and not confirmed on a first-party price page. The hosted Max is a different product from the open weights. The 17 GPU figure is one vendor's projection, not a demonstrated deployment, and 2.4T is total parameters with 95B active. That the frontier is priced for tenants, not owners, is the author's take.

The model card says Max is the official version with extra features, so a hosted versus self-hosted price comparison is not like for like. Alibaba Cloud's Model Studio page, updated 11 September 2026, carried no price lines in the text read, so the 2 and 6 dollar prices stay unconfirmed first-party. A third-party post calls the open model and Max the same brain sold two ways, which conflicts with the card's wording and is not resolved. Another post gives a 4.89 TB checkpoint against the 2.4 TB FP8 figure; the cause was not checked and the discrepancy stays open.

Related work

Watch next

  • Alibaba Cloud's pricing page and a deployment report with measured utilization.

Sources

  1. Self-host Qwen3.8-Max 2.4T MoE infrastructure requirements (Contra Collective, 18 Sep 2026)contracollective.com
  2. Qwen3.8-2.4T-A95B-FP8 (Hugging Face)huggingface.co
  3. Qwen3.8-2.4T-A95B on Model Studio (Alibaba Cloud)alibabacloud.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 10:46 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/self-hosting-a-2-4t-open-moe-only-Ddf0C1HDFoZ" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Self-hosting a 2.4t open moe only pays off near full utilization. a september 18 breakdown of…"></iframe>

More notes