← Founder Notes
Archive

The gguf quantization rules just got replaced by measurements. bartowski's per-tensor maps, out…

Yethikrishna ROriginal on Threads

the gguf quantization rules just got replaced by measurements. bartowski's per-tensor maps, out september 10, assign precision inside gguf files from measured sensitivity instead of general heuristics.

the 4-bit rule died quietly.

Context

Bartowski's post on Hugging Face of September 10, 2026 says upstream llama.cpp quantization uses model-agnostic heuristics, for example bumping the first and last eighth of layers, and describes a framework that measured per-tensor sensitivity to build per-tensor layout maps for GGUF quantization. The metric is KL divergence against the bf16 model using llama-perplexity on wikitext-2-raw at context 512, lower being better. The work was mainly on Qwen models with validation on others. Findings include that token embeddings dominate sensitivity, depth is U-shaped, and per bit the small attention projections and ffn_up and ssm_out are most sensitive while ffn_gate rarely merits extra bits.

How it compares

The results are tied to that metric, dataset, context length and the tested models, and are one author's measurements. Replacement within his own maps fits the framing, but the post proposes per-tensor layout maps as an approach, and no evidence was inspected that upstream defaults changed or that upstream adopted them. The 4-bit rule is not a rule named in the text read. Whether the maps are released as files or tooling was not extracted. The 96 hour and 1,000 plus configuration counts come from a secondary snippet. The gguf quantization rules just got replaced by measurements and the 4-bit rule died quietly are the note author's framing and take.

Watch next

  • Upstream llama.cpp adoption and released maps.

Sources

  1. Hugging Face: per-tensor layout maps for GGUF quantization (September 10, 2026)huggingface.co

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 23 September 2026 at 03:48 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-gguf-quantization-rules-just-got-replaced-by-DdmymagCCe4" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The gguf quantization rules just got replaced by measurements. bartowski's per-tensor maps, out…"></iframe>

More notes