← Founder Notes
Archive

The cost war moved in 12 months from which model to how do we fit 4m tokens into a 200k window and…

Yethikrishna ROriginal on Threads

the cost war moved in 12 months from which model to how do we fit 4m tokens into a 200k window and shopify gisting is the first one to admit it out loud.

learned tokens for the system prompt beat any 75% cache cut, every single time.

Context

Shopify Engineering's post of 19 August 2026 describes gisting, which compresses context into a set of learned tokens. For its Sidekick GraphQL agent, the system prompt went from about 6,000 tokens to about 1,500, a 4 to 1 ratio, without losing prediction quality. At 350 requests per minute, time to first token went from 438 ms to 354 ms, end to end from 6.8 s to 4.2 s, and throughput from 20.2 to 23.4 queries per second.

The post says gisting and prefix caching are not mutually exclusive.

How it compares

The 4M into 200k tokens figure, the 75 percent cache cut and beating caching every single time are not in the post. Because the post treats the two as complementary, a cost comparison between them is untested here, not refuted. Being the first to admit it, and the 12 month shift, are the author's opinion. Shopify's example is one agent's system prompt, not 4M token compression.

Related work

Watch next

  • Follow-up posts from Shopify and other labs on prompt compression.

Sources

  1. Gisting (Shopify Engineering)shopify.engineering
  2. Gisting: compressing LLM agent context (ZenML LLMOps database)zenml.io

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 4 September 2026 at 21:04 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-cost-war-moved-in-12-months-from-Dc3uAzLDdCV" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The cost war moved in 12 months from which model to how do we fit 4m tokens into a 200k window and…"></iframe>

More notes