← Founder Notes
Archive

The cheapest performance win in production is still the prompt. shopify compressed its sidekick…

Yethikrishna ROriginal on Threads

the cheapest performance win in production is still the prompt. shopify compressed its sidekick agent's 6,000-token system prompt to 1,500 learned gist tokens with no measured quality loss, cutting end-to-end latency from 6.8 to 4.2 seconds and raising throughput 16%.

model updates are noise next to a 4x context cut.

Context

Shopify Engineering's post of 19 August 2026 describes compressing the Sidekick GraphQL agent's system prompt of about 6,000 tokens to about 1,500 learned gist tokens, a 4 to 1 ratio, using knowledge distillation where the teacher sees the full prompt and the student sees the gist tokens. At 350 requests per minute it reports median time to first token from 438 ms to 354 ms, median end-to-end latency from 6.8 s to 4.2 s, and throughput from 20.2 to 23.4 queries per second, with 14 percent fewer GPUs on that workload. It says 4 to 1 was the optimal ratio for its domain, other domains vary, and quality degraded beyond it.

How it compares

All figures are Shopify's own, under one workload at 350 requests per minute. Without losing prediction quality is the post's wording, and the quality metric read is distillation fidelity, not end-to-end task success. The 4x is the system prompt only, not total context. Gisting is trained: it needs learned embeddings for a specific model and prompt, so it is a training and serving change and not a plain prompt edit. The post itself says gisting and prefix caching compound. That the prompt is the cheapest win and that model updates are noise are the author's opinions.

Related work

Watch next

  • How the quality metric is defined, and results from other teams.

Sources

  1. Gisting (Shopify Engineering, 19 Aug 2026)shopify.engineering

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 03:19 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-cheapest-performance-win-in-production-is-still-DdfA7iJl6VY" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The cheapest performance win in production is still the prompt. shopify compressed its sidekick…"></iframe>

More notes