The cheapest performance win in production is still the prompt. shopify compressed its sidekick…
the cheapest performance win in production is still the prompt. shopify compressed its sidekick agent's 6,000-token system prompt to 1,500 learned gist tokens with no measured quality loss, cutting end-to-end latency from 6.8 to 4.2 seconds and raising throughput 16%.
model updates are noise next to a 4x context cut.
Context
Shopify Engineering's post of 19 August 2026 describes compressing the Sidekick GraphQL agent's system prompt of about 6,000 tokens to about 1,500 learned gist tokens, a 4 to 1 ratio, using knowledge distillation where the teacher sees the full prompt and the student sees the gist tokens. At 350 requests per minute it reports median time to first token from 438 ms to 354 ms, median end-to-end latency from 6.8 s to 4.2 s, and throughput from 20.2 to 23.4 queries per second, with 14 percent fewer GPUs on that workload. It says 4 to 1 was the optimal ratio for its domain, other domains vary, and quality degraded beyond it.
All figures are Shopify's own, under one workload at 350 requests per minute. Without losing prediction quality is the post's wording, and the quality metric read is distillation fidelity, not end-to-end task success. The 4x is the system prompt only, not total context. Gisting is trained: it needs learned embeddings for a specific model and prompt, so it is a training and serving change and not a plain prompt edit. The post itself says gisting and prefix caching compound. That the prompt is the cheapest win and that model updates are noise are the author's opinions.
Related work
- Earlier note on Shopify gist tokens ↗Same Shopify gisting post.
Watch next
- How the quality metric is defined, and results from other teams.
Sources
- Gisting (Shopify Engineering, 19 Aug 2026)shopify.engineering
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 03:19 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →