A change invisible to humans cuts prompt injection success from 61% to 10%. a september 17 paper…
a change invisible to humans cuts prompt injection success from 61% to 10%. a september 17 paper found destyling text into plain formatting collapses the model's role confusion, the mechanism behind most injection attacks.
defenses may live in typography, not sandboxes.
Context
The paper Prompt Injection as Role Confusion (arXiv 2603.12277, search metadata dated 27 June 2026, an ICML 2026 poster) says models perceive the source of text from how it is written and not from the tag-specified role, and traces prompt injection to that. A later paper, Render Before Reading: Visual Rendering as a Prompt Injection Defense (arXiv 2609.36121, 28 September 2026), proposes a rendering-based defense. A 2024 spotlighting paper (arXiv 2403.14720) marks untrusted input with formatting transforms.
The 17 September paper with destyling and the 61 to 10 percent figures was not found in three searches, so the figures are unverified and unsupported, not refuted. The June paper predates 17 September and is a different, earlier source: it supports the idea that style imitation makes injections work, and it studies attack mechanics and not a destyling defense. No evidence was seen for the mechanism behind most injection attacks. The rendering paper and the spotlighting paper are different techniques and were not read in full. Defenses may live in typography, not sandboxes is the author's take.
Watch next
- The exact title, tested models and attack set of the 17 September paper.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 01:38 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →