The safety gatekeeper is the weak link. check point's puzzlemask, disclosed september 13, hides…
the safety gatekeeper is the weak link. check point's puzzlemask, disclosed september 13, hides policy-violating payloads in plain english and slipped past all four commercial filters it tested, with 100% bypass and over 90% payload extraction.
cheap filters fail exactly when it matters.
Context
Check Point Research's PuzzleMask post of 10 September 2026 describes wrapping a hidden payload in plain-English prose so that a quick LLM policy gatekeeper marks it safe while a stronger target model can recover it. They tested 23 crafted prompts from an automated LLM-assisted pipeline against four gatekeeper models, gpt-4o-mini-2024-07-18, gpt-oss-safeguard:20b, claude-3-haiku-20240307 and llama-guard3, with llama-guard3 tested on 5 prompts, and the post argues for defense in depth such as monitoring outputs or chain of thought.
The 100 percent bypass and over 90 percent extraction are first-party but are small-sample results from the researchers' own tests on one pipeline. Disclosed September 13 is not supported, since the post and Check Point's blog are dated 10 September. The four are LLMs used as quick gatekeepers, including open-weight models and older versions, and the post does not call them commercial filters or products, so this is not a test of vendors' deployed guardrail products. The safety gatekeeper is the weak link is the author's framing.
Watch next
- Any vendor response, larger-scale tests and Check Point's list of affected products.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 11:37 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →