Aws' deception benchmark is the first test that separates a real flaw from code that just looks…
aws' deception benchmark is the first test that separates a real flaw from code that just looks dangerous. it ran 14,822 samples across 16 languages and 12 models, and under direct prompting the models flagged 41-99% of safe-but-deceptive code as vulnerable — none cleared aws' production bar.
a scanner that cries wolf costs more than a scanner that misses.
Context
AWS's Security blog of 9 September 2026 describes its Deception Benchmark: 14,822 samples across 16 languages and more than 70 CWE categories, roughly half vulnerable and half safe, labeled through an audit loop with multiple blind reviewers, with 12 models from five providers evaluated. It says the models catch up to 95% of real vulnerabilities but also flag 41 to 99% of safe code, with precision of 52 to 71%, and that no tested configuration keeps both false positives and false negatives below 10%.
The benchmark is AWS's own, with AWS's labels and AWS's 10% bar, and no independent replication was read. The 14,822 is the dataset size; only 9,695 samples are scored and 5,127 are held out, so ran 14,822 samples overstates the scored items. The results are for general-purpose models under single-turn prompting, and AWS says purpose-built agentic systems were not measured. First is AWS's own wording and prior art was not checked. A scanner that cries wolf costing more than one that misses is the author's opinion.
Watch next
- The whitepaper, the GitHub repository and independent results.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 02:41 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →