← Founder Notes
Archive

Aws' deception benchmark is the first test that separates a real flaw from code that just looks…

Yethikrishna ROriginal on Threads

aws' deception benchmark is the first test that separates a real flaw from code that just looks dangerous. it ran 14,822 samples across 16 languages and 12 models, and under direct prompting the models flagged 41-99% of safe-but-deceptive code as vulnerable — none cleared aws' production bar.

a scanner that cries wolf costs more than a scanner that misses.

Context

AWS's Security blog of 9 September 2026 describes its Deception Benchmark: 14,822 samples across 16 languages and more than 70 CWE categories, roughly half vulnerable and half safe, labeled through an audit loop with multiple blind reviewers, with 12 models from five providers evaluated. It says the models catch up to 95% of real vulnerabilities but also flag 41 to 99% of safe code, with precision of 52 to 71%, and that no tested configuration keeps both false positives and false negatives below 10%.

How it compares

The benchmark is AWS's own, with AWS's labels and AWS's 10% bar, and no independent replication was read. The 14,822 is the dataset size; only 9,695 samples are scored and 5,127 are held out, so ran 14,822 samples overstates the scored items. The results are for general-purpose models under single-turn prompting, and AWS says purpose-built agentic systems were not measured. First is AWS's own wording and prior art was not checked. A scanner that cries wolf costing more than one that misses is the author's opinion.

Watch next

  • The whitepaper, the GitHub repository and independent results.

Sources

  1. The state of AI for security (AWS Security Blog, 9 Sep 2026)aws.amazon.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 02:41 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/aws-deception-benchmark-is-the-first-test-that-Dde8gZ-kWIq" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Aws' deception benchmark is the first test that separates a real flaw from code that just looks…"></iframe>

More notes