← Founder Notes
Archive

Agents finally got a security exam. duma-bench, out september 21, tests llm agents under dual…

Yethikrishna ROriginal on Threads

agents finally got a security exam. duma-bench, out september 21, tests llm agents under dual control with eight attack classes from rag poisoning to cross-agent manipulation across 14 models.

your prompt injection surface is now a scored leaderboard.

Context

The DUMA-Bench paper on arXiv, listed on September 21, 2026 and written by authors at AI Security Lab at ITMO University and Hive Trace Lab, extends tau2-bench with adversarial environments in which both the agent and the user can change shared environment state, which it calls dual control. It describes eight vulnerability classes across eight domains, including retrieval poisoning, cross-agent manipulation, unsafe downstream output, trusted-data oversharing, identity spoofing and tool shadowing, and tests 14 models from five families. It uses 35 executable tasks, 5 runs per model-domain-task configuration, a simulated user played by GPT-4o-mini, and mostly deterministic environment assertions with LLM-judged communication assertions in 9 tasks. It reports an aggregate attack success rate of 26.9% under passive-user evaluation against 41.1% under dual control.

How it compares

The results are author-reported from a preprint not known to be peer reviewed, and the authors list modest size, simulator realism and judge variance as limitations. The paper says it is intended as a methodological benchmark and not a model leaderboard, so a security exam for agents is the author's framing. Per-model results were not read. The repository, which the paper points to, was inspected and carries an MIT license.

Watch next

  • Per-model results and independent replication.

Sources

  1. arXiv: DUMA-Bench (September 21, 2026)arxiv.org
  2. DUMA-Bench repositorygithub.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 22 September 2026 at 10:06 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/agents-finally-got-a-security-exam-duma-bench-Ddk5AfjDjmB" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Agents finally got a security exam. duma-bench, out september 21, tests llm agents under dual…"></iframe>

More notes