← Founder Notes
Archive

A september paper found over 50 percent inconsistency in hallucination benchmarks like med-halt,…

Yethikrishna ROriginal on Threads

a september paper found over 50 percent inconsistency in hallucination benchmarks like med-halt, and argues detection tools measure consistency, not correctness, while rag can even add new errors. the field is fighting the wrong target: fluent consistent answers pass the detectors, true ones fail them.

verification is about the claim, not the confidence.

Context

The arXiv paper Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity (2602.00723) says its analysis reveals significant multiplicity, over 50% inconsistency in benchmarks like Med-HALT. It shows that on Med-HALT accuracy varies under 0.5% across prompt structure changes while more than 50% of questions get different facts depending on prompt structure, across 8 benchmarks and 16 models from six families. It says hallucination detection techniques predominantly detect consistency, not correctness, and that mitigation such as RAG introduces additional inconsistencies through prompt sensitivity.

How it compares

The claims match the paper, but the arXiv number indicates February 2026 and no September version was found, so the note's September paper does not match. Inconsistency means answers change under prompt variants such as shuffled option order, not general accuracy loss, and the over 50% is for Med-HALT-type settings, not every benchmark. RAG can add new errors loosens the paper's additional inconsistencies. Fluent consistent answers pass the detectors, true ones fail them overstates it. Two September arXiv items on medical evaluation and hallucination detection surfaced in search and were not shown to be the note's source.

Related work

Watch next

  • A September version or follow-up, if one exists.

Sources

  1. Rethinking Hallucinations (arXiv 2602.00723)arxiv.org

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 19:03 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/a-september-paper-found-over-50-percent-inconsistency-DdbjTg1CJJ1" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="A september paper found over 50 percent inconsistency in hallucination benchmarks like med-halt,…"></iframe>

More notes