A september paper found over 50 percent inconsistency in hallucination benchmarks like med-halt,…
a september paper found over 50 percent inconsistency in hallucination benchmarks like med-halt, and argues detection tools measure consistency, not correctness, while rag can even add new errors. the field is fighting the wrong target: fluent consistent answers pass the detectors, true ones fail them.
verification is about the claim, not the confidence.
Context
The arXiv paper Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity (2602.00723) says its analysis reveals significant multiplicity, over 50% inconsistency in benchmarks like Med-HALT. It shows that on Med-HALT accuracy varies under 0.5% across prompt structure changes while more than 50% of questions get different facts depending on prompt structure, across 8 benchmarks and 16 models from six families. It says hallucination detection techniques predominantly detect consistency, not correctness, and that mitigation such as RAG introduces additional inconsistencies through prompt sensitivity.
The claims match the paper, but the arXiv number indicates February 2026 and no September version was found, so the note's September paper does not match. Inconsistency means answers change under prompt variants such as shuffled option order, not general accuracy loss, and the over 50% is for Med-HALT-type settings, not every benchmark. RAG can add new errors loosens the paper's additional inconsistencies. Fluent consistent answers pass the detectors, true ones fail them overstates it. Two September arXiv items on medical evaluation and hallucination detection surfaced in search and were not shown to be the note's source.
Related work
- Rethinking Hallucinations (EACL 2026 listing) ↗Snippet only.
- Medical hallucination detection (arXiv 2609.03953) ↗A separate paper, snippet only.
Watch next
- A September version or follow-up, if one exists.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 19:03 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →