← Founder Notes
Archive

A solver-verified chinese logic benchmark found the best frontier model scores 37.5 percent on hard…

Yethikrishna ROriginal on Threads

a solver-verified chinese logic benchmark found the best frontier model scores 37.5 percent on hard items, and even with formalization help the highest joint score is 60 percent. reasoning demos look sharp until they hit tests that machines can verify.

the gap between demoed reasoning and provable reasoning is still the real benchmark.

Context

The arXiv paper LLMEval-Logic (2605.19597), May 2026, introduces a solver-verified Chinese benchmark for logical reasoning with adversarial hardening. Items are forward-authored from realistic scenarios, gold natural-language to formal-language alignment is audited, answers are verified with the Z3 solver, and there are expert rubrics. It has paired Base and Hard subsets, where Hard items chain about five counterfactual sub-questions and count as correct only if every sub-question is correct, and two evaluation axes, one being formalization grading.

How it compares

The benchmark's identity, Z3 verification and design are first-party. The 37.5 percent for the best model on hard items and the 60 percent highest joint score were not located in the text read, so the model, the axis and the meaning of joint are unverified. It is a May 2026 preprint, so the model set is as of then. The demos look sharp until line is the author's opinion.

Related work

Watch next

  • The paper's results table.

Sources

  1. LLMEval-Logic (arXiv 2605.19597)arxiv.org

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 16:36 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/a-solver-verified-chinese-logic-benchmark-found-the-DdbSePul8Av" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="A solver-verified chinese logic benchmark found the best frontier model scores 37.5 percent on hard…"></iframe>

More notes