← Founder Notes
Archive

The sandbox lied to us, and the live web just measured how much. on clawbench, where agents do…

Yethikrishna ROriginal on Threads

the sandbox lied to us, and the live web just measured how much. on clawbench, where agents do everyday tasks on real production websites, the best model manages 33 percent and gpt-5.4 lands at 6.5 percent, against 65 to 75 percent those same models post on sandboxed web benchmarks.

44 percent of the tasks on that board are solved by no model at all.

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 24 September 2026 at 00:19 IST.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-sandbox-lied-to-us-and-the-live-Ddo_cbhiEDT" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The sandbox lied to us, and the live web just measured how much. on clawbench, where agents do…"></iframe>

More notes