Mastodon Feed: Post

Mastodon Feed

Boosted by pixxl:
jonny@neuromatch.social ("jonny (nonvenomous)") wrote:

So hang on. The screenshot test says I can pay 10 cents for a robot to move my mouse to a button?

That can't be right, it must be 10 cents to run the whole benchmark, there are only 1500 samples per run.

But the output tokens are all less than 1000... So either it can move my mouse with fewer than 1 token on average... Or...... 10 cents to move my mouse is really what is being shown here.............

[gpt-6 achieving 90% accuracy between 7 and 11 cents, gpt-5.5 60-80% accuracy between 3 and 4 cents] ScreenSpot-Pro tests whether models can locate the correct interface element in high-resolution screenshots of professional software. GPT‑6 Astra sets a new high with 92.7% accuracy.
[gpt-6 at 90% accuracy between 0 and 750 output tokens. gpt-5.6 at 60-80% between 100 and 1000 output tokens]
In this paper, we introduce ScreenSpot-Pro, a novel GUI grounding benchmark that includes 1,581 instructions, each in a unique screenshot. They are sourced from 23 applications in 5 types of industries, as well as common usages in 3 operating systems