Hacker Newsnew | past | comments | ask | show | jobs | submit | mpavlov's commentslogin

You mean all the possible states of the each element of the interface (buttons, forms, content blocks, etc.)?


No, not to that level of detail. There's a higher-level set of states that all other states should be derivable from.


Unfortunately, mainly because there're only self-reported and partial numbers for the benchmarks that matter the most.


> If we really want to benchmark the ability of models to use human UIs to solve problems, then perhaps we need to choose benchmarks that don’t have APIs available such that the model cannot get “creative” in any way and must use the UI as part of the task. More simply, maybe the model isn’t the problem; maybe the benchmark designer is.

That's a valid point, yet it's hard to blame authors of OSWorld and ALE. They created an env for benchmarking long horizon task completion to be as close to real computer as possible. And for this goal CLI/API access is generally useful, yet when the model not defaults to it for the majority of subtasks.

There're benchmarks that would measure UI literacy (Webgames Benchmark is one). But they are far from the task we want to benchmark in the end.


Sure, “blame” is a strong word. My point isn’t that anyone deserves actual blame, but just that the models are proving creative. We already know that they will comment out unit tests to make them “pass,” for instance. If we want to focus on a particular skill or behavior, we need the tests to highly constrain the model. If we don’t create our benchmarks like that, shame on us.


There's an anecdotal paper 'How We Broke Top AI Agent Benchmarks: And What Comes Next' https://moogician.github.io/blog/2026/trustworthy-benchmarks...


Haha, one day, one day...


(author of PokerBattle is here)

Well, you're not wrong :) Vercel is not the one to blame here, it's my skill issue. Entire thing was vibecoded by me — product manager with no production dev experience. Not to promote vibecoding, but I couldn't do it myself the other way.


sorry i was mean


(author of PokerBattle here)

You right, results and numbers are mainly for entertainment purposes. This sample size would allow to analyze main reasoning failure modes and how often they occur.


(author of PokerBattle here)

Haven't seen it before, thanks Are you affiliated with them?


(author of PokerBattle here)

I think it would've completely crush them (like any other solver-based solution). Poker is safe for now :)


(author of PokerBattle here)

I noticed the same and think that you're absolutely right. I've thought about adding their current hand / draw, but it was too close to the event to test it properly.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: