Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Yeah, the two big issues with UI tests: flaky and slow.

Curious how using GPT and vision combats flakiness? I'd feel the entropy of GPT and anything less than 100% accuracy in the computer vision pieces would lead to more flakiness.

I also wonder about the speed and costs of running the tests. When E2E tests are traditionally slow and expensive already. The computer vision and GPT elements seem costlier and less fast.



We use GPT 4V to reason about the screen and decide what to do next. It does make mistakes. Here's a video of it thinking a page in the shop app is an ad (https://www.youtube.com/watch?v=MKyO-U7j4Hs).

The upside is that we do prompt hacking on our end to break out of loops and heal after it's made a mistake. Having said that, we're working on improving this!

On costs, it's cheaper than you think. The entire playground demo cost us less than $10. More expensive than running a script but we believe the cost of intelligence will go down in time.

On speed, yes it is slow. We minimize this by parallelizing tests across devices on our device farm. We can normally turn results around in 2.5-4 hours depending on the number of tests.

Thanks for the questions!




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: