Three right diagnoses, three times worse
We have a tool, Lens, that opens our pages in a real browser and checks that they look correct. It took screenshots of the page and elements on the page.
Then an error occurred that it couldn’t see: in our own chat, the view jumped on its own without anyone touching anything.
A picture cannot show movement. We read the code instead, found an explanation that fit three times — and made it worse all three times. We didn’t lack hours. We lacked an instrument.
What we built
Lens can now record. But it’s not the video that finds the error. It’s a measurement: did the page move without anyone touching it? A jerk after a click is just a page scrolling. A jerk without one is an error. Only afterwards are the images clipped out — to look at, not to search through.
We deliberately did not give 600 images to an AI and ask if it saw anything. A model asked to find something will find it — even when there’s nothing there. The most important feature is that it can say no.
What it delivered
Five hours of code reading gave three wrong answers. A hundred seconds of recording found the error on the first try — plus one we hadn’t seen.
It doesn’t record your screen, but runs its own browser.