StrangerTest

AI user testing

AI user testing: what it can — and cannot — tell you

AI agents can now use a live product the way a first-time visitor does. That is genuinely useful, and it has limits worth knowing before you rely on it.

What AI user testing is

AI user testing sends AI agents — browser agents — through a real product as if they were people using it for the first time. Each agent is given a persona and a reason to be there, not a script. It looks at the screen, decides what to do, clicks and types, gets confused, recovers or gives up. What it did, and where it struggled, is the result.

You will also see it called synthetic user testing or AI usability testing. The terms overlap; what matters is whether the agents actually operate the product, or only talk about it.

How it differs from what you may already use

Scripted end-to-end tests

Asks
Does this route still work?
How
A script follows steps someone wrote: click this, type that, expect this.
Can’t see
Whether a newcomer would ever have taken that route. The script already knows the way.

Synthetic interviews and surveys

Asks
What might these people say or prefer?
How
A model answers questions as a simulated respondent.
Can’t see
What happens when someone actually tries the product. An opinion about a product is not an attempt to use it.

Heuristic UX audits, human or AI

Asks
Does this page follow good practice?
How
An expert, or a model, reviews screens against principles.
Can’t see
Behaviour. A page can pass every heuristic and still lose the person trying to do one specific thing.

AI user testing

Asks
What happens when a first-time user with a goal tries this?
How
AI agents with a persona and a goal use the live product in a real browser, choosing their own actions.
Can’t see
Real human motivation, emotion and money. And its agents can be wrong, so its evidence has to be checked.

These are complements, not rivals. Scripted tests guard the routes you know about; AI user testing asks whether a stranger would find them.

Where AI user testing is useful

  • Before launch, when you have no users yet and cannot watch anyone.
  • After a change to signup, onboarding or pricing, to see whether a newcomer still gets through.
  • When you are too close to the product to see the first-time experience — which, after a few weeks of building, is everyone on the team.
  • Repeatedly and cheaply: the same kind of test can run again after every fix.

Where real people remain necessary

  • Whether anyone wants the product, and whether they will pay for it.
  • Your actual conversion rate. Simulated visits do not predict it.
  • Emotional response, trust built over time, accessibility needs and lived experience.
  • Behaviour that depends on real data, real colleagues or real money.

An AI user is a model with a goal, not a customer. Treat its visit as evidence of what a first-time user could understand and accomplish — not as a forecast of what your customers will do.

Why browser evidence matters

A browser agent is also a model operating software, and both can fail. It can click the wrong element, type into a label instead of its field, be turned away by bot protection, or describe an outcome it never reached. From inside the test, each of those looks exactly like a product problem.

So the useful question about any AI user-testing result is not only “what did the agent say?” but “what does the browser record show?” A tool that reports the agent’s narration as fact will, sooner or later, blame your product for its own mistakes.

A question worth asking any tool

When the AI user fails, how do you know whether the product failed or the agent did?

How StrangerTest approaches it

StrangerTest is built on two constraints. Its users arrive cold: they know who they are and why they came, and nothing about how the product is meant to work. And a separate judge decides what counts: the users’ narration is checked against the browser record, and anything our own test caused is excluded and listed rather than reported.

Read how StrangerTest works, the evidence standard, or the difference between synthetic users that act and synthetic users that answer.

See a sample report

Will a stranger actually get it?

Paste your product’s address. Three independent first-time users try it while you watch. Free, and no account.