Evidence standard
The StrangerTest Evidence Standard
An AI user saying something happened does not make it true.
AI users are useful because they can explore products independently. They are also models operating browsers, and both the model and the browser automation can make mistakes.
StrangerTest is designed around that fact. These are the rules a user’s experience has to pass before it becomes a finding about your product.
01
Browser state outranks narration
If a user says “I submitted the form” but the recorded browser state does not show a submission, the claim is not accepted as fact.
The browser’s record says what happened. The user’s narration says what they meant, expected and made of it — which is valuable, and is never proof.
02
Our failure is not your failure
Our hand’s uncertainty, execution failures, safety refusals, bot-protection redirects and test-environment limits cannot become product findings.
They are removed from the evidence of every finding. A finding left with no evidence is set aside, and everything removed is listed in the report under “What we didn’t count”.
Why this rule exists
In the first live scan, several users reported that a product’s trial links were broken. They worked. Our click had landed on a different element. Several users meeting our defect had been counted as several witnesses.
03
“I didn’t find it” does not mean “it doesn’t exist”
Every finding says what kind of claim it is: an observed fact; this user’s experience of this visit; a findability problem, which needs users who looked in the reasonable places; or a product-wide absence, which needs at least two users who independently searched for it.
The report uses the narrowest statement the evidence supports.
Why this rule exists
A user scrolled a pricing page, didn’t see security information, and left. The report said the information was missing. It was in a collapsed FAQ on that same page, and in the footer.
04
Guessed routes are not product evidence
If an AI user guesses an address such as /security and receives a 404, that proves the guessed address was wrong. It does not prove security information is absent.
Addresses a user typed in themselves are marked as guesses, and a guess that fails cannot support a finding.
05
Hidden content is not reviewed content
Collapsed sections, tabs and menus the user never opened are not treated as inspected. Scrolling a page is not reading what its closed sections hide.
The record lists what each user actually opened. A user who opened nothing has seen the visible page, not the full page.
06
Outcomes require evidence
“I completed the setup” must match the actual browser record: the actions taken and the screen the visit ended on.
A claim to have made, submitted or completed something that the record does not support is marked as not established. It cannot support a finding, and it is never quoted.
Why this rule exists
A user claimed it had finished a multi-step diagram, with arrows. The final screen showed one unfinished shape and an empty text box.
07
Test-caused quotes do not become customer blame
A user’s sentence can be a true account of their visit while the cause was StrangerTest. If a persona’s negative statement rests on a step our own machinery caused, it does not appear as a quote about your product.
It stays in the record as their account of what our test caused. A decision or a feeling about the product is always theirs, and can be quoted.
Why this rule exists
A user said it couldn’t get text into a form’s first field. Our hand had been typing at the field’s label.
08
Independent observations matter
Several users meeting the same problem strengthens the evidence only when their experiences are genuinely independent — not when they are all consequences of the same limitation of the test.
When users met the same behaviour for what may be one shared cause, the report says so instead of counting them separately.
Why this rule exists
An earlier design sent a second, blind user to re-check single observations. Checked by hand against the live pages, the false headlines were ones that re-check had “confirmed”; issues that recurred naturally across users held up.
09
Severity is not evidence strength
How much a problem would cost you and how sure we are that it happened are two different things, and the report shows them separately.
A potentially serious issue observed once is still observed once. Its severity does not lend it certainty.
10
“We could not establish it” is a valid result
StrangerTest is allowed to abstain. When the evidence is weak, the report says so. When nothing significant came up, the report says that.
A re-test can be inconclusive, and says so rather than guessing “fixed”. A verdict of fixed on weak evidence is worse than no verdict, because you would ship believing it.
The objective is not to produce the maximum number of findings. It is to avoid making the founder debug the tester.
Every rule above is applied before you see a report, and every observation it excluded is listed in that report with the reason. See it working in the sample report, or read how the rules came about in StrangerTest Research.
How StrangerTest worksWill a stranger actually get it?
Paste your product’s address. Three independent first-time users try it while you watch. Free, and no account.