StrangerTest

Changelog

Changelog

Major changes to what a scan does and how far its report can be trusted.

Dates are when each change landed in StrangerTest’s codebase. Smaller fixes are not listed. The reasoning behind several entries is in Research.

  1. Public methodology, evidence standard and research

    The method, the ten evidence rules and our research notes are published.

    Why it matters You can see how a finding is judged before you trust one.

    Evidence standard

  2. Per-user identities and email verification

    Each Full Scan user whose mission needs an account signs up with its own generated identity and name. Email verification is judged after your product responds, and a user counts as signed in only when the browser has actually got inside.

    Why it matters Users no longer invalidate each other’s sign-in codes, and coverage reports what really happened.

  3. Full Scan runs the same user as the Free Scan

    Paid, re-test and follow-up users now use the same first-time-user model as the free scan. A paid user differs only in its mission and what it is allowed to do.

    Why it matters One behaviour model, tested on every free scan, now carries the paid run too.

  4. Evidence-safe founder quotes

    A user’s sentence that rests on a step our own test caused is no longer shown under “In their words”.

    Why it matters A quote in your report is never the user blaming your product for our mistake.

  5. Outcome claims checked against the browser

    A user’s claim to have made, submitted or completed something is kept only if the recorded actions and final screen support it.

    Why it matters A report cannot tell you a flow works because an AI user said it finished.

    Research note

  6. Findability is not absence; excluded evidence never supports a finding

    Each finding states what kind of claim it is. Guessed addresses, unopened sections, our hand’s uncertainty and our refusals are stripped from every finding’s evidence and listed as not counted.

    Why it matters “One user didn’t find it” can no longer be reported as “it doesn’t exist”.

  7. Grounded browser interaction, and the record outranks narration

    The user describes what it wants to act on; a grounding step locates it on screen. Typing goes into the field a label belongs to, and the judge weighs what the browser recorded above what the user said.

    Why it matters Fewer mis-aimed actions, and the ones that remain are labelled as ours.

    Research note

  8. Natural Agent user model

    Users get who they are, why they came and their screen — and decide for themselves what to do and when to leave. They are not told to find problems.

    Why it matters What they run into is what a first-time user runs into, not what a tester goes looking for.

  9. Visual-first perception

    Users reason from the rendered screen, before and after each action, instead of from the page’s hidden structure. Hover, drag and canvas interaction became possible.

    Why it matters The user only knows what a person at that browser could see.

  10. Bot protection and redirects reported as test limits

    A server turning the automated browser away, or a bot-protection check, is recorded as a limit of the test and never circumvented.

    Why it matters Your main action is not reported as broken because it blocked our browser.

  11. Email verification inside a Full Scan

    Confirmation codes and links sent by your product are read from a test inbox and used, so users get past “check your email” into onboarding.

    Why it matters The Full Scan reaches the part of the product that decides whether new users stay.

  12. Authenticated Full Scan and fix verification

    The Full Scan builds on the free scan and can sign up inside your product, with a password permission scoped to the exact generated values. Re-tests report each fix as fixed, improved, still present or inconclusive.

    Why it matters Testing continues past the signup form, and a fix is confirmed by new users who know nothing about it.

  13. Independent users, one judge

    Layers of scoring rules were replaced by independent users and one reviewer who reads every session, with the browser’s facts kept separate from interpretation.

    Why it matters A user’s genuine experience is no longer deleted by a keyword match, and findings say how sure they are in words.

  14. Automation failures cannot become findings

    Every candidate is checked for whether our automation or infrastructure caused it before it can reach a report.

    Why it matters Built the day our first live scan reported working links as broken.

    Research note

Will a stranger actually get it?

Paste your product’s address. Three independent first-time users try it while you watch. Free, and no account.