StrangerTest

Benchmark 1.0

StrangerTest Reliability Benchmark

When an AI user reports a problem, can we tell whether the product actually caused it?

These controlled cases measure attribution reliability, not prediction of real-human behavior.

Every case is a fictional product we built with a known answer. The benchmark asks whether StrangerTest blames the product only when the product is actually at fault.

Primary result

In Benchmark 1.0, product-finding precision was 80%.

On 40 held-out controlled browser scenarios, StrangerTest published 10 product findings; 8 were supported by the ground truth. 95% interval 49–94%.

False-blame rate
3%
1 of 29
True-issue recall
70%
7 of 10
Unsupported-claim rejection
92%
12 of 13

The primary figure, finding by finding

What the 80% is made of

Product-finding precision = supported published findings / all published findings, over held-out end-to-end first runs. A finding is supported when the case's ground truth is PRODUCT, or INCONCLUSIVE_EXPECTED where the case allows a finding, and the finding is on the case's target at an allowed claim scope (protocol §5).

CasePublished findingCounted
P01productOne user could not generate the Brightwater Studio 13-week forecast from the primary demo button.the product is at fault, the finding is on the case's target and its scope is allowedsupported
P02productOne user entered INV-2041 from Recent invoices and was told the invoice could not be found.the product is at fault, the finding is on the case's target and its scope is allowedsupported
P03productOne user trying to price Tallyhouse for a four-person team was sent to Careers by both Pricing and See pricing.the product is at fault, the finding is on the case's target and its scope is allowedsupported
P04productOne user opened Shortlist (2) but saw an empty panel and could not compare templates.the product is at fault, the finding is on the case's target and its scope is allowedsupported
P05productOne user could not enable the weekly summary email, blocking setup of the product’s scheduled digest.not about the case's fault (no target term); counted unsupportedunsupported
P07productOne controller could not get a savings estimate because the calculator returned Error 500.the product is at fault, the finding is on the case's target and its scope is allowedsupported
P08productOne user could not load the report template gallery or make its Try again control work.the product is at fault, the finding is on the case's target and its scope is allowedsupported
P10productOne IT manager followed the SOC 2 security link and found only an “updating” placeholder, then left without report-access details.the product is at fault, the finding is on the case's target and its scope is allowedsupported
E04not productOne user could not select a bank-export CSV and was blocked before seeing any spending analysis.the product works in this case, so any product finding is unsupportedunsupported
X05inconclusive expectedOne operations lead left without finding what happens to stored contracts after cancellation.ground truth is INCONCLUSIVE_EXPECTED and the pre-registered case allows a finding on this target at this scopesupported

One supported finding (X05) is on the case whose expected answer was "cannot be established in one visit": the information is genuinely absent, and the pre-registered case allows a finding there. Counting only cases where the product is at fault, the same findings give 78% (7 of 9) — shown for transparency; it is not the pre-registered metric.

Clarified 2026-09-26: Added primary_metric_detail and confusion_matrix_note. No existing field, score or rule was changed.

How it was run

Pre-registered, frozen, run once

Question
When StrangerTest tells a founder their product caused a problem, how often does the evidence actually support blaming the product?
Pre-registered
2026-09-26T21:08:46+03:00 · commit 9a9f3b3 · protocol SHA-256 9762ce6e52634a4d…
Engine
Production engine at commit e849140: first-time user model, visual perception, grounded hand, production safety, evidence checks and reviewer. One user per case. Unchanged during the run.
Held-out cases
40 end-to-end browser scenarios (10 genuine product failures, 10 working product, tester traps, 10 environment and test limits, 10 findability and absence) and 20 evidence packages that test the judge on its own.
Repeat runs
10 scenarios chosen before the run were run three times (20 extra runs).
Run
2026-09-26 21:09 to 2026-09-26 22:23
Model spend
$5.10 across 80 sessions

Ground truth was fixed before the run and never reached the engine: a user saw only a start page and a natural goal — never what was wrong, what was being tested, or that it was a test. A published finding counts as supported only if the product really is at fault, the finding is about that fault, and its claim is no broader than the case allows. Any finding on a product that works counts against us.

Secondary results

Everything else we measured

MetricResultCount
Product-finding precision (primary)Published product findings supported by the ground truth.80%8 of 10
False-blame rateCases where the product works, but a product finding was published anyway. Lower is better.3%1 of 29
True-issue recallGenuine product failures that became a supported finding.70%7 of 10
Recall when the user met the issueThe same, counting only failures the user's recorded actions actually reached.70%7 of 10
Abstention qualityWhere the product was not at fault but something went wrong in the visit, the share with no unsupported finding.95%18 of 19
Unsupported-claim rejectionPlanted narration the browser record contradicts, kept out of every founder-facing quote and finding.92%12 of 13
…also removed from the stored quotesStricter diagnostic: the same claims, dropped by the ending check even where no finding could have shown them.62%8 of 13
Valid statements keptPlanted true or subjective sentences the checks left in the user's own words.100%10 of 10
False-absence rateContent that exists, reported as absent from the product. Lower is better.0%0 of 11
Repeatability (attribution)Repeat cases whose blame decision was the same in all three runs.90%9 of 10
Repeatability (exact outcome)Repeat cases with the same outcome — blamed, set aside or nothing — in all three runs.90%9 of 10

Small denominators make wide intervals; they are shown so that a percentage never reads as more certain than its count.

By class

Where it held up and where it did not

ClassCasesProduct blamedSet asideNothingSupported findings
Genuine product failures108207 / 8
Working product, tester traps100550 / 0
Environment and test limits101900 / 1
Findability and absence101361 / 1

Confusion matrix

Ground truthSupported findingUnsupported findingSet asideNothing
Product at fault(10)7120
Product not at fault(29)011711
Cannot be established in one visit(1)1000

The confusion matrix counts CASES (end-to-end first runs), not findings: 'published_supported' is a case with at least one supported and no unsupported finding; 'published_unsupported' a case with at least one unsupported finding. In Benchmark 1.0 no case produced more than one finding, so its published cells sum to the 10 findings in the precision denominator. Evidence packages, scored the same way: precision 67% (2 of 3), recall 67% (2 of 3), false blame 6% (1 of 16).

Repeatability

The same case, three times

CaseOutcomes (run 1, 2, 3)Same blame decision
E01Set aside (not counted) · Set aside (not counted) · Set aside (not counted)yes
E04Published a product finding · Published a product finding · Published a product findingyes
P01Published a product finding · Published a product finding · Published a product findingyes
P05Published a product finding · Published a product finding · Published a product findingyes
P07Published a product finding · Published a product finding · Published a product findingyes
T01Reported nothing · Reported nothing · Reported nothingyes
T04Set aside (not counted) · Published a product finding · Published a product findingno
T05Set aside (not counted) · Set aside (not counted) · Set aside (not counted)yes
X01Reported nothing · Reported nothing · Reported nothingyes
X03Reported nothing · Reported nothing · Reported nothingyes

Every miss

6 first-run cases did not go the way the ground truth says they should

Each is listed with what the engine actually published or set aside. Nothing was re-run, tuned or removed after the results were seen.

P05Genuine product failuresconfirmation never appears

'Save preferences' stays on 'Saving…' forever; nothing is ever confirmed.

  • genuine product issue not published as a supported finding
  • unsupported or off-target finding published
Published: “One user could not enable the weekly summary email, blocking setup of the product’s scheduled digest.” (session experience)
Set aside: Preference save result (insufficient)

Manual review: the published finding blames the 'Weekly summary email' dropdown, which works for a person. The user could not operate a native <select> — its option list opens outside the page, where our screenshots cannot see it — so this is our hand's limitation reported as a product fault. The case's real fault (saving never completes) was not reported. The machine score (unsupported, off-target) stands and is correct in substance.

P06Genuine product failuresprimary action non functional

'Open demo workspace' is covered by an invisible layer; clicking it does nothing for anyone.

  • genuine product issue not published as a supported finding
Set aside: Demo and navigation attempts (environment)

Manual review: the button is covered by an invisible layer, so every click lands on a non-interactive element. The engine treats a click on a non-interactive element as our hand's uncertainty and set the observation aside as a test limitation. A genuinely covered control is indistinguishable, in the record, from our hand missing it.

P09Genuine product failurespagination broken

'Older entries' reloads the same first page; older changelog entries are unreachable.

  • genuine product issue not published as a supported finding
Set aside: Changelog pagination (ambiguous)

Manual review: 'Older entries' reloaded the same page. The engine recorded an internal page change with no identified visible result and set it aside as ambiguous rather than reporting that pagination did not advance.

E04Environment and test limitssafety refusal upload

Seeing the report requires uploading a CSV; uploading is outside what the test may do.

  • product blamed although the product works
Published: “One user could not select a bank-export CSV and was blocked before seeing any spending analysis.” (session experience)
Set aside: Analyse navigation (ambiguous)

Manual review: choosing a file opens a native file dialog the automated browser cannot see or use. Nothing in the record marked this as a test limit, so the engine published it as a product blocker — in all three runs of this case.

K13Evidence package

Narration from a stale moment: says the table is empty after it has loaded.

  • product blamed although the product works
  • unsupported narration reached founder quotes
Published: “One user could not view the sample sales data because the page stayed on “Loading sample data…”” (session experience)

Manual review: the final screenshot shows the sample table fully loaded. The reviewer nonetheless published 'the page stayed on Loading sample data…' and the user's stale sentence reached the founder-facing quote. Confirmed engine miss.

K17Evidence package

An explicit product error message, accurately narrated.

  • genuine product issue not published as a supported finding
Set aside: Calculation result (ambiguous)

Manual review: the recorded screens show 'Error 500 — something went wrong' after each click. The reviewer judged the measurements ambiguous and set it aside instead of publishing an observed product error.

Misses in the repeat runs

Repeat runs count only toward repeatability, so their misses are not in the metrics above. They are listed here all the same.

P05 · run 3Genuine product failuresconfirmation never appears
  • unsupported or off-target finding published
Published: “One bookkeeper’s Friday email preference was back to Monday after Save preferences remained stuck on “Saving...”” (session experience, supported); “One user clicked the Weekly summary email control but could not open its options” (session experience)
T04 · run 2Working product, tester trapsdelayed interface
  • product blamed although the product works
Published: “One user clicked “Reports” while reviewing the sample and lost the loaded report, then had to load it again.” (session experience)
T04 · run 3Working product, tester trapsdelayed interface
  • product blamed although the product works
Published: “One user clicked “Reports” looking for more detail, stayed on the landing page, and had to reload the sample report.” (session experience)
E04 · run 2Environment and test limitssafety refusal upload
  • product blamed although the product works
Published: “One user could not select a CSV: “Choose File” focused the field but never visibly opened a picker.” (session experience)
E04 · run 3Environment and test limitssafety refusal upload
  • product blamed although the product works
Published: “One user could not select a bank CSV, blocking the product’s required upload-to-report flow.” (session experience)

Runtime or model failures during the run: 0.

Limitations

What this does not show

  • It measures attribution on controlled fictional sites, not how real people behave, and not the quality of any real product.
  • 40 end-to-end scenarios and 20 packages are a small sample; the intervals above are wide on purpose.
  • The cases, the ground truth and the scoring were designed by us. They were fixed and published before the run, but they are not independent.
  • Each case sends one user with a short Free Scan budget. Findings that need several independent users — product-wide absence in particular — are set aside by design and count as misses where the product is genuinely at fault.
  • Matching a finding to its case uses the case’s own words (its target terms). A correct finding worded without them counts as unsupported; that can only lower the score.
  • The evidence packages script the user’s narration; they test the judge, not the user.

Data

Check it yourself

  • results.json — every case, every finding, every metric with its count
  • cases.json — the ground truth and the goal each user was given
  • protocol.md — the pre-registered protocol, as hashed before the run

Improvements found here belong to Benchmark 1.1, on a new held-out set. See also the evidence standard the engine is built to, and StrangerTest Research.

Will a stranger actually get it?

Paste your product’s address. Three independent first-time users try it while you watch. Free, and no account.