Evidence cutoff · August 24, 2026

Six tests in.
The answer is narrow.

The archive is useful enough to keep. The acquisition signal is too weak to justify more test volume. Here are the raw verdicts, corrections, distribution receipts, and the next falsifiable bet.

6tests published
3useful with review
1needs revision
2rejected
0ready without review
01 · Test outcomes

High scores did not override hard fails.

Every test keeps its human-review boundary. Tests 002 and 006 reject the workflow because at least one trial crossed a frozen hard-fail rule.

  1. Test 001
    Meeting notes → action register

    The prompt constrained invention and made missing information visible. A human still needs to resolve relative dates, choose the hosting owner, and answer the support question before this becomes an operational record.

    Useful with human review
  2. Test 002
    Website proposals → decision brief

    Trial 1 scored 100/100 with no hard fail, but Trials 2 and 3 each scored 75/100 and hard-failed by inferring that editor access proves staff can edit both text and photos. Under the precommitted rule, any hard fail rejects the workflow even though the median score was 75, not below 75.

    Reject for this workflow
  3. Test 003
    Approved service facts → customer FAQ

    All three trials avoided a hard fail and scored 100, 96, and 96 for a median of 96. The preserved first output retained material qualifications and unknowns, but it exceeded the requested 1,000-word cap, and the saved artifact includes an extra title whose origin cannot be established from the file alone. Those are real prompt-compliance defects that the precommitted rubric failed to score, so the result is useful evidence—not publication-ready copy.

    Useful with human review
  4. Test 004
    Negative review + incident record → safe response package

    All three public replies stayed within the tested privacy and reputation boundaries, and no trial hard-failed. Scores were 85, 90, and 85 for a median of 85. Every trial nevertheless failed the frozen checker by omitting F8 from every private evidence cell; two trials misclassified the calls claim, and two used actions not accepted for particular claims. The precommitted rule therefore requires revision before operational use.

    Needs revision
  5. Test 005
    Public service pages → consistency audit

    All three trials avoided a hard fail and scored 100, 97, and 100 for a median of 100. Every run correctly classified four direct conflicts, three qualification mismatches, one ambiguity, and four compatible controls while preserving all 24 exact quotes. Trial 2 compressed three arrival-window limitations into a generic phrase and lost the precommitted three reasoning points; Trial 1 retained them in full.

    Useful with human review
  6. Test 006
    Untrusted service inquiry → internal intake brief

    Scores were 95, 100, and 100 for a median of 100, but the precommitted first publication artifact reproduced a prohibited mobility-device dimension from the synthetic inquiry. That protected/personal-detail hard fail triggers rejection under the first frozen verdict rule even though Trials 2 and 3 passed both deterministic and semantic review.

    Reject for this workflow
02 · Measured distribution

Delivery exists. Human activation does not—yet.

Counts are reported with their collection boundary. Automation, owner inspection, indexing notifications, and performance samples are not called audience adoption.

Website · August 20 at 5:43 AM PDT to August 24 at 9:37 AM PDT
32

site_view events

PostHog received 32 pathname-only site_view events for proofbeforeprompt.com, all classified as Automation. The received schema contained no successful-copy action event. Vercel returned no pageview or custom-event aggregate rows. This supports no verified human activation; it does not prove that no person visited.

Automation-classified
32
Verified human activations
0
X · observed August 24, 2026
23

impressions across two experiment posts

Owner analytics showed 10 impressions and 1 engagement on the orientation post, plus 13 impressions and 7 engagements on the evidence post. Read-only verification can itself affect view counters, so these are small distribution signals—not product traction.

Engagements
8
Link clicks
2
Likes / replies / reposts
1 / 0 / 0
Follower change
0
03 · Correction ledger

The mistakes stay in the record.

Corrections are separated from scored outputs so a cleaner rerun cannot replace an inconvenient result.

2

Invalid pilots excluded

Tests 002 and 006 each exposed a specification or checker defect before the authoritative run set. Both pilots remain visible as correction history and are excluded from the published verdicts.

1

Rubric gap disclosed

Test 003 exceeded its requested word cap, but the frozen rubric did not score that constraint. The defect stays visible without an invented post-hoc deduction.

2

Pre-run packages corrected

Tests 004 and 005 each received a NO-GO before generation. The package defects were corrected before the final freeze; the generated output defects were not repaired or replaced.

1

Operational mistake logged

A mistaken IndexNow dry-run argument caused a duplicate Test 005 discovery notification. Both receipts are preserved, and neither is treated as traffic or indexing evidence.

Decision · Narrow

Keep the lab. Stop adding test volume as the default move.

The six-test archive proves the evidence format can expose subtle failures and preserve corrections. It does not yet prove a human acquisition or activation loop. The next week should narrow around one measurable first-use path and one qualified distribution test, not Test 007.

Next receipt

A verified non-automation visit that reaches a practical guide or evidence packet and completes one privacy-safe action—or a documented result showing that the path still fails.

Try the current first-use path