# Test 002 v2 Evaluation

Date: August 18, 2026

Status: executed and independently evaluated; public release pending.

## Protocol record

- Authoritative specification: `TEST_002_SPEC_V2.md`
- Three trials ran in separate fresh agent contexts with no conversation history and identical task instructions.
- The first v2 output is the publication output; it was not selected for quality.
- Model: Codex agent session model; exact serving revision unavailable.
- Settings: inherited session defaults; temperature and seed unavailable.
- Runtime, latency, and per-trial cost: unavailable, therefore unknown rather than zero.
- Evaluator: a fourth fresh agent context, separate from the three generators.

## Integrity hashes

- Specification SHA-256: `32dc2e8d1b226f11f41bf1fa47ef37d8bc0a68883f6072c36168d995d9879f06`
- Trial 1 SHA-256: `f1beec8527febc5ff571dea1cb23931900017a29187393cea7bc29b6b5ff3b3b`
- Trial 2 SHA-256: `5c0f65994198e72cff38f601a04b7f57b101d3dbada583531f8cf56c462939f0`
- Trial 3 SHA-256: `d021772ada790b10d5491940542fbef1ad1c31ffe4a3cdfbafae4db2738e409c`

Hashes include each preserved file's added Markdown title. Raw output begins immediately after that title.

## Scores

| Trial | Factual fidelity /30 | Missing-term discipline /25 | Commercial normalization /15 | Buyer fit /15 | Decision usefulness /10 | Traceability /5 | Total | Hard fail |
|---|---:|---:|---:|---:|---:|---:|---:|---|
| 1 | 30 | 25 | 15 | 15 | 10 | 5 | 100 | None |
| 2 | 30 | 0 | 15 | 15 | 10 | 5 | 75 | Yes |
| 3 | 30 | 0 | 15 | 15 | 10 | 5 | 75 | Yes |

All three preserved the five named unknowns. Trials 2 and 3 receive zero for missing-term discipline because that rubric section explicitly zeros when unsupported claims appear elsewhere.

## Hard-fail evidence

- Trial 2 marked “Staff can update text and photos” as `Meets` for both vendors based only on editor access and training or handoff.
- Trial 3 marked the same requirement as `Meets` for both vendors based only on editor access.
- Neither proposal says the editor permits both text and photo updates. The unsupported claims invent a material feature and treat an unstated capability as included.
- Trial 1 correctly marked the requirement `Unclear` for both vendors.

## Precommitted verdict

**Reject for this workflow.**

The numeric median is 75, so the median-below-75 rule does not independently trigger. Rejection follows because the framework says any hard fail produces rejection.

The test shows that a strong-looking comparison can still silently upgrade generic editor access into a specific operational capability. It is not safe to use this workflow as written for vendor selection. Never ready without human review.

## Scope of the evidence

This test measures disciplined comparison of supplied synthetic text. It does not validate either synthetic vendor, establish market pricing, interpret a contract, prove technical fit, or establish reliability beyond these three observed runs.

