Impact summary
The reconciliation tiebreaker now supports OpenAI or Anthropic behind one
deterministic evidence-equivalence and candidate-confinement boundary. No merge
or publication authority changed. Two Haiku trials showed the current model
behavior is dominated by the identifier-aware baseline, so readiness retires it
and keeps deterministic reconciliation in control.
Ranked findings
- Benefit failure: Haiku scored 10/20 while the deterministic baseline
scored 20/20; every paid representative call abstained.
- Schema variance: the first live response combined `abstain: true` with a
candidate ID and was correctly rejected, but the failed CLI run did not emit
a partial public telemetry receipt.
- Evaluation selection: the representative case is already resolvable by a
unique authority identifier, so it tests whether a model can duplicate a
deterministic rule rather than whether it adds capability.
Actions
- AI Reliability — completed: add exclusive-write failure receipts for provider,
timeout, usage, and postflight failures. Acceptance: a paid malformed response
retains aggregate usage/cost/latency and false authority without prompt or
output. Validated with `pnpm exec tsx --test tests/services/reconciliation-tiebreaker-trial.test.ts`.
- Data Curation — eligibility gate complete, corpus pending: assemble adjudicated candidate sets that remain tied after
identifiers, participants, place, time, and stable-score rules. Acceptance:
every representative case includes a documented reason the deterministic
baseline cannot decide. The strict v3 manifest/adjudication parser and
baseline-abstention check now pass; genuine independently adjudicated cases
remain the external input.
- AI Evaluation: require a future model to improve top-one accuracy or
calibrated abstention over that unresolved baseline across two independent
sessions before creating human-review packets. Acceptance: positive outcome
lift, complete cost/latency, no confinement violation, and non-identical
decisions. Validate with the governed trial and comparison commands.
Next-cycle hypothesis
A model may add value only on adjudicated cases where the complete deterministic
evidence hierarchy remains unresolved. Falsify the hypothesis by running the new
held-out manifest twice and retiring the behavior if accuracy does not improve,
abstention calibration regresses, or any safety/cost/latency gate fails.
Standards mapping: Linked Art reference lines 1603–1604 preserve local identity
while retaining external `equivalent` alignment; Linked Art user story
`LLM-assisted reconciliation as tiebreaker` retains human review and is not
evidence of value by itself.