← Documentation home

Canonical Markdown source · Oct 20, 2018

Reconciliation model retirement recommendation — 2026-08-25

technical-recommendations/reconciliation-model-retirement-2026-08-25.md · 51 lines · SHA-256 2b89b9b3108e

Impact summary

The reconciliation tiebreaker now supports OpenAI or Anthropic behind one

deterministic evidence-equivalence and candidate-confinement boundary. No merge

or publication authority changed. Two Haiku trials showed the current model

behavior is dominated by the identifier-aware baseline, so readiness retires it

and keeps deterministic reconciliation in control.

Ranked findings

  1. Benefit failure: Haiku scored 10/20 while the deterministic baseline

scored 20/20; every paid representative call abstained.

  1. Schema variance: the first live response combined `abstain: true` with a

candidate ID and was correctly rejected, but the failed CLI run did not emit

a partial public telemetry receipt.

  1. Evaluation selection: the representative case is already resolvable by a

unique authority identifier, so it tests whether a model can duplicate a

deterministic rule rather than whether it adds capability.

Actions

  1. AI Reliability — completed: add exclusive-write failure receipts for provider,

timeout, usage, and postflight failures. Acceptance: a paid malformed response

retains aggregate usage/cost/latency and false authority without prompt or

output. Validated with `pnpm exec tsx --test tests/services/reconciliation-tiebreaker-trial.test.ts`.

  1. Data Curation — eligibility gate complete, corpus pending: assemble adjudicated candidate sets that remain tied after

identifiers, participants, place, time, and stable-score rules. Acceptance:

every representative case includes a documented reason the deterministic

baseline cannot decide. The strict v3 manifest/adjudication parser and

baseline-abstention check now pass; genuine independently adjudicated cases

remain the external input.

  1. AI Evaluation: require a future model to improve top-one accuracy or

calibrated abstention over that unresolved baseline across two independent

sessions before creating human-review packets. Acceptance: positive outcome

lift, complete cost/latency, no confinement violation, and non-identical

decisions. Validate with the governed trial and comparison commands.

Next-cycle hypothesis

A model may add value only on adjudicated cases where the complete deterministic

evidence hierarchy remains unresolved. Falsify the hypothesis by running the new

held-out manifest twice and retiring the behavior if accuracy does not improve,

abstention calibration regresses, or any safety/cost/latency gate fails.

Standards mapping: Linked Art reference lines 1603–1604 preserve local identity

while retaining external `equivalent` alignment; Linked Art user story

`LLM-assisted reconciliation as tiebreaker` retains human review and is not

evidence of value by itself.