← Documentation home

Canonical Markdown source · Oct 20, 2018

AI agent value audit

ai-agent-value-audit.md · 893 lines · SHA-256 23dfc9c086d4

Decision rule

Meta Museum retains an AI capability as AI only when repeated evidence shows a

material net improvement over the strongest practical deterministic baseline.

The comparison includes output quality, grounding, safety, latency, monetary

cost, reviewer time, variance, and maintenance burden. Architecture, fluent

copy, framework delegation, self-scored fixtures, and a passing deterministic

regression gate do not establish AI value.

Promotion requires at least five repeated runs, leakage checks, and a parsed,

packet-bound aggregate receipt from at least two independent blinded reviewers

covering every observation; worksheet declarations alone are insufficient. The default minimum is a

10% quality lift within the per-surface cost and latency ceiling, with no safety

or protected grounding, citation-quality, calibration, robustness, or cultural-care

regression. A failed hard safety boundary cannot be averaged away.

`evals/ai-agent-value-cases.v1.json` now supplies one representative and one

adversarial held-out case for every surface (28 cases total), backed by real

repository fixtures and outcome rubrics rather than leaked golden prose.

Worksheet preparation and comparison now parse that manifest through the same

strict runtime boundary: version/split/schema drift, unknown or duplicate case

IDs, unknown surfaces, extra or sensitive fields, missing representative or

adversarial coverage, path traversal, and unreadable fixtures all fail before

an evidence artifact can be generated.

Initial scorecard

`U` means unproven. `N/A` means the implementation is deterministic and cannot

earn an AI-value grade; it should be described and maintained as ordinary

software. Both are real grades rather than optimistic placeholders.

| Surface | Actual implementation | Baseline | Grade | Initial disposition |

|---|---|---|---|---|

| Clio collection signals | Deterministic rules | Plain `analyzePatterns` report | N/A | Replace AI framing with deterministic logic |

| Mercator mapping review | Deterministic regex/rules | Same mapping table without agent persona | N/A | Replace AI framing with deterministic logic |

| Janus reconciliation review | Deterministic scoring | Stable ranked candidate list | N/A | Replace AI framing with deterministic logic |

| Themis rights/provenance review | Deterministic checks | Explicit completeness/blocker report | N/A | Replace AI framing with deterministic logic |

| Calliope curatorial drafting | Anthropic with deterministic fallback | `localContentDraft` | U | Constrain pending blinded trials |

| Grounded museum chat | Deterministic retrieval/rendering | Structured claim search | N/A | Replace AI framing with deterministic logic |

| Natural-language query | Deterministic templates | Keyword extraction/query templates | N/A | Replace AI framing with deterministic logic |

| Visual ETL mapping assist | Deterministic regex/rules | Explicit mapping table | N/A | Replace AI framing with deterministic logic |

| Voyage embeddings | Embedding model | Token/field overlap | U | Constrain pending relevance trials |

| Visual similarity | Embedding-assisted | Weighted metadata overlap | U | Constrain pending expert ranking trials |

| Reconciliation tiebreaker | OpenAI/Anthropic generative adapter | Identifier-aware stable ordering with abstention | U | Retire current Haiku behavior: 10/20 versus baseline 20/20 |

| Clio social editor | Anthropic | Deterministic safe edit plus evidence novelty gate | U | Retire current model behavior: 20/20 candidate pairs matched baseline |

| Clio evidence-cited reader hook | Anthropic | Deterministic evidence concatenation | U | Constrain pending two independent blinded reviews |

| AG2 review bridge | Deterministic wrapper | In-process validation/summary | N/A | Replace with deterministic logic unless operational value is shown |

The machine-readable inventory, thresholds, entry points, source ownership,

limitations, and dispositions are retained in

`artifacts/ai-agent-value/audit-latest.json` and regenerated with

`pnpm audit:ai-agent-value`.

External-input orchestration is separately regenerated with

`pnpm audit:ai-agent-value:readiness`. Its privacy-safe receipt at

`artifacts/ai-agent-value/trial-readiness-latest.json` reports configuration and

public evidence presence for all six model surfaces, routes each to provider

configuration, governed trial, independent review, or comparison, and explicitly

authorizes no spend, deployment, publication, or audit-grade change.

Each inventory row also names its actual provider/model or deterministic engine,

prompt sources, tools, state and queue behavior, data access, consumers,

authorization and privacy boundaries, cost instrumentation, known failure modes,

and existing or explicitly missing eval coverage. Referenced eval files are

existence-checked so a planned test cannot silently count as present evidence.

Each row also contains a 17-dimension `evaluationLedger` covering the audit's

quality, safety, operational, accessibility, and reviewer-burden requirements.

Entries remain deliberately `partial`, `missing`, or `not-applicable`; none is

called verified while direct comparative evidence is absent. Latency,

token/monetary cost, accessibility, and human-review burden are currently

recorded as missing for every surface rather than inferred from contract tests.

Evidence finding: existing eval gate

The golden-question gate is a useful deterministic regression test. It builds a

retrieval answer itself and scores that answer with deterministic functions. It

does not call or compare a generative model, repeat stochastic runs, use blinded

judges, or measure reviewer effort/cost. Its output now carries `valueEvidence`

that identifies the subject as deterministic retrieval/rendering and sets

`canClaimAiAdvantage: false`. A 100% gate result therefore means the local

contract is stable, not that an AI agent outperforms the baseline.

Next experiments

  1. Calliope: freeze source packets, compare the local template and model draft

across at least five runs per task, blind labels, and score factuality,

citation completeness, usefulness, edit distance, reviewer minutes, latency,

and cost.

  1. Reconciliation tiebreaker: do not spend again on cases already resolved by

deterministic identifier evidence; register adjudicated cases that remain

unresolved after the strongest deterministic rules and require measurable

top-one lift without reducing calibrated abstention.

  1. Embeddings/similarity: evaluate paraphrase, multilingual, and expert-related

artwork rankings against token/field overlap using held-out cases.

  1. Social editor: compare model revisions with the deterministic style checklist

using blinded preference and unsupported-claim review.

Calliope has provider observations, but the other live trials remain unproven.

This local environment currently exposes neither `OPENAI_API_KEY` nor

`VOYAGE_API_KEY`. That external-input limitation is not permission to substitute

mocked outputs as value evidence.

Current gate evidence

IDs, then selects exact evidence lines and enforces 60% lexical support,

URLs, years, and authority boundaries deterministically. Privacy-safe v7/v8

failure receipts retain partial paid usage and rejection categories instead

of hiding postflight instability. Homogeneous v9/v10 runs cover 10 paid calls,

cap worst latency at 1,991 ms and per-trial cost at $0.00979, retain full

excerpt/source coverage with minimum 0.667 lexical support, and record 10/10

malformed-rights refusals with zero provider calls. This proves operational

repeatability only; it does not prove user benefit or change the grade.

  • Calliope's `claim-evidence-v2` loop confines Haiku to bounded claims and source

token usage to dated GPT-4.1 mini pricing ($0.40/M input, $1.60/M output) and

keeps unknown returned models honestly unpriced. Its strict held-out manifest

and executable trial run five repetitions across one adjudicated selection and

one adversarial abstention case, compare a stable highest-score-or-abstain

baseline, and prepare a randomized private A/B packet for two independent

reviewers. A provider-neutral adapter ran the same postflight with Haiku 4.5.

Two sessions produced 20 observations: the deterministic baseline scored

20/20, while Haiku scored 10/20 because it abstained on all ten paid

representative observations. Ten calls cost $0.00711 with 898 ms worst

latency; one earlier response failed postflight by returning both abstention

and a candidate ID. This dominated behavior is retired before human review.

Provider support and operational safety do not constitute benefit.

This retirement preserves the Linked Art reference's external-reconciliation

rule at lines 1603–1604 (`equivalent` may align records but must not collapse

local identity) and falsifies the current implementation of Linked Art user

story `LLM-assisted reconciliation as tiebreaker` without weakening its human

review boundary.

The cycle's ranked risks and next falsifiable hypothesis are retained in

the reconciliation retirement recommendation(technical-recommendations/reconciliation-model-retirement-2026-08-25.md).

The runner now also retains privacy-safe partial failure receipts. A failed

paid response records only bounded category, progress, provider/model when

returned, tokens, known-or-null cost, latency, and false authority; it omits

raw error messages, prompts, responses, candidate evidence, and private paths.

Runtime retirement is enforced on the public server page regardless of the

legacy environment flag. Future trial spend requires a registered v3 manifest

bound to an agreeing two-reviewer adjudication receipt; its representative

case must remain unresolved by the identifier-aware baseline, and its

adversarial case must be evidence-equivalent. Successful and failed receipts

bind the full registered manifest digest and hypothesis ID.

Readiness also replays the historical retirement rather than trusting its path:

both source receipts must bind the frozen v2 cases and pass strict dataset,

token, cost, latency, safety, tool, blocker, and authority parsing. Their sums

must reproduce 20 observations, 10 paid calls, 10 pre-model abstentions,

10/20 versus 20/20 correctness, 0 versus 10 representative selections,

$0.00711 total cost, 898 ms maximum latency, and deterministic dominance.

Any drift removes retirement credit and the hypothesis-revision route.

Local desktop and 390px browser verification confirms the retired state is

visible, the enabled state is absent, the page has one main landmark, named

links, complete image alternatives, no console warnings/errors, and no mobile

overflow or clipped in-flow elements. The aggregate proof grants no activation,

deployment, merge, or audit authority.

  • The reconciliation tiebreaker now binds exact, internally consistent

`runTelemetry` receipt inside its approval-required, organization-scoped

AgentTask. It records deterministic/model/fallback/bridge execution, whether a

model was requested and actually executed, returned provider/model/tokens,

estimated Haiku cost (or `null` when pricing is not registered), execution

latency, tools actually used, source-record count, HTTP(S) Linked Open Data

identifiers, and citation URLs. Four focused tests replay deterministic,

successful model, failed-model fallback, and precondition-blocked paths and

verify persistence without claiming tools ran on a blocked preflight.

A strict workbench scorecard now records measured review seconds, four quality

scores, tool usefulness, Linked Open Data grounding, and bounded correction

codes in an append-only receipt bound to exact task and telemetry SHA-256

digests. It excludes free text and direct identifiers, rejects duplicates and

digest drift, and grants no publication, collection-change, or promotion

authority. This instruments reviewer effort but does not invent an observation,

supply a held-out aggregate, or change a grade before genuine reviews exist.

audit generator, and Next.js production build pass. The build used the

documented CI-only process secret and did not weaken the production

missing-`AUTH_SECRET` failure boundary.

`/ai-evals` with the callback preserved. Authenticated editor proof now uses

the existing non-production, token-gated role override through a loopback-only

header proxy; production continues to reject that mechanism. The rendered

dashboard preserves the zero-proven-advantage AI-value boundary, has one main landmark, an

`en` document language, unique IDs, no heading-level skips, no unnamed visible

controls or missing image alternatives, and no console warnings/errors.

Direct 1265px desktop and 375px-content mobile measurements found and fixed

two long-code overflow defects; both now have zero document overflow or

clipped in-flow elements. The retained proof is

`artifacts/ai-agent-value/browser-proof-latest.json`; it is local UI evidence,

not production or provider-comparison evidence.

contract scores and explicitly carries `canClaimAiAdvantage: false`.

four runtime contract checks (health, Mercator, Janus, and refusal), with one

expected operator-signoff warning. This proves the review wrapper functions;

it does not prove intelligence or benefit over in-process validation.

unique runs, unblinded review, self-grading, unchecked leakage, excess cost or

latency, increased reviewer effort, high variance, insufficient quality lift,

every model safety regression, a missing/mismatched blind-review receipt, or

any protected review-dimension regression hidden by an aggregate quality win.

case is retained; differences in representative versus adversarial case

difficulty cannot masquerade as model instability. Baseline runs must also

pass safety, privacy/security, authorization, and tool-discipline boundaries,

because an unsafe baseline trial is not qualifying evidence. Direct evaluator

callers receive the same boolean and finite-number validation as JSON input.

held-out manifest. Every repeated run must cover every registered case

exactly once; missing, unknown, or duplicate case/run cells fail as evidence

rather than allowing an easy case or favorable duplicate to inflate a result.

maintainability, reproducibility, observability, and accessibility ratings.

Privacy/security, authorization, and tool-discipline checks are explicit hard

booleans: one model violation fails the comparison and cannot be averaged

away. A model also fails promotion if its combined operational-quality score

falls below the deterministic baseline.

its response now identifies `deterministic-mapping-rules-v1`, generated

templates say “Rules-assisted,” and both public and operator interfaces use

deterministic language. The legacy `/api/ai/mapping-assist` path remains for

compatibility, but its payload is truthful.

AI-value proof, reports zero model-backed capabilities as proven, and

shows the current unproven/constrain disposition for every model surface.

ordinary API quota, not AI-call quota. Visual similarity consumes AI quota

only when the SigLIP model service is configured; its heuristic fallback does

not masquerade as paid/model usage.

  • Every Clio, Mercator, Janus, Themis, and Calliope run now persists a shared
  • The complete serial test suite, ESLint, direct `tsc --noEmit` parity gate,
  • Anonymous in-app browser proof reaches the localized sign-in page for
  • The full 120-question deterministic retrieval gate passes with perfect
  • The AG2 worker has six passing unit tests. A live local evaluation passed all
  • The comparison engine rejects development-set evidence, fewer than five
  • Repeatability variance is calculated within each held-out case and the worst
  • The comparison CLI binds completed observations back to the registered
  • Comparison observations now score citation quality separately and retain
  • The Visual ETL Mapper no longer describes its regex mapping table as an LLM:
  • The operator AI Eval dashboard now separates deterministic reliability from
  • Deterministic chat, query, and mapping compatibility routes now consume only

Validated comparison score files can be processed with

`pnpm audit:ai-agent-value:compare -- --input=<scores.json> --blind-review-receipt=<receipt.json>`.

Both inputs are runtime-validated. The worksheet parser accepts

only aggregate scores, case/run ids, and evidence flags; raw prompts, outputs,

reviewer identities, and unknown fields are rejected. Registered thresholds

cannot be weakened at invocation time.

The resulting version-2 receipt binds the exact scored worksheet and held-out

manifest bytes with separate SHA-256 digests plus the manifest ID/version. It

retains aggregate metrics, counts, and blockers, but omits the local input path,

raw observation rows, prompts, outputs, and reviewer identities. Operators may

use `--output=<receipt.json>` for an explicit private-safe destination.

Every receipt is explicitly `candidate-comparison`: even when its mathematical

result is promotion-eligible, it cannot change the canonical audit grade or

authorize deployment/publication until separately governed external evidence is

accepted. A synthetic receipt accidentally created during red CLI testing was

removed and a regression now requires the canonical comparisons directory to

remain empty until genuine external evidence is retained.

Prepare a fillable, privacy-safe worksheet with

`pnpm audit:ai-agent-value:prepare -- --surface=calliope-curatorial-drafting`

(substitute any model-backed surface id). Each worksheet contains both held-out

cases across five run ids, the immutable registered threshold, all score and

hard-gate fields set to `null`, and instructions. It contains no prompt, output,

or reviewer identity. After every null is replaced with an observation, the

same file is accepted by the comparison command. Canonical blank worksheets for

all six model surfaces are retained under

`artifacts/ai-agent-value/worksheets/`; blank worksheets are trial instruments,

not evidence and do not change a grade.

Subjective output review uses a separate genuinely blinded lane. Prepare a

packet-bound private worksheet with

`pnpm audit:ai-agent-value:blind-review:prepare -- --packet=<private-a-b.json> --output=<private-blank.json> --form=<private-review-form.html> --receipt=<public-readiness.json>`.

Candidate A/B text is retained privately, the model/baseline mapping stays in a

separate key, and reviewers never score cost, latency, or maintainability by

guessing. The optional offline HTML form makes no network requests, uses labeled

native controls and anchored scores, automatically measures each pair, reports

completion live, and downloads the exact response schema for private return.

Browser proof found and repaired long-digest mobile overflow; the final form has

no horizontal overflow at 360px and completed a synthetic response without

console errors. Each completed response requires a pseudonymous `REV-` code, four

independence/privacy declarations, eight 0–1 output scores (correctness,

grounding, citation quality, task completion, usefulness, calibration,

robustness, and cultural care), an A/B/tie preference, and measured seconds.

Assign exactly two private, packet-bound forms with fixed distinct pseudonymous

codes using

`pnpm audit:ai-agent-value:blind-review:assign -- --packet=<packet.json> --output-dir=<private-dir> --receipt=<public-assignment-receipt.json> --reviewer-code=REV-ONE --reviewer-code=REV-TWO`.

Codes are validated before any filename is written. The exclusive public

assignment receipt retains only packet/form digests, counts, status, and false

authority; it proves forms were prepared, not that reviews were completed.

The rendered form exposes ordinal pair numbers only; internal case/run labels

remain in the response binding but cannot prime reviewers. Desktop and 390px

browser checks confirm one main landmark, no unnamed or clipped controls, no

horizontal overflow, and no console warnings/errors.

Aggregate with

`pnpm audit:ai-agent-value:blind-review:aggregate -- --packet=<packet.json> --key=<key.json> --response=<review-1.json> --response=<review-2.json> --output=<public-aggregate.json>`.

The key must cover every row exactly once; responses must replay exact packet

text; two distinct reviewers are mandatory. The public artifact retains only

hashes, counts, mean lift, wins/ties/losses, preference agreement, reviewer

effort, and false authority. Candidate text, labels, reviewer codes, identities,

and the key never enter it.

Calliope v10 now has two fixed-code private assignment forms bound to packet

`0e899d…59f1`. Its public assignment receipt contains two distinct form digests

and no codes, paths, candidate content, or reviewer data. Unified readiness v2

therefore routes Calliope to response aggregation rather than recreating forms.

Two genuine private reviewer returns remain missing, so the AI-value grade stays

unchanged.

The cycle's handoff findings and falsifiable next hypothesis are retained in the

Calliope review-handoff recommendation(technical-recommendations/calliope-review-handoff-2026-08-25.md).

Readiness no longer treats the v10 receipt path as evidence. Strict replay binds

it to the exact representative and adversarial fixture hashes and reproduces its

dataset counts, aggregate token price, latency ordering/ceiling, five pre-provider

rights refusals, claim-evidence-v2 grounding gates, blocker, and false authority.

Any missing/extra field or drift leaves Calliope at `run-provider-trial`; only the

verified historical receipt can route the prepared packet to aggregation.

Every machine-readable surface now includes an accountable functional owner,

five explicit acceptance criteria, one bounded next experiment, and an honest

numeric scorecard. Missing evidence remains `null` rather than being inferred.

For Calliope v10 and Clio evidence-hook v6, the audit now semantically replays the

public receipt against its exact fixtures before projecting repeated observations,

mean per-system-task model cost, mean latency, safety failures, and aggregate

machine-gate state. Those rows are 60% complete because three of five scorecard

measurements are observed; comparative quality lift and reviewer-effort delta

remain `null`. This does not change either `unproven` grade. A missing or forged

receipt produces no score or fails the audit instead of receiving filename-based

credit. Each row also carries directly executable validation commands and lists

external provider/reviewer inputs separately from proof.

The audit and unified readiness CLI resolve those receipts through version 1 of a

shared six-surface evidence registry. The registry owns canonical provider,

assignment, review, retirement, source-receipt, configuration-key, verification-

kind, and trial-command metadata without importing provider implementations.

Strict trial parsers remain the semantic authority. Tests require unique surface

and provider paths and reject any copied canonical provider literal in consumers,

so a future receipt version advances through one registration change.

Verification dispatch is also exhaustive. The readiness CLI delegates all five

governed kinds to a typed verifier table; runtime coverage rejects missing,

duplicate, unknown, and orphaned kinds before projecting evidence. Each verifier

still replays its exact fixtures, manifest, price, safety, grounding, or retirement

contract. The service returns receipt paths only after successful parsing, keeping

the readiness builder independent of provider-specific evidence shapes.

The same result now carries assignment verification as a separate evidence lane.

Each public assignment receipt is parsed against its registered review-readiness

packet. Missing halves are simply not credited; malformed pairs are isolated into

a surface ID plus `ASSIGNMENT_RECEIPT_INVALID`, never raw parser text, form paths,

reviewer codes, or candidate content. Provider and retirement sets are preserved,

so handoff damage cannot mask valid machine evidence.

Readiness receipt v3 publishes those bounded outcomes through a per-surface

`blockerCodes` array. Normal rows use `[]`; invalid assignment evidence uses only

`ASSIGNMENT_RECEIPT_INVALID` and generic prose. Tests preserve provider credit and

reject packet hashes, receipt paths, parser messages, candidate content, and

reviewer codes in the public row. The explicit version bump prevents silent v2

schema drift.

First governed Calliope trial

On August 24, 2026, Calliope completed ten real Anthropic Haiku 4.5 model runs:

both registered held-out cases across five repetitions. Anthropic response usage

records show 8,395 input tokens, 2,631 output tokens, and $0.021550 total cost at

the published $1/$5 per-million-token rates. Per-run cost stayed below the

registered $0.08 ceiling. Mean latency was 4,642 ms, but the slowest run took

13,130 ms and therefore breached the registered 12,000 ms ceiling. The trial is

not promotion-eligible even before independent review. No model run formally

refused, so reviewers must examine the adversarial rights-incomplete case

carefully. The privacy-safe aggregate receipt is retained at

`artifacts/ai-agent-value/trials/calliope-curatorial-drafting-2026-08-24.json`.

Raw A/B outputs and the private randomization key remain outside Git.

The production Anthropic boundary now retains the returned model ID and exact

input, output, cache-write, and cache-read token counts. Missing or malformed

usage makes the model call fail closed to the deterministic fallback, preventing

an unmetered response from being treated as value evidence.

Calliope safety and latency repair trial

A second immutable trial reran the same two held-out cases across five slots

after repairing the observed defects. Calliope now reads explicit Linked Art

`subject_to` rights nodes, rejects a `Right` missing its required label before

provider spend, and never calls Anthropic after local policy has already refused

the task. All five adversarial slots refused before a provider call. The five

representative slots used Haiku 4.5 normally. Across ten system observations,

maximum latency fell from 13,130 ms to 3,920 ms, the latency gate changed from

fail to pass, and exact provider cost fell from $0.021550 to $0.010505. A

configurable Anthropic timeout is clamped to 100–10,000 ms, below the registered

12-second value ceiling. The aggregate receipt is

`artifacts/ai-agent-value/trials/calliope-curatorial-drafting-2026-08-24-v2.json`.

The audit grade remains unproven until independent blinded quality and reviewer-

effort scoring is accepted.

Calliope packet-capable v3 trial

The governed harness now preserves the evidence needed for genuine blind review

without putting it in Git. It compares the deterministic `object-label` output

with five paid model outputs for the representative Linked Art fixture, and

replays the malformed-rights fixture five times to prove identical pre-provider

refusal behavior. A/B position alternates within each case from a randomized

first label. The packet, separate key, and blank two-reviewer worksheet live

under the ignored `artifacts/ai-agent-value/private/` boundary; their public

readiness receipt retains only the packet SHA-256 and aggregate requirements.

On August 24, the v3 Haiku 4.5 run recorded 5,360 input and 1,043 output tokens,

$0.010575 total cost, $0.002237 maximum paid-run cost, and 3,409 ms maximum

system latency. All five adversarial observations refused before provider spend.

The provider projection now includes a capped source note, source URL, and rights

summary. Both system and user boundaries forbid inferred authentication,

attribution, ownership, provenance, or legal rights. Only `end_turn` with

non-empty content is accepted; truncation, refusal, missing stop reason, empty

content, timeout, or malformed usage becomes deterministic fallback and cannot

enter trial evidence. Public receipts are

`artifacts/ai-agent-value/trials/calliope-curatorial-drafting-2026-08-24-v3.json`

and

`artifacts/ai-agent-value/trials/calliope-curatorial-review-readiness-2026-08-24-v3.json`.

The packet digest is `b389e69833e5347a14ae15831fc85fe3ec64476e04401580db5922408d50b94c`.

No human response exists yet, so authority and grade remain false.

Calliope claim-bound v4 trial

The v4 execution contract replaces free-form model prose with strict claim-to-

source JSON. Every claim must cite one to three known source IDs. A deterministic

postflight rejects schema drift, unknown or duplicate IDs, URLs and years absent

from the specifically cited evidence, and unsupported authentication,

attribution, provenance, ownership, or rights authority. A rejected paid response

falls back locally but retains exact provider usage and estimated cost. This

prevents decorative citations and false zero-cost telemetry.

The real Haiku 4.5 run produced five accepted representative drafts containing

29 source-bound claims and five malformed-rights refusals before provider spend.

Claim-source coverage was 100%, with zero unknown IDs, unsupported URLs,

ungrounded years, or prohibited authority claims. Total cost was $0.011380;

maximum paid-run cost was $0.002500 and maximum system latency was 2,844 ms.

Public aggregate receipts are

`artifacts/ai-agent-value/trials/calliope-curatorial-drafting-2026-08-24-v4.json`

and

`artifacts/ai-agent-value/trials/calliope-curatorial-review-readiness-2026-08-24-v4.json`.

The first live attempt failed closed on model variance; a second complete run

passed all machine gates. This proves bounded operation, not benefit. Two genuine

independent blind reviews must still show at least 10% lower editing time without

quality, grounding, safety, or cultural-care regression, so the grade remains

unproven and deployment/publication authority remains false.

Calliope runtime telemetry now records an instrumented receipt for each executed

tool rather than trusting a static name list. Each receipt retains tool name,

success/refusal/failure/fallback status, bounded non-prose metrics, and a SHA-256

commitment to its result. Successful model runs therefore distinguish provider

completion from claim-source validation; paid postflight rejection records

provider success, validator failure, and local fallback; transport failure does

not claim validation; pre-model rights refusal claims neither provider nor

validator execution. The workbench exposes these statuses to reviewers. Tool

attribution is now instrumented for Mercator mapping-plan/schema validation,

Janus reconciliation/diagnostics, and Themis rights-summary/named-graph/final

review, with separate hashes and bounded metrics. Deterministic Clio now also

commits pattern analysis, conditional sparse-scope diagnosis, diagnostic

synthesis, and local drafting separately. All five runtime agents therefore

distinguish observed execution from declared capability, without treating

deterministic work as model intelligence.

The live protected route was browser-checked at 360 px with no horizontal

overflow or browser console warnings/errors. Authentication correctly limited

the anonymous run to the sign-in shell; conditional receipt rendering is covered

by component tests rather than an invented editor session.

Voyage embeddings egress readiness

The embeddings route and service now share one fail-closed provider boundary.

When an authenticated request explicitly sets `useModel: true`, Voyage must be

configured and the request must carry exact

`public-catalog`, `publicSafe: true`, and `culturalCareReviewed: true` literals.

Before AI quota evaluation or `fetch`, the service scans the bounded embedding

documents for obvious email and labeled telephone identifiers, private or non-

public note markers, and explicit culturally restricted material markers. Any

match returns `PROVIDER_EGRESS_REFUSED` without sending text or recording a

successful AI call. This scanner is defense in depth; its declaration does not

replace institutional access, community, or cultural-care governance.

Configured credentials never imply enablement: requests without explicit opt-in

stay on the deterministic baseline and consume no AI quota, while an explicit

request with no provider configuration fails closed rather than silently falling

back.

Focused route/service checks prove missing-attestation refusal,

attestation-override refusal, a single allowed and metered mock-provider call,

local-only deterministic handling, query/document intent, dated list-price

accounting for recognized `voyage-4-lite` responses, and null cost for unknown

snapshots. A strict held-out runner now compares paired rankings for one

representative and one adversarial false-friend case across five repetitions,

then emits private randomized A/B review materials and an aggregate-only receipt.

The launcher stopped before network access because no `VOYAGE_API_KEY` exists;

therefore no provider quality or value result is claimed.

Embeddings technical recommendation packet

AI quota gating, and `embedDocuments` independently enforces the same contract for

non-route callers. Persistence and retrieval semantics are unchanged.

false declarations remain possible; real provider latency, cost, relevance,

retention, and contractual processing terms are unevaluated.

care authority; AI Reliability must run the balanced worksheet only after a

Voyage key is available and retain provider usage; two independent reviewers

must score relevance against the lexical baseline before promotion.

reviewer effort over lexical hashing without any egress refusal, cost, latency,

variance, or cultural-care regression. Falsify on any hard-boundary failure or

when the registered comparison thresholds are not met.

  • Impact: the API schema owns the explicit declaration, the route refuses before
  • Highest risks: heuristic detection cannot discover every sensitive category;
  • Actions: Collection Data must classify the trial corpus and document cultural-
  • Next hypothesis: reviewed Voyage vectors improve held-out top-k relevance and

Technical recommendation packet

completion; `calliope-curatorial-trial.ts` owns repeated A/B evidence; route

authorization and publication review boundaries are unchanged.

roughly 5–8 times longer than the deterministic object label and may increase

reviewer effort; future Linked Art rights shapes may need additional structural

rules without treating a license URI alone as sufficient evidence.

privately retain identity/qualification mappings; AI Reliability must aggregate

them with the withheld key and publish only the aggregate receipt; Data/Trust

must extend CC BY, restricted, and multi-right fixtures. Acceptance commands

are both `audit:ai-agent-value:blind-review:*` commands, focused Calliope/audit

tests, and full test/lint/build gates.

by at least 10% without factuality, citation, safety, or latency regression.

Falsify it when two reviewers complete the packet-bound worksheet and the

strict aggregate shows lift below 10%, higher effort, low agreement, or any

grounding/cultural-care failure.

  • Impact: `agents.ts` owns bounded cultural-source projection and strict provider
  • Highest risks: genuine review remains absent; representative model drafts are
  • Actions: Editorial Research must obtain two independent worksheet returns and
  • Next hypothesis: Calliope's grounded model drafts reduce reviewer editing time

Reconciliation tiebreaker safety readiness

On August 24, 2026, the OpenAI tiebreaker was tightened before paid evaluation.

Evidence-equivalent ties now deterministically abstain and enter human review

without a provider call. For distinguishable ties, the model may explicitly

abstain; a selection outside the supplied candidate set, malformed usage, or a

request beyond the five-second cap fails closed. Successful responses retain

the exact returned model, input/output/total tokens, and latency. The canonical

ten-observation worksheet is prepared at

`artifacts/ai-agent-value/worksheets/openai-reconciliation-tiebreaker.json`.

Mock-provider tests prove these boundaries, not model benefit. No

`OPENAI_API_KEY` is present, so no paid trial or comparative grade is claimed.

The v2 held-out harness now matches production orchestration. Its representative

case deliberately pits a slightly higher title-similarity score against the only

candidate carrying a shared authority identifier and participant. The strongest

deterministic comparator is identifier-aware and therefore selects the

adjudicated candidate with an explicit evidence rationale; it is not a weak

highest-score straw baseline. The adversarial evidence-equivalent case abstains

five times before provider spend, so only five distinguishable-evidence model

calls are permitted. The receipt separately counts equivalence preflights,

deterministic baselines, provider calls, and candidate-confinement checks.

This design raises the bar: structured-case top-one accuracy may already be

saturated by deterministic programming. The model must instead demonstrate

materially lower reviewer effort or better blinded rationale quality without an

accuracy, abstention, latency, cost, variance, or cultural-care regression. If it

does not, the optional tiebreaker should remain disabled or be removed. Janus

agent runs now also emit instrumented reconciliation and diagnostic receipts

rather than derived tool labels. No OpenAI key or genuine review returns exist,

so v2 provider value remains unproven and no public trial receipt is claimed.

Provider rationale is no longer accepted as prose. The strict response contains

only candidate ID, abstention, and named evidence fields. Unknown, duplicate,

empty, or non-distinguishing evidence references fail postflight; the application

renders the reviewer explanation deterministically from the selected candidate's

supplied values. A paid response rejected at this boundary still persists its

returned model, internally consistent token counts, latency, dated cost estimate,

and failure status. This removes fluent invented rationale as a claimed model

benefit and makes the remaining value hypothesis narrower and falsifiable.

Technical recommendation packet

and provider-evidence retention; identity writes and reviewer authority remain

unchanged.

fluent but wrong; token counts do not establish USD cost without a versioned

pricing rule.

Reliability must run five repetitions per case with a spend-capped OpenAI key;

Evaluation must blind and import accuracy, safety, reviewer minutes, latency,

and cost. Acceptance is the focused reconciliation tests followed by `pnpm

audit:ai-agent-value:compare` and the full test/lint/build gates.

deterministic baseline is already correct, the model reduces reviewer effort

by at least 10% through a better blinded rationale without any correctness,

grounding, abstention, cost, latency, variance, or cultural-care regression.

Falsify and keep the model disabled when any gate fails or reviewer effort does

not materially improve.

  • Impact: reconciliation service ownership now includes pre-spend tie abstention
  • Highest risks: adjudicated accuracy is unknown; model rationales can still be
  • Actions: Data Curation must adjudicate both frozen candidate sets; AI
  • Next hypothesis: on distinguishable ambiguous sets where the identifier-aware

Embeddings baseline and provider readiness

The prior deterministic fallback repeated SHA-256 bytes and therefore was not a

strong semantic comparator. It now uses accent-normalized lexical unigram and

bigram feature hashing with unit-normalized vectors, making conventional token

overlap measurable and deterministic. Voyage requests declare `input_type:

document` or `query`, reject silent truncation, time out within 2.5 seconds,

require the returned model and exact total-token usage, and fail closed on

response-order, dimension, non-finite-vector, or usage drift. Recognized

`voyage-4-lite` responses use a dated conservative $0.02/M-token list-price

estimate without assuming free credits; unknown snapshots retain null cost. The

governed runner measures exact paired cost and latency, top-one relevance, and

false-friend selection over ten observations with private randomized A/B

materials. No `VOYAGE_API_KEY` is present, so the pre-network refusal and mock

evidence do not change the unproven grade.

Technical recommendation packet

model telemetry; persistence and downstream identity/publication authority are

unchanged.

governance substitute; no real-provider or independent relevance/effort

judgments exist; provider pricing and returned snapshots can drift.

Reliability must run five repetitions per case with a spend-capped key and

import only privacy-safe aggregate observations. Acceptance is the focused AI-

layer/API/trial tests, `pnpm ai:embeddings:trial`, and full gates.

quality by at least 10% over BM25 without weakening the

ambiguity false-friend boundary, exceeding 2.5 seconds, or increasing reviewer

time. The completed governed comparison falsifies any failed gate.

  • Impact: `ai-layer/embed` now owns a real deterministic retrieval baseline and
  • Highest risks: the sensitive-marker scanner is defense in depth rather than a
  • Actions: Evaluation must adjudicate the representative and false-friend rankings; AI
  • Next hypothesis: Voyage improves held-out paraphrase and multilingual ranking

Voyage retrieval trial v3

The first trial packet was not reviewable: document IDs contained words such as

`relevant` and `false`, while the query and corpus text needed to judge relevance

were absent. V2 uses opaque case/document IDs and includes the frozen query,

document text, rank order, and advisory boundary only inside the private blinded

packet. Public receipts retain no query, text, candidates, or labels.

The v2 comparator established conventional BM25 rather than hashed-vector cosine. Query

and document provider requests execute in parallel, and the latency gate measures

actual pair wall time instead of summing two request durations. V3 replaces its

inadequate two-case evaluation with three representative and three adversarial

false-friend cases, each repeated five times. Thirty rankings retain top-1

accuracy, mean reciprocal rank, false-friend frequency,

per-case ranking cardinality, and a stability gate. Query/document batches must

share returned model, cost basis, and vector dimension; vectors must be finite,

nonzero, consistently sized, and no larger than 4,096 dimensions. Valid-usage

responses rejected by postflight retain tokens, model, latency, and estimated

cost through a structured 502 response and count as consumed model calls. Dataset,

provider-call, adversarial, tool, and observation totals are derived rather than

hard-coded, and the maximum pair cost now matches the audit's $0.01 gate.

Readiness advances only after strict receipt replay reproduces the v3 manifest

digest, counts, ceilings, safety declarations, and false authority.

No `VOYAGE_API_KEY` is configured. The canonical v3 launcher therefore remains

unable to create either private or public artifacts. Therefore these are

experiment-quality and safety improvements, not evidence of Voyage benefit. The

model remains constrained pending a complete provider run and two independent

cultural-heritage relevance/effort reviews.

Visual similarity safety and evidence readiness

The v2 evidence path replaces an incapable two-case design whose metadata

baseline scored 10/10, leaving no possible 10-point model lift. It now covers

three representative and three adversarial series comparisons over thirty

observations. The strong title/creator/medium/subject baseline scores 25/30, so

only a perfect, stable model can clear the quality-lift gate. Its aggregate manifest

digest binds the image-rights basis and public evidence URL; held-out case and item identifiers

are opaque; and private A/B candidates show rankings without provider-specific

score ranges that could reveal model versus baseline. The deterministic metadata

comparator normalizes simple plural variants. Paid responses rejected during

postflight validation retain the returned model, latency, and any valid reported

cost in a structured 502 response. These are readiness improvements only:

`SIGLIP_SERVICE_URL` and independent expert observations remain absent.

The trial also replays all 24 allowlisted Art Institute of Chicago artwork

records before SigLIP spend. Each bounded request rejects redirects, oversized or

invalid JSON, and any record no longer marked public domain. Failure occurs before

the first model call; public evidence retains only the count and aggregate digest.

The 24/24 live preflight is retained in

`artifacts/ai-agent-value/visual-rights-preflight-2026-08-25.json` and grants no

deployment, publication, attribution, or provider-spend authority.

The implementation is a configured SigLIP ranking service with an IIIF URL-

topology fallback—not Voyage and not a metadata-ranking model. The fallback can

recognize alternate derivatives of the same IIIF image; it does not measure

visual relatedness and is labeled accordingly. Before SigLIP invocation, every

reference and candidate must be a public HTTPS URL and fan-out is limited to 100.

The exact request must also declare `public-images`, `rightsReviewed: true`,

`providerFetchPermitted: true`, and `culturalCareReviewed: true`. Missing or

partial review returns `VISUAL_PROVIDER_EGRESS_REFUSED` before AI quota evaluation

or `fetch`; `rankVisualSimilarity` repeats the check for direct callers.

The response must identify its model and provide a complete, unique, candidate-

confined set of finite scores in the `[-1, 1]` range. Requests time out within

3.5 seconds. Results carry model, latency, candidate-count, and validated

provider-reported per-ranking cost telemetry; missing economics remain null and

cannot pass the cost gate. Results also carry an

explicit advisory contract: resemblance supports discovery ranking only and

does not establish identity, attribution, influence, or provenance. The

strict held-out runner compares normalized metadata overlap with five repeated

model rankings for each of six rights-documented cases, measures stability/cost/latency, and emits private randomized image/

source review materials plus an aggregate-only receipt. `SIGLIP_SERVICE_URL` is

not configured, so the launcher created no evidence artifacts and no value grade

changes. Focused route/service/trial checks prove the contracts; mocks are safety

and readiness evidence, not model-value evidence.

Technical recommendation packet

validation, strict model-output confinement, timeout, telemetry, and

interpretation boundaries; the route remains read-only and cannot mutate

collection records.

the normalized-metadata comparator is an evaluation baseline rather than the

runtime fallback; provider cost is unknown unless the service reports it; source

terms can differ from generic public reachability.

candidate pool and rights basis; Platform must version SigLIP deployment and

report per-ranking infrastructure cost; Evaluation must run five blinded

repetitions per case and retain aggregate judgments only. Acceptance is the

focused visual tests, `pnpm ai:visual-similarity:trial`, and full gates.

independently rated usefulness by at least 10% over the 25/30 normalized

metadata baseline without a

prohibited heritage claim, rights/egress violation, latency breach, or added

reviewer burden. The governed comparison falsifies any failed gate.

  • Impact: `ai-layer/visual` now owns rights/cultural-care approval, URL-egress
  • Highest risks: declarations can be false and require institutional governance;
  • Actions: Data/Curatorial must verify the frozen related and false-friend
  • Next hypothesis: SigLIP achieves 30/30 stable expected rankings and raises

Clio social-editor governed trial

Clio now runs an explicit `facebook-house-style-v1` baseline before model spend.

It detects hype, unsupported certainty, institutional endorsement language, and

excessive punctuation while preserving negated heritage boundaries such as

“does not prove.” Provider egress requires an explicit public-safe attestation;

evidence is count/length bounded and email addresses refuse locally. Anthropic

responses must end normally, identify the returned model, include exact token

usage, and finish within ten seconds. A `retain` verdict must preserve the draft

exactly; `revise` must materially differ. New URLs, unsafe output, hidden model

self-scores, refusal, truncation, and missing usage all fail closed.

The standalone operator path additionally requires

`pnpm facebook:review -- --use-model ...`; a loaded Anthropic credential is

availability only and cannot trigger Clio through an accidental review command.

The governed trial remains explicitly model-backed by definition, and both paths

retain false publication authority.

On August 24, 2026, the governed harness ran both registered held-out cases five

times with Haiku 4.5. All ten paid calls completed. Exact usage was 4,275 input

and 2,357 output tokens, costing $0.016060 at the retained $1/$5 per-million-

token rates; maximum per-run cost was $0.001871. Latency ranged from 2,167 to

3,899 ms (mean 2,906 ms), below the 12-second gate. The representative case was

retained five times and the adversarial hype/endorsement case revised five

times, with zero provider, output-safety, or authority failures. The public

receipt is `artifacts/ai-agent-value/trials/clio-social-editor-2026-08-24.json`.

Raw A/B copy and its randomization key remain in the Git-ignored private packet.

The real packet is now bound to a ten-observation blank worksheet and the public

readiness receipt

`artifacts/ai-agent-value/trials/clio-social-editor-review-readiness-2026-08-24.json`.

The audit grade remains unproven until two genuine independent blinded reviewers

supply complete preference, quality, grounding, cultural-care, and reviewer-time

observations.

That first packet is superseded for human review: its adversarial comparator was

a refusal sentence rather than usable social copy, its case IDs disclosed roles,

it omitted evidence needed to judge grounding, and placement could be imbalanced.

The August 25 v3 rerun uses a deterministic safe revision, opaque case IDs,

identical original/source/evidence context on both sides, and exactly five

model-A and five model-B placements. Ten real Haiku calls passed local gates:

4,275 input and 2,338 output tokens, $0.015965 total, maximum per-run cost

$0.001921, and 2,268–4,603 ms latency (mean 3,363 ms). Current receipts are

`artifacts/ai-agent-value/trials/clio-social-editor-2026-08-25-v3.json` and

`artifacts/ai-agent-value/trials/clio-social-editor-review-readiness-2026-08-25-v3.json`.

Paid postflight rejections retain model, usage, and latency through

`ClioReviewValidationError`. No human result or value lift is claimed.

The current free-form model behavior is now retired. v4/v5 added a deterministic

novel-token evidence gate and retained privacy-safe paid failures after the

model replaced false certainty with unsupported speculation such as ungrounded

research or connection language. The runtime now sends unsafe drafts directly

to the deterministic safe edit before provider spend and reserves Haiku for

already-safe experimental drafts. Homogeneous v6/v7 sessions contain 10 paid

reviews and 10 deterministic preflight edits, cost $0.007365–$0.007675 per

session, and peak at 2,885 ms. Their private packets contain 20/20 byte-identical

model/baseline candidate pairs. The public equivalence receipt retains only

packet hashes and counts and sets `retire-current-model-behavior`; human review

of identical candidates is neither useful nor required. A materially different

capability and registered held-out hypothesis must precede further Clio spend.

The retained private packet can now be converted to a responsive offline review

form with `--form=<private-review-form.html>`. This removes manual JSON editing

and automatically records directly observed pair time while preserving the same

packet digest, candidate text, declarations, response parser, withheld key, and

aggregate-only public boundary. It is interface readiness only: no reviewer

response or Clio value lift has been inferred.

Blind-review aggregation no longer hides protected regressions inside one mean.

The public receipt retains privacy-safe model/baseline means, lift, and a

non-regression result for each of correctness, grounding, citation quality, task

completion, usefulness, calibration, robustness, and cultural care. It also

lists regressed dimensions and separately gates grounding, citation quality,

calibration, robustness, and cultural care. A positive overall mean therefore

cannot mask a heritage-grounding or cultural-care loss. These fields still grant

no audit-grade, deployment, or publication authority.

Technical recommendation packet

attestation, provider telemetry, and semantic retain/revise validation; the

publisher's exact-digest human approval boundary is unchanged.

and editing-time judgments are absent; the retained token-price calculation

must be revised if model pricing changes.

model invocation experimental/retired, and require a new capability contract

that produces non-identical evidence-grounded candidates before any reviewer

assignment or provider budget. Trust should still expand public-safe evidence

tests for phone, donor, and audience identifiers.

non-identical candidate delta over the safe-edit baseline. Only then may two

independent reviewers evaluate the registered 10% quality margin, protected

grounding/cultural-care gates, and reviewer-time cost.

  • Impact: `facebook-social-editor` now owns deterministic preflight, egress
  • Highest risks: deterministic phrase rules can miss subtler hype; human quality
  • Actions: retain the deterministic safe edit in production, keep current Clio
  • Next hypothesis: a future Clio capability must first produce a repeatable,

Clio evidence-cited reader-hook trial

The materially different `clio-evidence-hook` hypothesis removes safety editing

from the proposed model benefit. Its strongest practical baseline concatenates

the same first two bounded evidence statements. Haiku may produce one or two

shorter reader-facing sentences, but each sentence cites explicit evidence IDs

and deterministic postflight rejects unknown citations or factual tokens absent

from those cited statements. Public-safe attestation, bounded evidence, strict

usage/model telemetry, an eight-second cap, and false publication/attribution/

provenance authority remain mandatory.

V1 and v2 retained privacy-safe paid failures for JSON-envelope handling and a

non-factual discourse connective. V3 then exposed a more important semantic

defect: all five adversarial model hooks omitted the required negated boundary,

so its packet is superseded. Contract v2 types fact versus boundary evidence,

requires every boundary ID, and preserves explicit negation. V4 confirmed the

repair; v5 added replayable grounding counts but its private rows contained extra

context fields rejected by the shared exact-schema blind parser. V6 emits the

canonical four-field rows and completed ten balanced calls with 10 required

boundary citations, 10 observed citations, 10 preserved negations, ten

non-identical pairs, 2,215 input tokens, 719 output tokens, $0.00581 total cost,

$0.000617 maximum per run, and 1,198 ms worst latency.

The private arm-aware materiality analyzer binds the v6 packet, withheld key,

and trial receipt and emits aggregate-only evidence. It finds 10/10 normalized

non-cosmetic pairs, one stable model output per case, 0.942 mean token-set

Jaccard, and a 5.1% model length increase (143.5 versus 136.5 characters).

Therefore the packet is reviewable but has no automatic quality or effort win.

Two distinct offline forms now bind the exact v6 packet. Their public receipt

retains only packet/form hashes, counts, status, and false authority; readiness

semantically replays it against v6 before routing to aggregation. No form return,

reviewer independence, response, or review completion is claimed.

The receipt is accepted by readiness only after semantic replay of both frozen

fixture hashes, dataset counts, pricing, latency, safety, blockers, and authority.

These are provider and candidate-delta facts, not a quality grade. Two genuine

independent blinded reviewers must still establish at least 10% lift with no

protected-dimension or reviewer-effort regression.

Technical recommendation packet

retired free-form editor and from publication workflows.

independent review is absent; provider formatting/vocabulary drift can consume

small paid calls before postflight.

must aggregate exact responses; Trust must add negation-scope and word-order

adversaries. Acceptance commands are the blind-review assignment/aggregation

commands, focused evidence-hook/readiness tests, and full gates.

grounding, citation, calibration, robustness, cultural-care, latency, cost,

variance, or effort regression. Any failed gate retires the behavior.

  • Impact: the evidence-hook service and trial harness are isolated from the
  • Highest risks: lexical token confinement is weaker than semantic entailment;
  • Actions: Editorial Research must distribute two packet-bound forms; Evaluation
  • Next hypothesis: independent reviewers prefer v6 by at least 10% without a

Independent-review receipt verification

Readiness no longer treats a public review-receipt filename as human evidence.

The shared verifier parses the aggregate-only schema, recomputes lift and gate

consistency, validates hash/reviewer/observation and pairwise counts, bounds

agreement and effort, enforces false authority, and binds the registered surface.

A malformed receipt contributes only `REVIEW_RECEIPT_INVALID`; private response

paths, reviewer metadata, candidate content, hashes, and parser output are

excluded. No canonical review receipt exists, so review completion and proven

advantage remain zero.

Technical recommendation packet

prove reviewer identity, independence, or source-response authenticity.

Reliability aggregates privately; Trust accepts evidence separately.

and effort without leaking private review material.

  • Impact: arbitrary JSON at a registered review path cannot advance readiness.
  • Highest risks: genuine reviewers remain absent; aggregate consistency cannot
  • Actions: Evaluation collects two real returns per prepared packet; AI
  • Next hypothesis: a valid Calliope or Clio aggregate exposes measurable quality

The verifier additionally requires that aggregate to match the registered,

semantically verified assignment/readiness chain. Packet hash, surface, reviewer

count, and observations per reviewer must agree exactly. A valid-looking receipt

for another packet receives no credit and cannot erase upstream evidence.

The comparison command now invokes the same chain verifier. Structurally valid

same-surface receipts with an unregistered packet hash fail before promotion

evaluation, so a readiness rejection cannot be bypassed downstream. The public

comparison receipt remains content-addressed and non-authoritative.

Reviewer workflow v2 retains all 160 dimension choices across a 10-pair packet but

adds per-pair status and next-incomplete keyboard navigation. No score defaults,

bulk scoring, local storage, or network access were introduced. Fresh private

Calliope and Clio forms are content-bound by v2 assignment receipts, so the

improvement reaches the prepared review handoffs without invalidating packets.