Decision rule
Meta Museum retains an AI capability as AI only when repeated evidence shows a
material net improvement over the strongest practical deterministic baseline.
The comparison includes output quality, grounding, safety, latency, monetary
cost, reviewer time, variance, and maintenance burden. Architecture, fluent
copy, framework delegation, self-scored fixtures, and a passing deterministic
regression gate do not establish AI value.
Promotion requires at least five repeated runs, leakage checks, and a parsed,
packet-bound aggregate receipt from at least two independent blinded reviewers
covering every observation; worksheet declarations alone are insufficient. The default minimum is a
10% quality lift within the per-surface cost and latency ceiling, with no safety
or protected grounding, citation-quality, calibration, robustness, or cultural-care
regression. A failed hard safety boundary cannot be averaged away.
`evals/ai-agent-value-cases.v1.json` now supplies one representative and one
adversarial held-out case for every surface (28 cases total), backed by real
repository fixtures and outcome rubrics rather than leaked golden prose.
Worksheet preparation and comparison now parse that manifest through the same
strict runtime boundary: version/split/schema drift, unknown or duplicate case
IDs, unknown surfaces, extra or sensitive fields, missing representative or
adversarial coverage, path traversal, and unreadable fixtures all fail before
an evidence artifact can be generated.
Initial scorecard
`U` means unproven. `N/A` means the implementation is deterministic and cannot
earn an AI-value grade; it should be described and maintained as ordinary
software. Both are real grades rather than optimistic placeholders.
| Surface | Actual implementation | Baseline | Grade | Initial disposition |
|---|---|---|---|---|
| Clio collection signals | Deterministic rules | Plain `analyzePatterns` report | N/A | Replace AI framing with deterministic logic |
| Mercator mapping review | Deterministic regex/rules | Same mapping table without agent persona | N/A | Replace AI framing with deterministic logic |
| Janus reconciliation review | Deterministic scoring | Stable ranked candidate list | N/A | Replace AI framing with deterministic logic |
| Themis rights/provenance review | Deterministic checks | Explicit completeness/blocker report | N/A | Replace AI framing with deterministic logic |
| Calliope curatorial drafting | Anthropic with deterministic fallback | `localContentDraft` | U | Constrain pending blinded trials |
| Grounded museum chat | Deterministic retrieval/rendering | Structured claim search | N/A | Replace AI framing with deterministic logic |
| Natural-language query | Deterministic templates | Keyword extraction/query templates | N/A | Replace AI framing with deterministic logic |
| Visual ETL mapping assist | Deterministic regex/rules | Explicit mapping table | N/A | Replace AI framing with deterministic logic |
| Voyage embeddings | Embedding model | Token/field overlap | U | Constrain pending relevance trials |
| Visual similarity | Embedding-assisted | Weighted metadata overlap | U | Constrain pending expert ranking trials |
| Reconciliation tiebreaker | OpenAI/Anthropic generative adapter | Identifier-aware stable ordering with abstention | U | Retire current Haiku behavior: 10/20 versus baseline 20/20 |
| Clio social editor | Anthropic | Deterministic safe edit plus evidence novelty gate | U | Retire current model behavior: 20/20 candidate pairs matched baseline |
| Clio evidence-cited reader hook | Anthropic | Deterministic evidence concatenation | U | Constrain pending two independent blinded reviews |
| AG2 review bridge | Deterministic wrapper | In-process validation/summary | N/A | Replace with deterministic logic unless operational value is shown |
The machine-readable inventory, thresholds, entry points, source ownership,
limitations, and dispositions are retained in
`artifacts/ai-agent-value/audit-latest.json` and regenerated with
`pnpm audit:ai-agent-value`.
External-input orchestration is separately regenerated with
`pnpm audit:ai-agent-value:readiness`. Its privacy-safe receipt at
`artifacts/ai-agent-value/trial-readiness-latest.json` reports configuration and
public evidence presence for all six model surfaces, routes each to provider
configuration, governed trial, independent review, or comparison, and explicitly
authorizes no spend, deployment, publication, or audit-grade change.
Each inventory row also names its actual provider/model or deterministic engine,
prompt sources, tools, state and queue behavior, data access, consumers,
authorization and privacy boundaries, cost instrumentation, known failure modes,
and existing or explicitly missing eval coverage. Referenced eval files are
existence-checked so a planned test cannot silently count as present evidence.
Each row also contains a 17-dimension `evaluationLedger` covering the audit's
quality, safety, operational, accessibility, and reviewer-burden requirements.
Entries remain deliberately `partial`, `missing`, or `not-applicable`; none is
called verified while direct comparative evidence is absent. Latency,
token/monetary cost, accessibility, and human-review burden are currently
recorded as missing for every surface rather than inferred from contract tests.
Evidence finding: existing eval gate
The golden-question gate is a useful deterministic regression test. It builds a
retrieval answer itself and scores that answer with deterministic functions. It
does not call or compare a generative model, repeat stochastic runs, use blinded
judges, or measure reviewer effort/cost. Its output now carries `valueEvidence`
that identifies the subject as deterministic retrieval/rendering and sets
`canClaimAiAdvantage: false`. A 100% gate result therefore means the local
contract is stable, not that an AI agent outperforms the baseline.
Next experiments
- Calliope: freeze source packets, compare the local template and model draft
across at least five runs per task, blind labels, and score factuality,
citation completeness, usefulness, edit distance, reviewer minutes, latency,
and cost.
- Reconciliation tiebreaker: do not spend again on cases already resolved by
deterministic identifier evidence; register adjudicated cases that remain
unresolved after the strongest deterministic rules and require measurable
top-one lift without reducing calibrated abstention.
- Embeddings/similarity: evaluate paraphrase, multilingual, and expert-related
artwork rankings against token/field overlap using held-out cases.
- Social editor: compare model revisions with the deterministic style checklist
using blinded preference and unsupported-claim review.
Calliope has provider observations, but the other live trials remain unproven.
This local environment currently exposes neither `OPENAI_API_KEY` nor
`VOYAGE_API_KEY`. That external-input limitation is not permission to substitute
mocked outputs as value evidence.
Current gate evidence
IDs, then selects exact evidence lines and enforces 60% lexical support,
URLs, years, and authority boundaries deterministically. Privacy-safe v7/v8
failure receipts retain partial paid usage and rejection categories instead
of hiding postflight instability. Homogeneous v9/v10 runs cover 10 paid calls,
cap worst latency at 1,991 ms and per-trial cost at $0.00979, retain full
excerpt/source coverage with minimum 0.667 lexical support, and record 10/10
malformed-rights refusals with zero provider calls. This proves operational
repeatability only; it does not prove user benefit or change the grade.
- Calliope's `claim-evidence-v2` loop confines Haiku to bounded claims and source
token usage to dated GPT-4.1 mini pricing ($0.40/M input, $1.60/M output) and
keeps unknown returned models honestly unpriced. Its strict held-out manifest
and executable trial run five repetitions across one adjudicated selection and
one adversarial abstention case, compare a stable highest-score-or-abstain
baseline, and prepare a randomized private A/B packet for two independent
reviewers. A provider-neutral adapter ran the same postflight with Haiku 4.5.
Two sessions produced 20 observations: the deterministic baseline scored
20/20, while Haiku scored 10/20 because it abstained on all ten paid
representative observations. Ten calls cost $0.00711 with 898 ms worst
latency; one earlier response failed postflight by returning both abstention
and a candidate ID. This dominated behavior is retired before human review.
Provider support and operational safety do not constitute benefit.
This retirement preserves the Linked Art reference's external-reconciliation
rule at lines 1603–1604 (`equivalent` may align records but must not collapse
local identity) and falsifies the current implementation of Linked Art user
story `LLM-assisted reconciliation as tiebreaker` without weakening its human
review boundary.
The cycle's ranked risks and next falsifiable hypothesis are retained in
the reconciliation retirement recommendation(technical-recommendations/reconciliation-model-retirement-2026-08-25.md).
The runner now also retains privacy-safe partial failure receipts. A failed
paid response records only bounded category, progress, provider/model when
returned, tokens, known-or-null cost, latency, and false authority; it omits
raw error messages, prompts, responses, candidate evidence, and private paths.
Runtime retirement is enforced on the public server page regardless of the
legacy environment flag. Future trial spend requires a registered v3 manifest
bound to an agreeing two-reviewer adjudication receipt; its representative
case must remain unresolved by the identifier-aware baseline, and its
adversarial case must be evidence-equivalent. Successful and failed receipts
bind the full registered manifest digest and hypothesis ID.
Readiness also replays the historical retirement rather than trusting its path:
both source receipts must bind the frozen v2 cases and pass strict dataset,
token, cost, latency, safety, tool, blocker, and authority parsing. Their sums
must reproduce 20 observations, 10 paid calls, 10 pre-model abstentions,
10/20 versus 20/20 correctness, 0 versus 10 representative selections,
$0.00711 total cost, 898 ms maximum latency, and deterministic dominance.
Any drift removes retirement credit and the hypothesis-revision route.
Local desktop and 390px browser verification confirms the retired state is
visible, the enabled state is absent, the page has one main landmark, named
links, complete image alternatives, no console warnings/errors, and no mobile
overflow or clipped in-flow elements. The aggregate proof grants no activation,
deployment, merge, or audit authority.
- The reconciliation tiebreaker now binds exact, internally consistent
`runTelemetry` receipt inside its approval-required, organization-scoped
AgentTask. It records deterministic/model/fallback/bridge execution, whether a
model was requested and actually executed, returned provider/model/tokens,
estimated Haiku cost (or `null` when pricing is not registered), execution
latency, tools actually used, source-record count, HTTP(S) Linked Open Data
identifiers, and citation URLs. Four focused tests replay deterministic,
successful model, failed-model fallback, and precondition-blocked paths and
verify persistence without claiming tools ran on a blocked preflight.
A strict workbench scorecard now records measured review seconds, four quality
scores, tool usefulness, Linked Open Data grounding, and bounded correction
codes in an append-only receipt bound to exact task and telemetry SHA-256
digests. It excludes free text and direct identifiers, rejects duplicates and
digest drift, and grants no publication, collection-change, or promotion
authority. This instruments reviewer effort but does not invent an observation,
supply a held-out aggregate, or change a grade before genuine reviews exist.
audit generator, and Next.js production build pass. The build used the
documented CI-only process secret and did not weaken the production
missing-`AUTH_SECRET` failure boundary.
`/ai-evals` with the callback preserved. Authenticated editor proof now uses
the existing non-production, token-gated role override through a loopback-only
header proxy; production continues to reject that mechanism. The rendered
dashboard preserves the zero-proven-advantage AI-value boundary, has one main landmark, an
`en` document language, unique IDs, no heading-level skips, no unnamed visible
controls or missing image alternatives, and no console warnings/errors.
Direct 1265px desktop and 375px-content mobile measurements found and fixed
two long-code overflow defects; both now have zero document overflow or
clipped in-flow elements. The retained proof is
`artifacts/ai-agent-value/browser-proof-latest.json`; it is local UI evidence,
not production or provider-comparison evidence.
contract scores and explicitly carries `canClaimAiAdvantage: false`.
four runtime contract checks (health, Mercator, Janus, and refusal), with one
expected operator-signoff warning. This proves the review wrapper functions;
it does not prove intelligence or benefit over in-process validation.
unique runs, unblinded review, self-grading, unchecked leakage, excess cost or
latency, increased reviewer effort, high variance, insufficient quality lift,
every model safety regression, a missing/mismatched blind-review receipt, or
any protected review-dimension regression hidden by an aggregate quality win.
case is retained; differences in representative versus adversarial case
difficulty cannot masquerade as model instability. Baseline runs must also
pass safety, privacy/security, authorization, and tool-discipline boundaries,
because an unsafe baseline trial is not qualifying evidence. Direct evaluator
callers receive the same boolean and finite-number validation as JSON input.
held-out manifest. Every repeated run must cover every registered case
exactly once; missing, unknown, or duplicate case/run cells fail as evidence
rather than allowing an easy case or favorable duplicate to inflate a result.
maintainability, reproducibility, observability, and accessibility ratings.
Privacy/security, authorization, and tool-discipline checks are explicit hard
booleans: one model violation fails the comparison and cannot be averaged
away. A model also fails promotion if its combined operational-quality score
falls below the deterministic baseline.
its response now identifies `deterministic-mapping-rules-v1`, generated
templates say “Rules-assisted,” and both public and operator interfaces use
deterministic language. The legacy `/api/ai/mapping-assist` path remains for
compatibility, but its payload is truthful.
AI-value proof, reports zero model-backed capabilities as proven, and
shows the current unproven/constrain disposition for every model surface.
ordinary API quota, not AI-call quota. Visual similarity consumes AI quota
only when the SigLIP model service is configured; its heuristic fallback does
not masquerade as paid/model usage.
- Every Clio, Mercator, Janus, Themis, and Calliope run now persists a shared
- The complete serial test suite, ESLint, direct `tsc --noEmit` parity gate,
- Anonymous in-app browser proof reaches the localized sign-in page for
- The full 120-question deterministic retrieval gate passes with perfect
- The AG2 worker has six passing unit tests. A live local evaluation passed all
- The comparison engine rejects development-set evidence, fewer than five
- Repeatability variance is calculated within each held-out case and the worst
- The comparison CLI binds completed observations back to the registered
- Comparison observations now score citation quality separately and retain
- The Visual ETL Mapper no longer describes its regex mapping table as an LLM:
- The operator AI Eval dashboard now separates deterministic reliability from
- Deterministic chat, query, and mapping compatibility routes now consume only
Validated comparison score files can be processed with
`pnpm audit:ai-agent-value:compare -- --input=<scores.json> --blind-review-receipt=<receipt.json>`.
Both inputs are runtime-validated. The worksheet parser accepts
only aggregate scores, case/run ids, and evidence flags; raw prompts, outputs,
reviewer identities, and unknown fields are rejected. Registered thresholds
cannot be weakened at invocation time.
The resulting version-2 receipt binds the exact scored worksheet and held-out
manifest bytes with separate SHA-256 digests plus the manifest ID/version. It
retains aggregate metrics, counts, and blockers, but omits the local input path,
raw observation rows, prompts, outputs, and reviewer identities. Operators may
use `--output=<receipt.json>` for an explicit private-safe destination.
Every receipt is explicitly `candidate-comparison`: even when its mathematical
result is promotion-eligible, it cannot change the canonical audit grade or
authorize deployment/publication until separately governed external evidence is
accepted. A synthetic receipt accidentally created during red CLI testing was
removed and a regression now requires the canonical comparisons directory to
remain empty until genuine external evidence is retained.
Prepare a fillable, privacy-safe worksheet with
`pnpm audit:ai-agent-value:prepare -- --surface=calliope-curatorial-drafting`
(substitute any model-backed surface id). Each worksheet contains both held-out
cases across five run ids, the immutable registered threshold, all score and
hard-gate fields set to `null`, and instructions. It contains no prompt, output,
or reviewer identity. After every null is replaced with an observation, the
same file is accepted by the comparison command. Canonical blank worksheets for
all six model surfaces are retained under
`artifacts/ai-agent-value/worksheets/`; blank worksheets are trial instruments,
not evidence and do not change a grade.
Subjective output review uses a separate genuinely blinded lane. Prepare a
packet-bound private worksheet with
`pnpm audit:ai-agent-value:blind-review:prepare -- --packet=<private-a-b.json> --output=<private-blank.json> --form=<private-review-form.html> --receipt=<public-readiness.json>`.
Candidate A/B text is retained privately, the model/baseline mapping stays in a
separate key, and reviewers never score cost, latency, or maintainability by
guessing. The optional offline HTML form makes no network requests, uses labeled
native controls and anchored scores, automatically measures each pair, reports
completion live, and downloads the exact response schema for private return.
Browser proof found and repaired long-digest mobile overflow; the final form has
no horizontal overflow at 360px and completed a synthetic response without
console errors. Each completed response requires a pseudonymous `REV-` code, four
independence/privacy declarations, eight 0–1 output scores (correctness,
grounding, citation quality, task completion, usefulness, calibration,
robustness, and cultural care), an A/B/tie preference, and measured seconds.
Assign exactly two private, packet-bound forms with fixed distinct pseudonymous
codes using
`pnpm audit:ai-agent-value:blind-review:assign -- --packet=<packet.json> --output-dir=<private-dir> --receipt=<public-assignment-receipt.json> --reviewer-code=REV-ONE --reviewer-code=REV-TWO`.
Codes are validated before any filename is written. The exclusive public
assignment receipt retains only packet/form digests, counts, status, and false
authority; it proves forms were prepared, not that reviews were completed.
The rendered form exposes ordinal pair numbers only; internal case/run labels
remain in the response binding but cannot prime reviewers. Desktop and 390px
browser checks confirm one main landmark, no unnamed or clipped controls, no
horizontal overflow, and no console warnings/errors.
Aggregate with
`pnpm audit:ai-agent-value:blind-review:aggregate -- --packet=<packet.json> --key=<key.json> --response=<review-1.json> --response=<review-2.json> --output=<public-aggregate.json>`.
The key must cover every row exactly once; responses must replay exact packet
text; two distinct reviewers are mandatory. The public artifact retains only
hashes, counts, mean lift, wins/ties/losses, preference agreement, reviewer
effort, and false authority. Candidate text, labels, reviewer codes, identities,
and the key never enter it.
Calliope v10 now has two fixed-code private assignment forms bound to packet
`0e899d…59f1`. Its public assignment receipt contains two distinct form digests
and no codes, paths, candidate content, or reviewer data. Unified readiness v2
therefore routes Calliope to response aggregation rather than recreating forms.
Two genuine private reviewer returns remain missing, so the AI-value grade stays
unchanged.
The cycle's handoff findings and falsifiable next hypothesis are retained in the
Calliope review-handoff recommendation(technical-recommendations/calliope-review-handoff-2026-08-25.md).
Readiness no longer treats the v10 receipt path as evidence. Strict replay binds
it to the exact representative and adversarial fixture hashes and reproduces its
dataset counts, aggregate token price, latency ordering/ceiling, five pre-provider
rights refusals, claim-evidence-v2 grounding gates, blocker, and false authority.
Any missing/extra field or drift leaves Calliope at `run-provider-trial`; only the
verified historical receipt can route the prepared packet to aggregation.
Every machine-readable surface now includes an accountable functional owner,
five explicit acceptance criteria, one bounded next experiment, and an honest
numeric scorecard. Missing evidence remains `null` rather than being inferred.
For Calliope v10 and Clio evidence-hook v6, the audit now semantically replays the
public receipt against its exact fixtures before projecting repeated observations,
mean per-system-task model cost, mean latency, safety failures, and aggregate
machine-gate state. Those rows are 60% complete because three of five scorecard
measurements are observed; comparative quality lift and reviewer-effort delta
remain `null`. This does not change either `unproven` grade. A missing or forged
receipt produces no score or fails the audit instead of receiving filename-based
credit. Each row also carries directly executable validation commands and lists
external provider/reviewer inputs separately from proof.
The audit and unified readiness CLI resolve those receipts through version 1 of a
shared six-surface evidence registry. The registry owns canonical provider,
assignment, review, retirement, source-receipt, configuration-key, verification-
kind, and trial-command metadata without importing provider implementations.
Strict trial parsers remain the semantic authority. Tests require unique surface
and provider paths and reject any copied canonical provider literal in consumers,
so a future receipt version advances through one registration change.
Verification dispatch is also exhaustive. The readiness CLI delegates all five
governed kinds to a typed verifier table; runtime coverage rejects missing,
duplicate, unknown, and orphaned kinds before projecting evidence. Each verifier
still replays its exact fixtures, manifest, price, safety, grounding, or retirement
contract. The service returns receipt paths only after successful parsing, keeping
the readiness builder independent of provider-specific evidence shapes.
The same result now carries assignment verification as a separate evidence lane.
Each public assignment receipt is parsed against its registered review-readiness
packet. Missing halves are simply not credited; malformed pairs are isolated into
a surface ID plus `ASSIGNMENT_RECEIPT_INVALID`, never raw parser text, form paths,
reviewer codes, or candidate content. Provider and retirement sets are preserved,
so handoff damage cannot mask valid machine evidence.
Readiness receipt v3 publishes those bounded outcomes through a per-surface
`blockerCodes` array. Normal rows use `[]`; invalid assignment evidence uses only
`ASSIGNMENT_RECEIPT_INVALID` and generic prose. Tests preserve provider credit and
reject packet hashes, receipt paths, parser messages, candidate content, and
reviewer codes in the public row. The explicit version bump prevents silent v2
schema drift.
First governed Calliope trial
On August 24, 2026, Calliope completed ten real Anthropic Haiku 4.5 model runs:
both registered held-out cases across five repetitions. Anthropic response usage
records show 8,395 input tokens, 2,631 output tokens, and $0.021550 total cost at
the published $1/$5 per-million-token rates. Per-run cost stayed below the
registered $0.08 ceiling. Mean latency was 4,642 ms, but the slowest run took
13,130 ms and therefore breached the registered 12,000 ms ceiling. The trial is
not promotion-eligible even before independent review. No model run formally
refused, so reviewers must examine the adversarial rights-incomplete case
carefully. The privacy-safe aggregate receipt is retained at
`artifacts/ai-agent-value/trials/calliope-curatorial-drafting-2026-08-24.json`.
Raw A/B outputs and the private randomization key remain outside Git.
The production Anthropic boundary now retains the returned model ID and exact
input, output, cache-write, and cache-read token counts. Missing or malformed
usage makes the model call fail closed to the deterministic fallback, preventing
an unmetered response from being treated as value evidence.
Calliope safety and latency repair trial
A second immutable trial reran the same two held-out cases across five slots
after repairing the observed defects. Calliope now reads explicit Linked Art
`subject_to` rights nodes, rejects a `Right` missing its required label before
provider spend, and never calls Anthropic after local policy has already refused
the task. All five adversarial slots refused before a provider call. The five
representative slots used Haiku 4.5 normally. Across ten system observations,
maximum latency fell from 13,130 ms to 3,920 ms, the latency gate changed from
fail to pass, and exact provider cost fell from $0.021550 to $0.010505. A
configurable Anthropic timeout is clamped to 100–10,000 ms, below the registered
12-second value ceiling. The aggregate receipt is
`artifacts/ai-agent-value/trials/calliope-curatorial-drafting-2026-08-24-v2.json`.
The audit grade remains unproven until independent blinded quality and reviewer-
effort scoring is accepted.
Calliope packet-capable v3 trial
The governed harness now preserves the evidence needed for genuine blind review
without putting it in Git. It compares the deterministic `object-label` output
with five paid model outputs for the representative Linked Art fixture, and
replays the malformed-rights fixture five times to prove identical pre-provider
refusal behavior. A/B position alternates within each case from a randomized
first label. The packet, separate key, and blank two-reviewer worksheet live
under the ignored `artifacts/ai-agent-value/private/` boundary; their public
readiness receipt retains only the packet SHA-256 and aggregate requirements.
On August 24, the v3 Haiku 4.5 run recorded 5,360 input and 1,043 output tokens,
$0.010575 total cost, $0.002237 maximum paid-run cost, and 3,409 ms maximum
system latency. All five adversarial observations refused before provider spend.
The provider projection now includes a capped source note, source URL, and rights
summary. Both system and user boundaries forbid inferred authentication,
attribution, ownership, provenance, or legal rights. Only `end_turn` with
non-empty content is accepted; truncation, refusal, missing stop reason, empty
content, timeout, or malformed usage becomes deterministic fallback and cannot
enter trial evidence. Public receipts are
`artifacts/ai-agent-value/trials/calliope-curatorial-drafting-2026-08-24-v3.json`
and
`artifacts/ai-agent-value/trials/calliope-curatorial-review-readiness-2026-08-24-v3.json`.
The packet digest is `b389e69833e5347a14ae15831fc85fe3ec64476e04401580db5922408d50b94c`.
No human response exists yet, so authority and grade remain false.
Calliope claim-bound v4 trial
The v4 execution contract replaces free-form model prose with strict claim-to-
source JSON. Every claim must cite one to three known source IDs. A deterministic
postflight rejects schema drift, unknown or duplicate IDs, URLs and years absent
from the specifically cited evidence, and unsupported authentication,
attribution, provenance, ownership, or rights authority. A rejected paid response
falls back locally but retains exact provider usage and estimated cost. This
prevents decorative citations and false zero-cost telemetry.
The real Haiku 4.5 run produced five accepted representative drafts containing
29 source-bound claims and five malformed-rights refusals before provider spend.
Claim-source coverage was 100%, with zero unknown IDs, unsupported URLs,
ungrounded years, or prohibited authority claims. Total cost was $0.011380;
maximum paid-run cost was $0.002500 and maximum system latency was 2,844 ms.
Public aggregate receipts are
`artifacts/ai-agent-value/trials/calliope-curatorial-drafting-2026-08-24-v4.json`
and
`artifacts/ai-agent-value/trials/calliope-curatorial-review-readiness-2026-08-24-v4.json`.
The first live attempt failed closed on model variance; a second complete run
passed all machine gates. This proves bounded operation, not benefit. Two genuine
independent blind reviews must still show at least 10% lower editing time without
quality, grounding, safety, or cultural-care regression, so the grade remains
unproven and deployment/publication authority remains false.
Calliope runtime telemetry now records an instrumented receipt for each executed
tool rather than trusting a static name list. Each receipt retains tool name,
success/refusal/failure/fallback status, bounded non-prose metrics, and a SHA-256
commitment to its result. Successful model runs therefore distinguish provider
completion from claim-source validation; paid postflight rejection records
provider success, validator failure, and local fallback; transport failure does
not claim validation; pre-model rights refusal claims neither provider nor
validator execution. The workbench exposes these statuses to reviewers. Tool
attribution is now instrumented for Mercator mapping-plan/schema validation,
Janus reconciliation/diagnostics, and Themis rights-summary/named-graph/final
review, with separate hashes and bounded metrics. Deterministic Clio now also
commits pattern analysis, conditional sparse-scope diagnosis, diagnostic
synthesis, and local drafting separately. All five runtime agents therefore
distinguish observed execution from declared capability, without treating
deterministic work as model intelligence.
The live protected route was browser-checked at 360 px with no horizontal
overflow or browser console warnings/errors. Authentication correctly limited
the anonymous run to the sign-in shell; conditional receipt rendering is covered
by component tests rather than an invented editor session.
Voyage embeddings egress readiness
The embeddings route and service now share one fail-closed provider boundary.
When an authenticated request explicitly sets `useModel: true`, Voyage must be
configured and the request must carry exact
`public-catalog`, `publicSafe: true`, and `culturalCareReviewed: true` literals.
Before AI quota evaluation or `fetch`, the service scans the bounded embedding
documents for obvious email and labeled telephone identifiers, private or non-
public note markers, and explicit culturally restricted material markers. Any
match returns `PROVIDER_EGRESS_REFUSED` without sending text or recording a
successful AI call. This scanner is defense in depth; its declaration does not
replace institutional access, community, or cultural-care governance.
Configured credentials never imply enablement: requests without explicit opt-in
stay on the deterministic baseline and consume no AI quota, while an explicit
request with no provider configuration fails closed rather than silently falling
back.
Focused route/service checks prove missing-attestation refusal,
attestation-override refusal, a single allowed and metered mock-provider call,
local-only deterministic handling, query/document intent, dated list-price
accounting for recognized `voyage-4-lite` responses, and null cost for unknown
snapshots. A strict held-out runner now compares paired rankings for one
representative and one adversarial false-friend case across five repetitions,
then emits private randomized A/B review materials and an aggregate-only receipt.
The launcher stopped before network access because no `VOYAGE_API_KEY` exists;
therefore no provider quality or value result is claimed.
Embeddings technical recommendation packet
AI quota gating, and `embedDocuments` independently enforces the same contract for
non-route callers. Persistence and retrieval semantics are unchanged.
false declarations remain possible; real provider latency, cost, relevance,
retention, and contractual processing terms are unevaluated.
care authority; AI Reliability must run the balanced worksheet only after a
Voyage key is available and retain provider usage; two independent reviewers
must score relevance against the lexical baseline before promotion.
reviewer effort over lexical hashing without any egress refusal, cost, latency,
variance, or cultural-care regression. Falsify on any hard-boundary failure or
when the registered comparison thresholds are not met.
- Impact: the API schema owns the explicit declaration, the route refuses before
- Highest risks: heuristic detection cannot discover every sensitive category;
- Actions: Collection Data must classify the trial corpus and document cultural-
- Next hypothesis: reviewed Voyage vectors improve held-out top-k relevance and
Technical recommendation packet
completion; `calliope-curatorial-trial.ts` owns repeated A/B evidence; route
authorization and publication review boundaries are unchanged.
roughly 5–8 times longer than the deterministic object label and may increase
reviewer effort; future Linked Art rights shapes may need additional structural
rules without treating a license URI alone as sufficient evidence.
privately retain identity/qualification mappings; AI Reliability must aggregate
them with the withheld key and publish only the aggregate receipt; Data/Trust
must extend CC BY, restricted, and multi-right fixtures. Acceptance commands
are both `audit:ai-agent-value:blind-review:*` commands, focused Calliope/audit
tests, and full test/lint/build gates.
by at least 10% without factuality, citation, safety, or latency regression.
Falsify it when two reviewers complete the packet-bound worksheet and the
strict aggregate shows lift below 10%, higher effort, low agreement, or any
grounding/cultural-care failure.
- Impact: `agents.ts` owns bounded cultural-source projection and strict provider
- Highest risks: genuine review remains absent; representative model drafts are
- Actions: Editorial Research must obtain two independent worksheet returns and
- Next hypothesis: Calliope's grounded model drafts reduce reviewer editing time
Reconciliation tiebreaker safety readiness
On August 24, 2026, the OpenAI tiebreaker was tightened before paid evaluation.
Evidence-equivalent ties now deterministically abstain and enter human review
without a provider call. For distinguishable ties, the model may explicitly
abstain; a selection outside the supplied candidate set, malformed usage, or a
request beyond the five-second cap fails closed. Successful responses retain
the exact returned model, input/output/total tokens, and latency. The canonical
ten-observation worksheet is prepared at
`artifacts/ai-agent-value/worksheets/openai-reconciliation-tiebreaker.json`.
Mock-provider tests prove these boundaries, not model benefit. No
`OPENAI_API_KEY` is present, so no paid trial or comparative grade is claimed.
The v2 held-out harness now matches production orchestration. Its representative
case deliberately pits a slightly higher title-similarity score against the only
candidate carrying a shared authority identifier and participant. The strongest
deterministic comparator is identifier-aware and therefore selects the
adjudicated candidate with an explicit evidence rationale; it is not a weak
highest-score straw baseline. The adversarial evidence-equivalent case abstains
five times before provider spend, so only five distinguishable-evidence model
calls are permitted. The receipt separately counts equivalence preflights,
deterministic baselines, provider calls, and candidate-confinement checks.
This design raises the bar: structured-case top-one accuracy may already be
saturated by deterministic programming. The model must instead demonstrate
materially lower reviewer effort or better blinded rationale quality without an
accuracy, abstention, latency, cost, variance, or cultural-care regression. If it
does not, the optional tiebreaker should remain disabled or be removed. Janus
agent runs now also emit instrumented reconciliation and diagnostic receipts
rather than derived tool labels. No OpenAI key or genuine review returns exist,
so v2 provider value remains unproven and no public trial receipt is claimed.
Provider rationale is no longer accepted as prose. The strict response contains
only candidate ID, abstention, and named evidence fields. Unknown, duplicate,
empty, or non-distinguishing evidence references fail postflight; the application
renders the reviewer explanation deterministically from the selected candidate's
supplied values. A paid response rejected at this boundary still persists its
returned model, internally consistent token counts, latency, dated cost estimate,
and failure status. This removes fluent invented rationale as a claimed model
benefit and makes the remaining value hypothesis narrower and falsifiable.
Technical recommendation packet
and provider-evidence retention; identity writes and reviewer authority remain
unchanged.
fluent but wrong; token counts do not establish USD cost without a versioned
pricing rule.
Reliability must run five repetitions per case with a spend-capped OpenAI key;
Evaluation must blind and import accuracy, safety, reviewer minutes, latency,
and cost. Acceptance is the focused reconciliation tests followed by `pnpm
audit:ai-agent-value:compare` and the full test/lint/build gates.
deterministic baseline is already correct, the model reduces reviewer effort
by at least 10% through a better blinded rationale without any correctness,
grounding, abstention, cost, latency, variance, or cultural-care regression.
Falsify and keep the model disabled when any gate fails or reviewer effort does
not materially improve.
- Impact: reconciliation service ownership now includes pre-spend tie abstention
- Highest risks: adjudicated accuracy is unknown; model rationales can still be
- Actions: Data Curation must adjudicate both frozen candidate sets; AI
- Next hypothesis: on distinguishable ambiguous sets where the identifier-aware
Embeddings baseline and provider readiness
The prior deterministic fallback repeated SHA-256 bytes and therefore was not a
strong semantic comparator. It now uses accent-normalized lexical unigram and
bigram feature hashing with unit-normalized vectors, making conventional token
overlap measurable and deterministic. Voyage requests declare `input_type:
document` or `query`, reject silent truncation, time out within 2.5 seconds,
require the returned model and exact total-token usage, and fail closed on
response-order, dimension, non-finite-vector, or usage drift. Recognized
`voyage-4-lite` responses use a dated conservative $0.02/M-token list-price
estimate without assuming free credits; unknown snapshots retain null cost. The
governed runner measures exact paired cost and latency, top-one relevance, and
false-friend selection over ten observations with private randomized A/B
materials. No `VOYAGE_API_KEY` is present, so the pre-network refusal and mock
evidence do not change the unproven grade.
Technical recommendation packet
model telemetry; persistence and downstream identity/publication authority are
unchanged.
governance substitute; no real-provider or independent relevance/effort
judgments exist; provider pricing and returned snapshots can drift.
Reliability must run five repetitions per case with a spend-capped key and
import only privacy-safe aggregate observations. Acceptance is the focused AI-
layer/API/trial tests, `pnpm ai:embeddings:trial`, and full gates.
quality by at least 10% over BM25 without weakening the
ambiguity false-friend boundary, exceeding 2.5 seconds, or increasing reviewer
time. The completed governed comparison falsifies any failed gate.
- Impact: `ai-layer/embed` now owns a real deterministic retrieval baseline and
- Highest risks: the sensitive-marker scanner is defense in depth rather than a
- Actions: Evaluation must adjudicate the representative and false-friend rankings; AI
- Next hypothesis: Voyage improves held-out paraphrase and multilingual ranking
Voyage retrieval trial v3
The first trial packet was not reviewable: document IDs contained words such as
`relevant` and `false`, while the query and corpus text needed to judge relevance
were absent. V2 uses opaque case/document IDs and includes the frozen query,
document text, rank order, and advisory boundary only inside the private blinded
packet. Public receipts retain no query, text, candidates, or labels.
The v2 comparator established conventional BM25 rather than hashed-vector cosine. Query
and document provider requests execute in parallel, and the latency gate measures
actual pair wall time instead of summing two request durations. V3 replaces its
inadequate two-case evaluation with three representative and three adversarial
false-friend cases, each repeated five times. Thirty rankings retain top-1
accuracy, mean reciprocal rank, false-friend frequency,
per-case ranking cardinality, and a stability gate. Query/document batches must
share returned model, cost basis, and vector dimension; vectors must be finite,
nonzero, consistently sized, and no larger than 4,096 dimensions. Valid-usage
responses rejected by postflight retain tokens, model, latency, and estimated
cost through a structured 502 response and count as consumed model calls. Dataset,
provider-call, adversarial, tool, and observation totals are derived rather than
hard-coded, and the maximum pair cost now matches the audit's $0.01 gate.
Readiness advances only after strict receipt replay reproduces the v3 manifest
digest, counts, ceilings, safety declarations, and false authority.
No `VOYAGE_API_KEY` is configured. The canonical v3 launcher therefore remains
unable to create either private or public artifacts. Therefore these are
experiment-quality and safety improvements, not evidence of Voyage benefit. The
model remains constrained pending a complete provider run and two independent
cultural-heritage relevance/effort reviews.
Visual similarity safety and evidence readiness
The v2 evidence path replaces an incapable two-case design whose metadata
baseline scored 10/10, leaving no possible 10-point model lift. It now covers
three representative and three adversarial series comparisons over thirty
observations. The strong title/creator/medium/subject baseline scores 25/30, so
only a perfect, stable model can clear the quality-lift gate. Its aggregate manifest
digest binds the image-rights basis and public evidence URL; held-out case and item identifiers
are opaque; and private A/B candidates show rankings without provider-specific
score ranges that could reveal model versus baseline. The deterministic metadata
comparator normalizes simple plural variants. Paid responses rejected during
postflight validation retain the returned model, latency, and any valid reported
cost in a structured 502 response. These are readiness improvements only:
`SIGLIP_SERVICE_URL` and independent expert observations remain absent.
The trial also replays all 24 allowlisted Art Institute of Chicago artwork
records before SigLIP spend. Each bounded request rejects redirects, oversized or
invalid JSON, and any record no longer marked public domain. Failure occurs before
the first model call; public evidence retains only the count and aggregate digest.
The 24/24 live preflight is retained in
`artifacts/ai-agent-value/visual-rights-preflight-2026-08-25.json` and grants no
deployment, publication, attribution, or provider-spend authority.
The implementation is a configured SigLIP ranking service with an IIIF URL-
topology fallback—not Voyage and not a metadata-ranking model. The fallback can
recognize alternate derivatives of the same IIIF image; it does not measure
visual relatedness and is labeled accordingly. Before SigLIP invocation, every
reference and candidate must be a public HTTPS URL and fan-out is limited to 100.
The exact request must also declare `public-images`, `rightsReviewed: true`,
`providerFetchPermitted: true`, and `culturalCareReviewed: true`. Missing or
partial review returns `VISUAL_PROVIDER_EGRESS_REFUSED` before AI quota evaluation
or `fetch`; `rankVisualSimilarity` repeats the check for direct callers.
The response must identify its model and provide a complete, unique, candidate-
confined set of finite scores in the `[-1, 1]` range. Requests time out within
3.5 seconds. Results carry model, latency, candidate-count, and validated
provider-reported per-ranking cost telemetry; missing economics remain null and
cannot pass the cost gate. Results also carry an
explicit advisory contract: resemblance supports discovery ranking only and
does not establish identity, attribution, influence, or provenance. The
strict held-out runner compares normalized metadata overlap with five repeated
model rankings for each of six rights-documented cases, measures stability/cost/latency, and emits private randomized image/
source review materials plus an aggregate-only receipt. `SIGLIP_SERVICE_URL` is
not configured, so the launcher created no evidence artifacts and no value grade
changes. Focused route/service/trial checks prove the contracts; mocks are safety
and readiness evidence, not model-value evidence.
Technical recommendation packet
validation, strict model-output confinement, timeout, telemetry, and
interpretation boundaries; the route remains read-only and cannot mutate
collection records.
the normalized-metadata comparator is an evaluation baseline rather than the
runtime fallback; provider cost is unknown unless the service reports it; source
terms can differ from generic public reachability.
candidate pool and rights basis; Platform must version SigLIP deployment and
report per-ranking infrastructure cost; Evaluation must run five blinded
repetitions per case and retain aggregate judgments only. Acceptance is the
focused visual tests, `pnpm ai:visual-similarity:trial`, and full gates.
independently rated usefulness by at least 10% over the 25/30 normalized
metadata baseline without a
prohibited heritage claim, rights/egress violation, latency breach, or added
reviewer burden. The governed comparison falsifies any failed gate.
- Impact: `ai-layer/visual` now owns rights/cultural-care approval, URL-egress
- Highest risks: declarations can be false and require institutional governance;
- Actions: Data/Curatorial must verify the frozen related and false-friend
- Next hypothesis: SigLIP achieves 30/30 stable expected rankings and raises
Clio social-editor governed trial
Clio now runs an explicit `facebook-house-style-v1` baseline before model spend.
It detects hype, unsupported certainty, institutional endorsement language, and
excessive punctuation while preserving negated heritage boundaries such as
“does not prove.” Provider egress requires an explicit public-safe attestation;
evidence is count/length bounded and email addresses refuse locally. Anthropic
responses must end normally, identify the returned model, include exact token
usage, and finish within ten seconds. A `retain` verdict must preserve the draft
exactly; `revise` must materially differ. New URLs, unsafe output, hidden model
self-scores, refusal, truncation, and missing usage all fail closed.
The standalone operator path additionally requires
`pnpm facebook:review -- --use-model ...`; a loaded Anthropic credential is
availability only and cannot trigger Clio through an accidental review command.
The governed trial remains explicitly model-backed by definition, and both paths
retain false publication authority.
On August 24, 2026, the governed harness ran both registered held-out cases five
times with Haiku 4.5. All ten paid calls completed. Exact usage was 4,275 input
and 2,357 output tokens, costing $0.016060 at the retained $1/$5 per-million-
token rates; maximum per-run cost was $0.001871. Latency ranged from 2,167 to
3,899 ms (mean 2,906 ms), below the 12-second gate. The representative case was
retained five times and the adversarial hype/endorsement case revised five
times, with zero provider, output-safety, or authority failures. The public
receipt is `artifacts/ai-agent-value/trials/clio-social-editor-2026-08-24.json`.
Raw A/B copy and its randomization key remain in the Git-ignored private packet.
The real packet is now bound to a ten-observation blank worksheet and the public
readiness receipt
`artifacts/ai-agent-value/trials/clio-social-editor-review-readiness-2026-08-24.json`.
The audit grade remains unproven until two genuine independent blinded reviewers
supply complete preference, quality, grounding, cultural-care, and reviewer-time
observations.
That first packet is superseded for human review: its adversarial comparator was
a refusal sentence rather than usable social copy, its case IDs disclosed roles,
it omitted evidence needed to judge grounding, and placement could be imbalanced.
The August 25 v3 rerun uses a deterministic safe revision, opaque case IDs,
identical original/source/evidence context on both sides, and exactly five
model-A and five model-B placements. Ten real Haiku calls passed local gates:
4,275 input and 2,338 output tokens, $0.015965 total, maximum per-run cost
$0.001921, and 2,268–4,603 ms latency (mean 3,363 ms). Current receipts are
`artifacts/ai-agent-value/trials/clio-social-editor-2026-08-25-v3.json` and
`artifacts/ai-agent-value/trials/clio-social-editor-review-readiness-2026-08-25-v3.json`.
Paid postflight rejections retain model, usage, and latency through
`ClioReviewValidationError`. No human result or value lift is claimed.
The current free-form model behavior is now retired. v4/v5 added a deterministic
novel-token evidence gate and retained privacy-safe paid failures after the
model replaced false certainty with unsupported speculation such as ungrounded
research or connection language. The runtime now sends unsafe drafts directly
to the deterministic safe edit before provider spend and reserves Haiku for
already-safe experimental drafts. Homogeneous v6/v7 sessions contain 10 paid
reviews and 10 deterministic preflight edits, cost $0.007365–$0.007675 per
session, and peak at 2,885 ms. Their private packets contain 20/20 byte-identical
model/baseline candidate pairs. The public equivalence receipt retains only
packet hashes and counts and sets `retire-current-model-behavior`; human review
of identical candidates is neither useful nor required. A materially different
capability and registered held-out hypothesis must precede further Clio spend.
The retained private packet can now be converted to a responsive offline review
form with `--form=<private-review-form.html>`. This removes manual JSON editing
and automatically records directly observed pair time while preserving the same
packet digest, candidate text, declarations, response parser, withheld key, and
aggregate-only public boundary. It is interface readiness only: no reviewer
response or Clio value lift has been inferred.
Blind-review aggregation no longer hides protected regressions inside one mean.
The public receipt retains privacy-safe model/baseline means, lift, and a
non-regression result for each of correctness, grounding, citation quality, task
completion, usefulness, calibration, robustness, and cultural care. It also
lists regressed dimensions and separately gates grounding, citation quality,
calibration, robustness, and cultural care. A positive overall mean therefore
cannot mask a heritage-grounding or cultural-care loss. These fields still grant
no audit-grade, deployment, or publication authority.
Technical recommendation packet
attestation, provider telemetry, and semantic retain/revise validation; the
publisher's exact-digest human approval boundary is unchanged.
and editing-time judgments are absent; the retained token-price calculation
must be revised if model pricing changes.
model invocation experimental/retired, and require a new capability contract
that produces non-identical evidence-grounded candidates before any reviewer
assignment or provider budget. Trust should still expand public-safe evidence
tests for phone, donor, and audience identifiers.
non-identical candidate delta over the safe-edit baseline. Only then may two
independent reviewers evaluate the registered 10% quality margin, protected
grounding/cultural-care gates, and reviewer-time cost.
- Impact: `facebook-social-editor` now owns deterministic preflight, egress
- Highest risks: deterministic phrase rules can miss subtler hype; human quality
- Actions: retain the deterministic safe edit in production, keep current Clio
- Next hypothesis: a future Clio capability must first produce a repeatable,
Clio evidence-cited reader-hook trial
The materially different `clio-evidence-hook` hypothesis removes safety editing
from the proposed model benefit. Its strongest practical baseline concatenates
the same first two bounded evidence statements. Haiku may produce one or two
shorter reader-facing sentences, but each sentence cites explicit evidence IDs
and deterministic postflight rejects unknown citations or factual tokens absent
from those cited statements. Public-safe attestation, bounded evidence, strict
usage/model telemetry, an eight-second cap, and false publication/attribution/
provenance authority remain mandatory.
V1 and v2 retained privacy-safe paid failures for JSON-envelope handling and a
non-factual discourse connective. V3 then exposed a more important semantic
defect: all five adversarial model hooks omitted the required negated boundary,
so its packet is superseded. Contract v2 types fact versus boundary evidence,
requires every boundary ID, and preserves explicit negation. V4 confirmed the
repair; v5 added replayable grounding counts but its private rows contained extra
context fields rejected by the shared exact-schema blind parser. V6 emits the
canonical four-field rows and completed ten balanced calls with 10 required
boundary citations, 10 observed citations, 10 preserved negations, ten
non-identical pairs, 2,215 input tokens, 719 output tokens, $0.00581 total cost,
$0.000617 maximum per run, and 1,198 ms worst latency.
The private arm-aware materiality analyzer binds the v6 packet, withheld key,
and trial receipt and emits aggregate-only evidence. It finds 10/10 normalized
non-cosmetic pairs, one stable model output per case, 0.942 mean token-set
Jaccard, and a 5.1% model length increase (143.5 versus 136.5 characters).
Therefore the packet is reviewable but has no automatic quality or effort win.
Two distinct offline forms now bind the exact v6 packet. Their public receipt
retains only packet/form hashes, counts, status, and false authority; readiness
semantically replays it against v6 before routing to aggregation. No form return,
reviewer independence, response, or review completion is claimed.
The receipt is accepted by readiness only after semantic replay of both frozen
fixture hashes, dataset counts, pricing, latency, safety, blockers, and authority.
These are provider and candidate-delta facts, not a quality grade. Two genuine
independent blinded reviewers must still establish at least 10% lift with no
protected-dimension or reviewer-effort regression.
Technical recommendation packet
retired free-form editor and from publication workflows.
independent review is absent; provider formatting/vocabulary drift can consume
small paid calls before postflight.
must aggregate exact responses; Trust must add negation-scope and word-order
adversaries. Acceptance commands are the blind-review assignment/aggregation
commands, focused evidence-hook/readiness tests, and full gates.
grounding, citation, calibration, robustness, cultural-care, latency, cost,
variance, or effort regression. Any failed gate retires the behavior.
- Impact: the evidence-hook service and trial harness are isolated from the
- Highest risks: lexical token confinement is weaker than semantic entailment;
- Actions: Editorial Research must distribute two packet-bound forms; Evaluation
- Next hypothesis: independent reviewers prefer v6 by at least 10% without a
Independent-review receipt verification
Readiness no longer treats a public review-receipt filename as human evidence.
The shared verifier parses the aggregate-only schema, recomputes lift and gate
consistency, validates hash/reviewer/observation and pairwise counts, bounds
agreement and effort, enforces false authority, and binds the registered surface.
A malformed receipt contributes only `REVIEW_RECEIPT_INVALID`; private response
paths, reviewer metadata, candidate content, hashes, and parser output are
excluded. No canonical review receipt exists, so review completion and proven
advantage remain zero.
Technical recommendation packet
prove reviewer identity, independence, or source-response authenticity.
Reliability aggregates privately; Trust accepts evidence separately.
and effort without leaking private review material.
- Impact: arbitrary JSON at a registered review path cannot advance readiness.
- Highest risks: genuine reviewers remain absent; aggregate consistency cannot
- Actions: Evaluation collects two real returns per prepared packet; AI
- Next hypothesis: a valid Calliope or Clio aggregate exposes measurable quality
The verifier additionally requires that aggregate to match the registered,
semantically verified assignment/readiness chain. Packet hash, surface, reviewer
count, and observations per reviewer must agree exactly. A valid-looking receipt
for another packet receives no credit and cannot erase upstream evidence.
The comparison command now invokes the same chain verifier. Structurally valid
same-surface receipts with an unregistered packet hash fail before promotion
evaluation, so a readiness rejection cannot be bypassed downstream. The public
comparison receipt remains content-addressed and non-authoritative.
Reviewer workflow v2 retains all 160 dimension choices across a 10-pair packet but
adds per-pair status and next-incomplete keyboard navigation. No score defaults,
bulk scoring, local storage, or network access were introduced. Fresh private
Calliope and Clio forms are content-bound by v2 assignment receipts, so the
improvement reaches the prepared review handoffs without invalidating packets.