← Documentation home

Canonical Markdown source · Oct 20, 2018

Meta Museum Roadmap

roadmap.md · 2,948 lines · SHA-256 cb6a996a755f

  • [ ] Prove or retire AI-agent value: inventory all 14 current agent/model-assisted surfaces, stop crediting eight deterministic heuristics/wrappers as AI advantage, and run repeated blinded comparisons for the six model-backed workflows against explicit deterministic baselines. The 28-case held-out representative/adversarial manifest and fail-closed comparison engine now cover repetition, complete registered case/run balance, duplicate rejection, leakage, quality lift, safety, cost, latency, within-case repeatability variance, and reviewer effort. Promotion additionally requires a parsed packet-bound receipt from at least two independent blinded reviewers covering every observation; worksheet booleans alone cannot qualify evidence, and grounding, citation quality, calibration, robustness, or cultural-care regression blocks promotion even when aggregate quality wins. The initial score remains 0 proven AI advantages; genuine reviewer evidence remains pending.
  • [x] Make agent-value evidence operationally auditable: attach prompt sources, actual provider/model or deterministic engine, tools, state/queues, data access, consumers, authorization/privacy boundaries, cost instrumentation, failure modes, and real-or-missing eval coverage to all 14 inventory rows; add a privacy-minimized comparison CLI with immutable registered thresholds; and replace false LLM language in the deterministic Visual ETL Mapper while preserving its compatibility route.
  • [x] Publish a machine-readable 17-dimension evidence ledger per surface and expose the deterministic-reliability versus AI-value boundary in the operator eval dashboard.
  • [x] Correct usage accounting so deterministic chat, query, mapping, and visual-similarity fallback requests do not consume AI-call quota; retain AI accounting only for actual model-service invocation.
  • [x] Expand comparative scoring to cover citation quality and operational quality explicitly, with non-compensating privacy/security, authorization, tool-discipline, and safety gates.
  • [x] Generate directly runnable, privacy-safe comparison worksheets for all six model-backed surfaces: two held-out cases, five repeated runs, immutable thresholds, null-until-observed scores, and no raw prompts, outputs, or reviewer identities. Preparation and comparison runtime-validate the entire manifest for exact coverage, unique IDs, strict fields, safe relative fixture paths, and readable fixtures before writing evidence.
  • [x] Make comparison receipts replayable without widening retained data: version 2 binds the exact scored worksheet and held-out manifest with SHA-256 plus manifest identity, retains only aggregate result/count/blocker fields, omits local paths and raw observation rows, and supports an explicit private-safe output destination.
  • [x] Measure Calliope operational repeatability across real provider sessions: after v7/v8 privacy-safe failure receipts exposed brittle model-authored excerpts, `claim-evidence-v2` assigns exact evidence-line selection to deterministic tooling and confines Haiku to bounded claims/source IDs. The homogeneous v9/v10 stability receipt spans 10 paid calls, 10/10 pre-provider rights refusals, full excerpt/source coverage, minimum 0.667 lexical support, $0.00979 maximum trial cost, and 1,991 ms worst latency. It explicitly grants no user-benefit, deployment, publication, or audit-grade authority while blind human review remains incomplete.
  • [x] Remove reviewer priming and assignment errors from the offline blind-review lane: reviewer-visible legends now expose only opaque pair ordinals rather than internal representative/adversarial case and run labels; a fail-closed command creates exactly two packet-bound forms with distinct fixed pseudonymous codes. Real desktop and 390px browser checks confirm one main landmark, labeled/unclipped controls, zero overflow, and zero console warnings/errors.
  • [x] Make Calliope's independent review handoff executable and auditable: validate both reviewer codes before filenames are constructed, generate two exclusive offline forms for the exact v10 packet, retain a public assignment receipt containing only packet/form digests and counts with false authority, and route unified readiness v2 to private-response aggregation instead of recreating assignments. Genuine reviewer responses remain external and unclaimed.
  • [x] Remove filename authority from Calliope provider evidence: readiness strictly replays v10 against both frozen fixture hashes, recomputed Anthropic token pricing, latency ordering, refusal safety, claim-evidence-v2 grounding, blockers, and false authority before exposing the reviewer aggregation route. Extra-key, count, fixture, cost, or authority drift fails closed.
  • [x] Prevent configured credentials from silently enabling unproven models: Voyage embeddings and SigLIP visual similarity now remain deterministic unless an authenticated request explicitly sets `useModel: true`; explicit model requests fail closed with `MODEL_PROVIDER_UNAVAILABLE` when configuration is absent, and AI quota is charged only for genuine opt-in model execution. Public health reports model availability rather than claiming model enablement.
  • [x] Align Clio with the same experimental authorization rule: the standalone Facebook review command now fails before file or provider access unless the operator supplies `--use-model`; the governed trial stays explicitly model-backed, while review output retains false publication authority and still requires exact-digest human approval.
  • [x] Retire Clio's current model behavior when evidence shows no benefit: v4/v5 retain paid speculation-laundering failures; deterministic preflight now handles unsafe drafts without provider spend; v6/v7 bind 10 paid safe-draft reviews and 10 zero-cost adversarial edits; and a privacy-safe equivalence receipt proves all 20 blinded candidate pairs exactly match the deterministic baseline. Readiness now requires a new model hypothesis rather than pointless reviewer assignment.
  • [x] Retire the current reconciliation model behavior when it underperforms: a provider-neutral adapter placed Haiku behind the same candidate-confinement and equivalence preflight as GPT-4.1 mini; two governed sessions retained 10 paid calls plus 10 no-call abstentions, but Haiku scored 10/20 against the identifier-aware deterministic baseline's 20/20 and had one earlier schema-inconsistent response. The aggregate retirement receipt blocks human review and further spend until a future hypothesis targets genuinely unresolved adjudicated cases.
  • [x] Make reconciliation retirement reproducible rather than filename-driven: readiness strictly parses both historical source receipts against the frozen v2 cases, then recomputes every retirement aggregate, deterministic-dominance predicate, disposition, next hypothesis, and false authority. Provider or retirement path presence alone no longer changes routing.
  • [x] Preserve reconciliation trial failures rather than losing paid evidence: trial-level aggregation now catches postflight, missing usage/cost, and provider failures; writes an exclusive aggregate-only receipt with partial progress, tokens, known-or-null cost, latency, and bounded category; excludes error text/prompts/responses/candidate evidence; and grants no retry, merge, deployment, publication, or audit authority.
  • [x] Enforce reconciliation retirement at runtime and before new spend: the public server page always uses deterministic reconciliation and ignores the legacy model flag; the trial CLI no longer embeds the dominated v2 corpus and constructs no provider adapter until a strict v3 manifest passes a matching digest-bound, agreeing two-reviewer adjudication receipt, a deterministic-baseline-abstention eligibility check, and an evidence-equivalent adversarial check. Historical v2 parsing remains solely for receipt replay/tests.
  • [x] Unify external-input orchestration across the six model surfaces: readiness v2 accepts alternative provider keys per surface, distinguishes assignment-ready from review-complete without retaining values or private paths, routes both Calliope and the evidence hook from semantically verified prepared forms to response aggregation, and routes the dominated Clio and reconciliation behaviors to hypothesis revision with false spend/deployment/publication/audit authority.
  • [x] Prevent synthetic comparison output from masquerading as canonical proof: remove the test-generated promotion receipt, label every receipt `candidate-comparison`, retain false audit-grade/deployment/publication authority plus required external-evidence acceptance, and keep the canonical comparisons directory empty until genuine governed evidence lands.
  • [x] Complete local authenticated browser and accessibility proof for the AI-value dashboard: anonymous users reach `/en/signin?callbackUrl=%2Fai-evals`; the existing non-production token-gated editor override renders `/en/ai-evals` through a loopback-only proxy while production remains fail-closed. Desktop and 375px-content checks preserve the zero-proven-value boundary, one main landmark, language/title, heading structure, unique IDs, named controls, image alternatives, zero console warnings/errors, and zero overflow/clipped elements after fixing long source and artifact path wrapping. Retain `artifacts/ai-agent-value/browser-proof-latest.json` as local UI evidence, not provider or deployment proof.
  • [x] Replace the public documentation dump with a curated, searchable hub: six maintained starting points, complete title/path discovery, readable canonical documents, source dates rather than deployment freshness, stable section navigation, checksums, machine endpoints, and responsive no-overflow layouts.
  • [x] Retain canonical production evidence for the homepage provider cycle: after pausing autoplay, fourteen manual advances produced fourteen distinct providers and fourteen distinct images from a 52-record rights-qualified pool with no duplicate URLs; preserve the point-in-time and no-partnership claim boundaries.
  • [x] Establish a site-wide audit ledger and authoritative App Router inventory: classify all 77 pages and 190 route handlers (351 HTTP methods), eliminate every unassigned page, expand browser checks to all noindex/success/workflow surfaces, centralize all 15 anonymous protected-route expectations, and align private research approval/preflight pages with the proxy role policy.
  • [x] Replace the homepage's generic gradient-and-card treatment with a museum-editorial visual system: warm paper, artwork-led composition, object-label metadata, serif display hierarchy, restrained rules, unboxed pathways, source visibility, responsive layouts without decorative glass effects, and WCAG-safe caption contrast.
  • [x] Stop the homepage hero from restarting with the same artwork: generate a fresh deterministic order per browser session, preserve provider round-robin balance, retain every unique image, and disclose measured image-ready coverage against the fourteen governed production connectors without equating connector availability with reusable visual coverage.
  • [x] Raise homepage visual coverage to all fourteen governed production collections with source-backed, high-resolution artwork candidates, explicit record/image/rights attribution, fail-closed rights qualification, provider-balanced ordering, deduplication, broken-image fallback, reduced-motion behavior, and measured live coverage.
  • [x] Rebuild the Connections index and first public journey as artwork-led editorial experiences: frame the index as an episodic chain reaction, lead the journey with a visual mystery and two clickable verified public-domain works, follow a six-turn from-to chain with an explicit bridge in every chapter, use concrete looking prompts and a visual hinge to vary the rhythm, label museum facts versus editorial inferences and confidence row by row, end at a bounded reveal, keep provenance and rights visible, and preserve the machine-readable research layer without presenting a generic dashboard.
  • [x] Refine the first Connection for reader choice and narrative momentum: add direct look/follow/audit entry paths, carry the clue forward between all six turns, subordinate repetitive claim rows inside labeled keyboard-accessible evidence disclosures, retain source links and fact/inference boundaries, and verify zero overflow or broken images at 1440px and 390px.
  • [x] Stabilize the hosted accessibility gate by auditing the optimized production build rather than recompiling all 150 route/viewport cases through the development server; explicitly trust the CI localhost host and enforce both requirements with a workflow regression.
  • [x] Bind the accessibility/overflow matrix to every canonical sitemap entry in addition to the explicit operational-route inventory, distinguish successful typed JSON/JSON-LD/schema resources from HTML documents, and pass 204/204 local route-viewport checks across 68 routes at desktop, 200%-equivalent, and mobile.
  • [x] Rebuild Artwork of the Day as a source-backed daily editorial surface with a high-resolution image, compact fact ledger, explicit collection/image/rights verification links, attribution, stable dated navigation, and a provider-first UTC sequence proven to cover all 14 verified providers without image repetition across a 14-day window.
  • [x] Browser-exercise the canonical homepage's paused provider-first round: 14/14 sources, 14 unique images, zero broken media, and visible rights context. Replace the 352px first-round Met thumbnail with a verified 1920px CC BY-SA derivative and lead every provider group with its manually verified display candidate while retaining unique stored records later in the rotation.
  • [x] Add a canonical public-interaction integrity gate and clear its first defect set: enumerate the route inventory plus sitemap, render 66 routes, distinguish 58 HTML pages from eight machine resources, require one main landmark, inspect 2,784 visible links, probe 898 unique same-origin targets, validate 38 fragments, retry bounded rate limits, stop correctly at external redirect handoffs, and retain source-route diagnostics. Repair two nested landmarks, route protected browser traffic through the resilient custom sign-in page, disclose unavailable OAuth configuration instead of returning Auth.js 500s, and replace two non-resolving curated artwork links with authoritative Getty and Met pages. The optimized local candidate passes with zero interaction failures.
  • [x] Replace the public professional-workspace card matrix with an intent-led institutional entry point: lead with collection, research, and pilot pathways; give every research/source route descriptive purpose and action copy; separate sign-in-required agent, organization, readiness, income-evidence, and worker controls inside an explicit operator boundary; preserve every registered professional destination; and verify one main landmark with zero horizontal overflow at desktop and 390px mobile.
  • [x] Rebuild the public External Evidence Ledger as a human-readable proof surface rather than an internal card dashboard: lead with the current verdict and exact claim boundary, explain verified facts versus open gaps versus live reachability, present every row in a non-compensating claim ledger, move raw receipts into keyboard-native disclosures, separate current production probes from adoption and revenue proof, compact the tombstone history, and state what evidence would change the verdict without altering any underlying status. Add production-length timestamp, live-probe wrapping, scrollbar-safe full-bleed, and explicit small-screen gutter regressions after direct canonical inspection caught otherwise hidden horizontal-overflow defects in dynamic evidence values and `100vw` viewport math.
  • [x] Rebuild the provenance-path assessment as a private editorial decision guide with visible possible outcomes, four-answer progress, a device-local recommendation receipt, explicit non-verdict boundary, preserved deterministic routing, and a desktop/mobile funnel test aligned with both live Stripe checkout and the intentionally disabled email course.
  • [ ] Complete the route-by-route public experience audit, including every actionable control, content state, responsive breakpoint, image-quality boundary, source link, search path, accessibility gate, and conversion path; retain canonical production evidence for each corrected slice.
  • [x] Add a verified, server-only Meta Museum Facebook Page connection and a supervised publisher that creates deterministic source-linked drafts, requires exact digest-bound human approval, keeps durable publishing disabled by default, sends credentials only in the authorization header, and writes token-free post receipts. The first release pairs human-sounding Sun & Rain Works attribution with a high-resolution 2400×1260 digital-cultural-heritage visual built around the Cleveland Museum of Art's faithful CC0 image of Van Gogh's Landscape with Wheelbarrow. Clio, the Greek-history-inspired social editor, performs a bounded evidence-only review, and the native Page photo flow binds the media URL into the approved digest while recording both photo and post IDs.
  • [x] Operationalize Clio's August 22–28 launch plan as a validated seven-post campaign package with exact-plan hashing, source sets, candidate local times, per-post success signals, privacy-safe Facebook organic attribution, deterministic draft digests, and explicit false scheduling/publication authority. Runtime artifacts remain local; each image and post still requires rights review and exact human approval.
  • [x] Finish the public My Collection experience with private browser-local saving, preserved museum sources and rights labels, relevance-ranked cross-provider discovery, a true total result limit, compact source guidance, and graceful provider-outage messaging that never exposes raw API failures.
  • [x] Reframe `/records` as an honest public collection-network page: separate fourteen governed production connectors from the retained normalization sample, expose every source from the canonical registry, remove image-less/held records from the visual gallery, prefer full-resolution assets, and preserve complete records through the dataset export.
  • [x] Replace the Stories placeholder with three complete, clickable, image-led editorial stories: Kōrin close looking, a bounded 1787–88 Houdon/Goya comparison, and uncertainty preserved in two Jacometto catalog titles. Every article separates observation from record facts, carries museum citations, materials and rights, emits Article JSON-LD, and appears in the sitemap.
  • [x] Give Stories an editorial reading hierarchy instead of three interchangeable cards: one artwork-led opening, two deliberately secondary paths, visible source and rights context on every entry, and keyboard-addressable chapter routes within every article. Verify the index and a representative story at desktop and 390px with no broken images or horizontal overflow.
  • [x] Apply the question-led Connections narrative system to every published Story: each turn now asks a concrete question, carries a visible clue-to-clue bridge, identifies its evidence class, and links directly to the museum records used in that turn. Rebuild the lead wave story as a five-turn Met/Cleveland trail whose surprise is a rejected direct-influence hypothesis: Cleveland's record points to Hokusai rather than Kōrin. Use verified 3400px CC0 Cleveland and 3811px Met Open Access images, preserve interpretation and rights boundaries, and verify both the index and story at desktop and 390px with zero broken media or overflow.
  • [x] Add a consent-gated, UTM-bound Stories-to-Research-Kit assist with a free-checklist alternative; retain clicks as funnel evidence only, never revenue.
  • [x] Rebuild the Research Kit offer as a conversion-focused editorial page: remove false card affordances, establish distinct purchase/terms/audience/contents/preview/boundary hierarchy, make all three sample tiles functional in-page links, verify desktop and 375px containment, and bind dedicated CSS into the commercial release digest.
  • [x] Rebuild `/support` as a public-value editorial journey instead of two generic card grids: lead with the evidence promise, link to three inspectable live outcomes, distinguish research/review/infrastructure work, keep terms adjacent to the offer, provide a useful contact path while checkout is inactive, and preserve the exact-digest production activation boundary. Desktop and 390px checks show zero overflow and no nested landmarks.
  • [x] Rebuild `/projects` as a public case study rather than an internal scorecard: lead with the cross-collection problem, correctly label the 25-second tour, provide a three-click live-product journey, explain reader outcomes before architecture, remove stale test/readiness scores, separate deployed capability from outside proof, add a collection-pilot path, and repair the broken architecture-overview link through the live docs renderer. Desktop and 390px checks show zero overflow and no nested landmarks.
  • [x] Repair the pnpm/action-setup v6 workflow conflict by matching every workflow to packageManager pnpm 10.34.5 and enforce the shared pin with a repository-wide regression test.
  • [ ] Renew the Meta Page token, re-run the read-only identity check, and obtain exact-digest human approval for the prepared Kōrin Rough Waves Facebook release packet; do not schedule or publish while OAuth code 190 persists.
  • [x] Implement and verify a Stripe-hosted one-time supporter Checkout Session in the Sun & Rain Works sandbox.
  • [x] Build the production supporter Checkout path behind an explicit disabled release flag, restricted live key, bounded one-time amounts, direct Stripe return verification, and a GET-only four-offer production conversion probe.
  • [x] Preserve sponsor and institutional-pilot intent through allowlisted contact routes and prefilled email subjects; measure consented contact opens as interest only, require the contextual paths in the GET-only production probe, and keep clicks separate from received inquiries, commitments, or revenue.
  • [ ] Obtain exact-digest approval for the supporter release, enable it only in Vercel Production, rerun the conversion probe to 4/4, and wait for genuine external payment plus settlement/fee evidence before claiming revenue.
  • [x] Persist genuine Stripe-delivered sandbox events idempotently, reproduce immutable receipts after restart, and enforce zero verified sandbox revenue.
  • [x] Build the approval-gated single-tenant Microsoft Graph outreach path with encrypted operational data, mailbox pinning, single-use human approvals, immutable send receipts, reply/opt-out suppression, and replay-safe subscription handling. Certificate-backed production reconnection, one controlled reply, and one explicit opt-out are verified against the operator-owned mailbox; suppression receipts are idempotent and unattended sending remains closed.
  • [ ] Renew human approval against the current Research Kit release digest in `docs/ops/research-kit-current-release-review-2026-08-22.md`(ops/research-kit-current-release-review-2026-08-22.md), then probe the exact release; the live account, terms/refunds, durable webhook destination, and sensitive production configuration gates are complete, but no external sale is yet evidenced.
  • [x] Repair the editor-only Org Scope Route Matrix layout: expand the desktop workspace to 96rem, contain the wide table rather than clipping the document, define stable readable columns and sticky headings, and provide semantic route cards for narrow screens with focused regression coverage.

This is the current, authoritative execution plan. It supersedes

development-roadmap.md(development-roadmap.md), which is retained as the legacy

pre-Next.js plan. Completed slice history lives in

progress/era-history.md(progress/era-history.md); the strict evidence checklist

lives in roadmap-to-10.md(roadmap-to-10.md); open engineering risks live in

risk-register.md(risk-register.md).

The architecture north star is

linked-art/LinkedArtSOTAWebApp.md(linked-art/LinkedArtSOTAWebApp.md). Provider,

schema, protocol, and validation work must map tests to fixture anchors in

linked-art/LinkedArtModel1.0-Reference.md(linked-art/LinkedArtModel1.0-Reference.md).

This document owns sequencing and stop/go decisions.

Immediate operational priorities — August 20, 2026

Public UI/UX A+ program: the canonical homepage baseline is now measured rather

than graded by impression alone (910 ms LCP, 0.00 CLS, Lighthouse accessibility

100, best practices 77, SEO 92). The first tested slice fixes localized canonical

metadata and the mobile-menu label mismatch; establishes a single primary hero

journey plus explicit audience paths; makes first-visit analytics choices compact

and even-handed; replaces generic conversion copy; reduces the footer from 25 to

18 links; and prevents oversized small-tablet artwork rows. The next gate is

shared page-header/card/form/state normalization across critical public routes,

followed by full quality gates and desktop/mobile production browser proof. The

baseline, boundaries, and acceptance matrix live in

the public UI/UX A+ program(product/ui-ux-a-plus.md).

The Research Commons public shell now follows that editorial system: a focused

question-led opening, four numbered research moves, an early device-local

workbench, three plainly differentiated help paths, and a separate standards

index replace the previous 17-link header and seven repeated cards. The full

18-journey desktop/mobile proof passes after aligning its Reports assertion with

the current evidence-desk heading; this changes presentation, not the existing

fail-closed research, publication, payment, or human-review boundaries.

The second tested slice separates buyer-facing pilot information from internal

sales operations: `/pilot` retains transparent pricing, scope, success metrics,

support limits, and a concrete conversation path, but no longer publishes the

activation ledger, prospect profiles, named outreach queue, or named target

accounts. Its 412-pixel mobile journey fell from roughly 24 viewports to eight

with no horizontal overflow or data table. The mobile footer now groups its

concise directory into two scannable columns while keeping the brand and public

promise full-width. Commercial-readiness controls now verify the public page's

new honest-status language without requiring internal operating evidence to be

rendered to buyers.

The shared launch accessibility gate now exercises 35 resolved routes at desktop,

a 200%-zoom-equivalent reflow width, and 390-pixel mobile (105 cases), and fails

on WCAG A/AA findings or horizontal overflow. The first expanded run found real mobile width defects in Insights,

the visual ETL mapper, and the Getty workspace; all three are now contained at

390 pixels. Production reruns remain part of the deployment gate. The browser

approval test now reads its generated capture configuration from the correct

download stream, restoring a warning-free lint gate and meaningful receipt

verification. Shared focus treatment now combines an explicit outline with the

existing halo, and reduced-motion users receive effectively instant transitions.

The retained browser gate also traverses the first six keyboard targets on six

critical public journeys and rejects missing, off-screen, or visually unmarked

focus.

Local production Core Web Vitals evidence now covers six critical routes plus a

stressed mobile interaction. Desktop LCP is 120–1,022 ms; stressed mobile LCP is

641–1,349 ms; menu INP is 104 ms; every measured CLS is 0.00. Explore's database

query dominates its 870–915 ms TTFB, so canonical production measurement remains

the decisive performance gate rather than extrapolating from localhost.

Auth host trust is now wired explicitly and fail-closed: only the literal

`AUTH_TRUST_HOST=true` enables forwarded-host trust. The retained performance run

was rebuilt with that opt-in and verified without localhost `UntrustedHost` noise.

The resulting full serial tests, warning-free lint, type diagnostics, and

216-route production build are green for the same local candidate.

The August 20 shipping rerun reconfirmed ESLint, the full serial suite, and the

production build using an ephemeral build-only `AUTH_SECRET`; no environment

secret was written to or staged from the repository.

The following public-journey slice normalizes Stories and Connections to the

shared page spacing, title/lede hierarchy, descriptive metadata, CTA language,

and readable feature width. Stories now labels its formats as previews rather

than implying published editorial content, Connections identifies its one

curated journey without inflating breadth, and Privacy no longer repeats the

analytics-event wording. The focused regressions, public trust suite, ESLint,

full serial suite, and 216-route build are green. Canonical production browser

proof and the five-session comprehension gate remain open.

The first canonical 105-case accessibility run passed every route except

`/iiif` at the reflow and mobile widths, where OpenSeadragon's HTML drawer

injected tile and navigator images without `alt` attributes. The deployed fix

labels each deep-zoom mount as a group and uses a scoped `MutationObserver` to

mark transient canvas fragments decorative. The subsequent canonical rerun

passed all 105 route-and-viewport cases with zero severe axe violations and zero

horizontal-overflow failures, closing this accessibility finding.

Canonical production mobile performance traces now pass the program thresholds

across Home, Explore, a representative artwork, Connections, Research Kit, and

Pilot: LCP is 615–2,489 ms, CLS is 0.00, and mobile-menu INP is 74 ms under Fast

4G and 4x CPU slowdown. The artwork detail is only 11 ms inside the LCP gate and

remains the watch item. Lighthouse SEO is now 100; its remaining best-practices

deduction is solely the documented third-party-cookie boundary on direct Met

artwork images, not a first-party failure.

Research Kit and support checkout returns now have explicit pending, invalid,

and sandbox-success presentation with assistive status/alert semantics and clear

recovery actions, without treating a redirect as payment evidence.

Commit `1706b595` is live on the canonical domains through Vercel deployment

`dpl_9RMjniY8bz56uADexLVHgTTYfkTm`; mobile probes pass all three state contracts

with zero horizontal overflow.

The privacy-safe PD0 evidence importer now encodes the exact A+ comprehension

gate: no more than 5,000 ms homepage exposure, separate offer and primary-action

results, five unique consented non-specialists, and at least four unassisted

joint passes. Generic task completion or `with-help` responses cannot satisfy it.

Linked Art community contribution: Meta Museum now has a conformance-tested

Ferdinand Bol/Rembrandt fixture that enriches the official Unmodeled

Relationships pattern with Getty's exact `ulan1102_student_of` predicate in

`assigned_property`. The case study and minimal upstream patch description are

prepared in

the ULAN relationship note(linked-art/ulan-associative-relationships.md) and

submission packet(ops/linked-art-ulan-upstream-contribution.md). The public

fork commit `84a7203` passes the upstream MkDocs build and pull request

linked-art/linked.art#808 is

open against `v1.1`. Community review and acceptance or merge—not submission or

Slack visibility—remain the evidence required for the A+ community-positioning

goal.

Revenue activation update (August 20): the retained Research Kit sandbox proof

still verifies payment, buyer-route delivery, and zero eligible test revenue.

The operator-owned live Stripe account is connected; a least-privilege `rk_live_`

key successfully accessed Checkout Sessions, and the enabled production webhook

targets `/api/stripe/webhook` with nine payment lifecycle events. Both secrets

are sensitive Vercel Production variables. The public offer and terms now state

immediate fulfillment, the 14-day refund policy, privacy boundary, and disabled

automatic-tax status. Owner approval is bound to release digest

`bfe54344d432de1a8a93a5b1a7219a01ea0b6ead749975dacd2babf97ac831eb` in

the launch approval(ops/research-kit-launch-approval-2026-08-20.md). The next

gate is exact-release deployment plus production probes; neither readiness nor

an operator rehearsal counts as revenue. A four-week attributable distribution packet is prepared for

August 21 through September 17; it authorizes no publication, contact, or spend.

The sponsor outreach gate remains closed with zero approved candidates and zero

messages sent.

Separately, the managed-pilot ledger records nine earlier first messages and no

replies; every recorded follow-up date has passed. The next pilot action is the

bounded, one-time, same-channel copy in

the August 20 follow-up packet(ops/pilot-follow-up-2026-08-20.md), subject to

accountable approval and suppression rules. Do not expand the cold list before

closing this follow-up cycle.

The shipped `main` release is healthy. Commit `3177c6c` built with Next.js

`16.2.12`, generated all 216 routes, and is live on the canonical production

domains through Vercel deployment `dpl_UZ46fBArKEAQAe4y35CcWpsrc28J`. The red

deployment in the dashboard is a separate Dependabot Preview build for

commit `6fa2828`; compilation and TypeScript passed, but page-data collection

correctly failed because `AUTH_SECRET` is scoped only to Production.

P0 — restore isolated preview verification

| Status | Owner | Action | Completion evidence |

|---|---|---|---|

| [ ] | Vercel operator | Generate a new high-entropy `AUTH_SECRET` for Preview only. Never copy or broaden the Production secret. | `vercel env ls` shows `AUTH_SECRET` for both environments without exposing either value. |

| [ ] | Vercel operator | Redeploy Dependabot commit `6fa2828` after the Preview secret exists. | The preview deployment reaches `Ready`; its log contains no missing-secret failure. |

| [ ] | Engineering | Test the dependency branch as an upgrade, not as an ordinary content preview: read the repository-bundled Next.js 16.3 guidance, then run diagnostic TypeScript, lint, the full serial suite, and a production-mode build. | All four gates pass on the exact dependency commit and the result is attached to the PR. |

| [ ] | Engineering | Review changes from Next.js `16.2.12` to `16.3.1`, React `19.2.4` to `19.2.8`, and the accompanying browser, database, telemetry, and type packages. | The PR records reviewed breaking/deprecation notes and either a merge decision or a specific rejection reason. |

| [ ] | Platform | Add a non-secret preview-environment readiness check so future preview failures identify missing required scopes before page collection. | A test or deployment check detects an absent Preview `AUTH_SECRET` while preserving the runtime fail-closed assertion. |

Stop/go rule: do not weaken `auth.ts`, introduce a shared development fallback

in production-mode builds, expose a secret in logs, or merge the dependency PR

merely because its preview becomes green. Merge only after the exact upgraded

branch passes the complete gate. This preview lane is operational maintenance;

it does not block or downgrade the healthy production release.

Next product evidence after P0

  1. Run the approved four-week provenance distribution experiment and retain

consented source/medium/campaign aggregates; do not interpret unavailable

analytics as zero.

  1. Complete the 50-case production AI evaluation and independent validation

program with three qualified researchers and one subject expert.

  1. Obtain real external reuse, qualified acquisition, and provider-confirmed

economics evidence before changing the current `3/11` A+ claim.

Autonomy A/A+ checkpoint — production feedback loops

The fail-closed A+ rubric and operator control surface are documented in

the supervised automation runbook(ops/supervised-automation-a-plus.md).

The revision-bound independent-review packet and strict response validator are

implemented; attributable external review remains the sole A+ evidence gap.

deployment. Ten public research, validation, feed, and revenue routes are

checked against `https://www.metamuseum.org`; every run writes a timestamped,

SHA-256-addressed receipt and uploads it even on failure.

receipts and a fail-closed GitHub issue alert. The monitor remains restricted

to configured HTTPS hosts, response bounds, and named semantic fields; drift

opens a publication hold and never updates claims automatically.

keys, dead-letter state, and operator escalation. Transitions use row locks and

rollback on invalid events; evidence completion always stops at an attributable

human-approval gate. Focused reducer/store/type tests pass. A production

database exercise and scheduled escalation check remain required before this

capability earns production-proven status.

Production proof v2 used Vercel's in-process environment runner, persisted a

row with monotonic timestamps, and stopped at `awaiting_human_approval` with

one visible escalation and no approval. The first run exposed a retrograde

timestamp defect; monotonic enforcement and attributable cancellation were

added, and that flawed row was cancelled rather than counted as proof.

recruitment approval; connect real commercial systems without fabricating

outcomes. Human approval remains mandatory for historical conclusions,

rights, publication, outreach, and financial commitments.

  • [x] Trigger canonical-domain smoke tests after a successful Vercel production
  • [x] Schedule the bounded connection source monitor weekly with immutable run
  • [x] Implement Postgres-backed orchestration state, retry budgets, idempotency
  • [ ] Complete genuine researcher/non-specialist validation after accountable

Autonomous first-profit loop

  • [x] Define the Research A+ evidence gate as eleven non-compensating technical and external requirements. The evaluator emits `A+` only when every requirement has fresh, attributable, non-synthetic proof; otherwise it emits `ungraded`, preserves exact next actions, and authorizes no publication, outreach, email, or payment action. The initial retained baseline is honestly `0/11`, making the remaining external validation and outcome work explicit rather than awarding points for feature breadth.
  • [x] Build the cohesive open Research Commons journey and bind its accessible desktop/mobile browser packet, public methodology/schema, contribution/correction/attribution contracts, and expert-review route to the A+ release evidence. The device-local question-to-packet workflow passes desktop/mobile accessibility and overflow proof, emits canonical digest-bound JSON, proposes only no-side-effect agent tasks, and links to versioned reports and independent review. The A+ command hashes the exact workflow, tests, browser receipt, schema, public policies, guides, and issue templates; local workflow, open-contract, and automation requirements now pass while external requirements remain missing.
  • [x] Replace manual GitHub expert-review transcription with a deterministic, privacy-safe Issue Form importer. It discards account metadata, validates exact public fields, labels, source-and-finding rows, and declarations, retains and replays the bounded issue body, and hands the candidate to separate facilitator verification without inferring qualification or validation.
  • [x] Replace manual public correction transcription with a deterministic, privacy-safe GitHub Issue Form importer. It discards account metadata, validates exact packet targets, public evidence, AI disclosure, and consent, retains and replays the bounded issue body, and emits an unaccepted correction envelope for independent human review without granting publication or impact authority.
  • [x] Connect signed correction-import artifacts directly to receipt-bound correction propagation. The propagation verifier replays the retained GitHub body and nested envelope before deriving a distinct unpublished packet candidate, removing manual extraction while preserving independent acceptance and publication boundaries.
  • [x] Add deterministic mixed GitHub researcher-intake batching. The offline router classifies up to 100 exported expert reviews and corrections by exact label contracts, rejects ambiguous or duplicate issues, strips account metadata through the specialized importers, writes non-overwriting routed artifacts, and replays the complete safe batch manifest before exposing distinct facilitator next actions.
  • [x] Make mixed-intake handoffs directly executable without manual filenames or config editing. Every replayed item carries a deterministic artifact filename and exact command; direct correction propagation verifies retained packet/import file receipts, rejects overwrite and drift, replays the imported Issue Form, and emits only an unpublished candidate awaiting a distinct human decision.
  • [x] Add privacy-minimizing GitHub Discussion category forms for bounded questions, reproduction notes, source discoveries, identity-collision warnings, research gaps, and reversible workflow proposals. The forms require public evidence, AI disclosure, limitations, privacy, and authority boundaries; because the repository Discussion URL currently returns 404, the activation guide forbids public linking until a maintainer enables matching categories and captures attributable rendered-form receipts.
  • [x] Expose the structured public correction Issue Form beside the device-local correction builder. Desktop/mobile Chromium prove the exact off-site URL, new-tab isolation, aggregate form-open event, privacy and non-acceptance copy, accessibility, and no horizontal overflow without leaving the site; the same proof caught and corrected an unrelated 1.92:1 quiet-header-link contrast regression.
  • [x] Complete correction-form funnel instrumentation end to end. The aggregate event is allowlisted, normalized from identifier-free provider exports, retained in the reproducible organic report, and displayed separately in the private cockpit; boundaries and tests prevent an open from becoming a returned or accepted correction, qualified action, validation, impact, revenue, or income.
  • [ ] Run the production AI evaluation and independent validation program: at least 50 cases, three qualified researchers, one subject expert, a materially corrected versioned investigation, frozen-baseline improvement, and external citation/reuse/correction evidence.
  • [x] Project the retained Rosenberg multi-source dossier into deterministic ranked claim-level results with evidence class, confidence, direct HTTPS citations, retrieval timestamps, uncertainty, rights boundaries, and an explicit no-live-search notice. Focused service and page tests protect query-bounded negative findings and the legal/restitution/publication refusal boundary. General multi-archive live search and external validation remain open.
  • [x] Restore the canonical local compiler/lint/test/build chain after strict fixture typing drift. `pnpm typecheck:diagnostic`, `pnpm lint`, the full serial `pnpm test`, and `pnpm build` pass; the build uses the documented CI-only ephemeral auth secret and retains the production missing-secret failure contract.
  • [x] Deliver the first public `/research/query` federation slice: ordinary-language input executes against the current Linked Art museum index, matches the reviewed Rosenberg retained archive package, ranks claim-level citations/uncertainty/rights, and exposes executed, retained, unsupported, and disabled lane status. API validation rejects malformed, oversized, contact-bearing, and raw-query-shaped input. The 16-journey desktop/mobile Research Commons proof passes axe and overflow checks. This is not yet general live multi-archive search; additional approved archive packages/connectors and production deployment proof remain open.
  • [x] Reassess the French and BnF connector boundary against current official documentation. POP/Rose Valland remains manual-import-only because the official open-data inventory does not list that collection and no collection-specific machine/rate contract is documented; public application code is not treated as permission. BnF SRU is documented and Open-Licence eligible as a future bounded discovery connector, but remains unimplemented pending explicit approval to add a new external API.
  • [x] Close the natural-language deployment-proof gap: the exact-release open Research Commons handoff, approval, bounded capture, replay verifier, tests, and operator documentation now require eighteen fixed probes—seventeen public resources including `/research/query`, plus one fixed privacy-safe `POST /api/research/query`. The POST receipt binds its request digest and must replay as ranked Linked Art claims with uncertainty, direct HTTPS citations, executed museum and retained archive lanes, and false consequential authority. A missing page, malformed answer, citation-free claim, or elevated authority cannot satisfy the technical open-release requirement.
  • [x] Register `POST /api/research/query` in the organization-scope route matrix after the full release gate identified the omission. The route is classified as a provenance read over scoped museum records, with the immutable reviewed archive package unable to expose sibling-scope museum records; route coverage, API, and matrix tests pass.
  • [x] Promote dossier disagreements into first-class federated results. `/api/research/query` now returns ranked `conflict` and `not-comparable` entries with exact comparison keys, source lists, direct citations, retrieval times, explicit unresolved boundaries, and false resolution/identity/legal-title/restitution/publication authority. Unsupported archive questions return no inherited contradictions. The public UI renders them separately from claims and refusals; the desktop/mobile 16-journey proof validates the unresolved boundary with zero axe or overflow defects. The production POST receipt now rejects missing or malformed contradiction evidence.
  • [x] Bind deployment-smoke latency to the exact release rather than relying only on transport timeout. Every open-release probe retains a non-negative duration and fails capture/replay above 15 seconds; both fixed natural-language POSTs have a stricter 5-second ceiling. Invalid, missing, negative, non-finite, rehashed, or slow duration evidence is rejected. These production smokes remain explicitly insufficient for the separate 50-case p95/cost evaluation.
  • [x] Production-prove the refusal path separately from the supported answer. The 19-probe exact-release contract adds a second digest-bound privacy-safe POST whose archive lane must be `unsupported`, whose refusal must state no reviewed package matches, and whose result must contain neither inherited contradictions nor archive claims. Capture/replay rejects a rehashed retained-match lane, contradiction, or archive claim. The public refusal journey passes on desktop/mobile, bringing the Research Commons proof to 18 journeys with zero axe or overflow defects.
  • [x] Make multi-source breadth explicit and machine-verifiable. Every federated response now includes a normalized source-coverage ledger mapping source ID/URL, museum or Nazi-era lane, evidence classes, claim and contradiction ranks, and `indexed-result` versus `retained-reviewed-evidence` execution. All rows state `liveQueried: false`; unsupported archive questions expose no archive coverage. The production supported-answer receipt requires at least three distinct retained Nazi-era sources including primary-document and official-catalogue evidence, while replay rejects hidden breadth, duplicate IDs, live-query mislabeling, or a refusal response that leaks archive coverage. Desktop/mobile proof renders and verifies the ledger.
  • [x] Build and browser-prove the approval-gated organic acquisition foundation: canonical robots and sitemap contracts, one flagship provenance workflow, ten search-intent guides, Article and Product/Offer JSON-LD, a no-email-required checklist, and a deterministic device-local assessment routing visitors to free, kit, or bounded-review paths. Encrypted optional double-opt-in lead storage, a digest-approved and suppression-safe nurture worker, bounded human-review inquiry scopes, deterministic PII-rejecting provider-export ingestion, reproducible net-income reporting, a retained production baseline, and exact activation plus 30/60/90-day plans are implemented. The editor-only Organic Income Operator Cockpit and private no-store API add release/source-digest validation, 35-day freshness alerts, assessment starts/completions/recommendations, unknown-not-zero semantics, sandbox exclusion, activation states, exact next actions, and deterministic period packets. Release-bound tests explicitly exercise assessment routing plus indexing, conversion, delivery, and settlement alert states. Desktop/mobile Chromium prove both public funnel and authenticated cockpit accessibility and no-overflow behavior. Production publication, provider verification, lead capture, email sending, and live checkout remain separately blocked on attributable approval and real receipts.
  • [x] Package the organic foundation into a reproducible four-week provenance distribution experiment. Consented guide and funnel events now retain only sanitized source, medium, and campaign attribution; guide views are session-deduplicated; eight niche touchpoint links cover Linked Art, museum/provenance partners, research communities, newsletters, and owned surfaces. The content-addressed packet defines one primary funnel, supporting signals, minimum denominators, learning thresholds, and an explicit boundary that authorizes no external publishing, contact, promotion, or spend. See the experiment runbook(ops/provenance-distribution-experiment.md).

checkout probe, approved/rights-safe content selection, fraud/privacy/rights

stops, a USD 50 total cost ceiling, and a strict settled-external-revenue

definition. Synthetic payments, self-funding, clicks, unpaid invoices, and

receipt-free revenue fail closed.

excluding the unapproved Rosenberg report, research listings, Connections

previews, validation studies, and package-test pages.

merchant identity, price/cadence, cancellation, refund, and privacy terms.

The August 11 live probe confirms checkout is inactive, so no revenue can yet

be collected and the loop correctly reports `configure-production-checkout`.

payments, import immutable commercial receipts and attributable costs, and

continue until independently reproducible verified net profit is at least

USD 1. No profit is currently claimed.

exports from payment, CRM, contract, and invoice systems. Imports are bound to

the complete export SHA-256, emit deterministic revenue/contract IDs, reject

PII and duplicate external records, require operator attestation for qualified

conversations, require attributable prior financial approval for contracts

and invoices, and explicitly authorize no financial action. Actual provider

credentials, production exports, and revenue outcomes remain unproved.

  • [x] Add the first direct-to-buyer product: a USD 29 Provenance Research Kit with a transparent offer page, fixed-price Stripe Checkout Session, product-specific metadata, consent-gated checkout-start analytics, and a versioned Markdown workbook. Fulfillment retrieves the Checkout Session server-side and refuses missing, unpaid, wrong-price, wrong-currency, or wrong-product sessions; production checkout remains inactive until a restricted live key and operator-approved commercial terms are configured.
  • [x] Extend the commercial evidence model with distinct research-kit checkout/payment events, processor fees, gross and net kit revenue, a content-addressed sandbox proof command, and a release-digest-bound launch-readiness handoff. Sandbox proof uses the authenticated Stripe CLI, verifies the exact paid test-mode offer metadata, and exercises both the fulfillment handler and buyer-facing HTTP download route without trusting a stale local key. The handoff recomputes retained proof integrity and binds the conversion and revenue-model sources into its release digest. It performs no activation and remains blocked until live restricted-key configuration and attributable merchant/refund/privacy/tax/fulfillment approval exist.
  • [x] Add a deterministic daily loop with immutable receipts, a canonical live
  • [x] Add `/sitemap.xml` for approved public utility and revenue surfaces while
  • [ ] Configure an operator-owned production checkout whose hosted page states
  • [ ] Accumulate genuine eligible sessions and provider-confirmed settled
  • [x] Add the commercial-system ingestion boundary for normalized production

The repository also includes a self-contained, leadership-facing architecture

and evaluation snapshot at

`metamuseum-project-architecture-overview.html`(../metamuseum-project-architecture-overview.html).

It summarizes the implemented stack and workflows while preserving the strict

boundary between local engineering capability and unearned production,

audience, or revenue claims.

The homepage trust-page regression test now accepts the repository's supported

one-provider fixture as well as multi-provider data by matching the rendered

`museum source` / `museum sources` grammar; CI no longer fails when managed

storage contains records from exactly one provider.

A+ public-content program

  • [x] Replace the graph's 483-row relationship wall with a searchable, progressively disclosed index. The default page now exposes 24 relationship rows and 55 main-region links instead of 973, retains all source/relationship/target links on demand, and reflows its Cytoscape canvas without horizontal overflow at a 390 px viewport.
  • [x] Reframe `/entities` as public cultural discovery and replace its 477-row DOM wall with eight-item server-addressable windows per entity type. The measured route now renders 66 rows and 90 main-region links instead of 477/491, adds people/place/material question paths, retains search/facets/tenant scope and previous/next access to every entry, states the source-identity boundary, and passes a 390 px filtered-page overflow check.
  • [x] Reframe `/insights` from an operational analysis dashboard into an evidence-bounded artwork journey. Three reader questions lead into recorded dates, places, and relationships; each visualization identifies museum facts or interface summaries, rejects unsupported influence and movement claims, and exposes recorded places as keyboard-operable filters when place data exists. Desktop and 390 px checks show no horizontal overflow.
  • [x] Reframe `/issues` from a 100-row operational triage queue into a public standards decision trail. The route now explains how Linked Art proposals evolve, separates GitHub source facts from Meta Museum interface readings, keeps all 321 issues searchable in twelve-item pages, focuses the selected interpretation for keyboard users, and reduces the default main-region load from 102 links/103 buttons to 14 links/15 buttons without removing source access.
  • [x] Replace `/roadmap`'s internal slice ledger and filesystem-derived date with a current public plan organized as Now, Next, and Not yet proven. The route derives its completed public-content count from the authoritative roadmap, links five inspectable discovery surfaces, retains maintained-source and structured-JSON access, and explicitly refuses to turn technical capability into claims of adoption, endorsement, revenue, or impact. Desktop and 390 px layouts remain overflow-free.
  • [x] Make the Rosenberg research deployment self-contained: Vercel retains the five exact tracked JSON inputs imported by production code while continuing to exclude all other generated evidence, with a clean packaging regression test.
  • [x] Complete Stage 1 of the external provenance-research network: six versioned, validated research-only source records now cover scope, geography, dates, searchable fields, access, language, reuse rights, primary-document availability, discovery provenance, and fail-closed automation. See the registry contract(product/provenance-research-source-registry.md).
  • [x] Complete Stage 2 as a safe manual federated-research lane: dated source-specific access assessments keep all six sources manual-only; bounded queries, raw-envelope SHA-256 receipts, Rosenberg collision controls, agreement/conflict/not-comparable matrices, fixtures, provenance-governance ownership, and release-ineligible outputs are enforced. See the Stage 2 workflow(product/provenance-federated-research.md).
  • [x] Complete Stage 3 with a shared fail-closed network contract and one officially supported connector: Getty Provenance Index SPARQL is bounded, CC0-attributed, allowlisted, timeout/retry/size/result limited, kill-switched, receipt-backed, operator-gated, and permanently research-only. A live exact-label smoke passed; five unverified sources remain explicitly disabled. See the connector runbook(product/provenance-network-connectors.md).
  • [x] Expand Getty Rosenberg discovery with fixed, bounded templates for Linked Art names/identifiers, actor types, objects, title-transfer sales activities, date ranges, and known entity IDs. The first retained live run finds one Getty Group and 25 sales-activity leads; review and primary-document linkage remain open.
  • [ ] Complete the decisive Rosenberg research benchmark across at least three sources, including manual receipts for the five non-network sources, primary documents, a reviewed dossier, an uncertainty-first interface, and genuine researcher/non-specialist reproducibility testing.
  • Publishing surface: the homepage now exposes a reader-first research spotlight, and the reusable report leads with the finding, significance, changed understanding, and unresolved question before the audit layer. A six-beat trail distinguishes documented February/May correspondence and the RBS identity anchor from the later-reported September handover, shows the bounded V.A.8 correction, and adds the official POP record's two hashed object images without calling them photograph 3492 or handover proof. Permanent HTML, report JSON, JSON Feed 1.1, and RSS 2.0 keep AI-assisted `research-leads` separate from researcher-controlled `reviewed-findings`; the latter remains empty until attributable review. Feed items expose researcher-decision state, limitations, missing evidence, an explicit validation boundary, and exact methodology, correction, and contribution routes. Four shared participation contracts route agents and researchers into reproduction, correction, method challenge, or independent review with exact evidence and privacy requirements, while denying publication, contact, novelty, legal, payment, impact, and income authority. The report page renders the same contracts as responsive evidence cards with direct, instrumented actions; browser proof covers accessibility, no horizontal overflow, exact GitHub targets, privacy, authority, and the unchanged noindex state. That proof exposed and removed a nested duplicate `main` landmark inherited from the report page. Canonical metadata, article social metadata, and citation-aware `Report` JSON-LD now present the proposed answer only as an under-review suggested answer; desktop/mobile proof parses the live JSON-LD and confirms no accepted answer, while indexing remains unavailable until publication approval. The pending value-release candidate binds the exact changed service and page digests. Canonical research routes bypass locale rewriting so dynamic report slugs and citations do not 404. Research pages, API routes, and the novelty command are registered in operational ownership controls. Email and webhook delivery remain future work.
  • Approval-bound organic handoff: SEO candidates now require the publication envelope's validation release digest to equal the current research release, closing stale-but-valid approval reuse. A replay-verified ready candidate can produce one non-overwriting, privacy-safe organic-publication manifest with exact canonical, robots, sitemap, report-channel, structured-data hash, citation count, internal-link, feed, implementation-target, and verification expectations. A separate no-write verification mode rebuilds the complete manifest from the retained SEO candidate and rejects semantically forged state even when an attacker recomputes the outer hash. The manifest retains no private validation rows and authorizes no file write, publication, deployment, indexing, sitemap submission, outreach, or payment.
  • Reports-index gate: `/research/reports` now remains canonical but `noindex,follow` and absent from the sitemap while no attributable publication-approved reviewed finding exists. Its state-honest `CollectionPage` JSON-LD exposes lead/reviewed counts and boundaries. Only reports satisfying publication-approved status, reviewed-findings channel, explicit approval, a dated accountable reviewer, and a retained publication decision with approver code, public HTTPS evidence, source-report/validation SHA-256 bindings, approval time, and four affirmative content decisions can activate hub/report sitemap entries; toggled flags without the decision fail. Both hub and detail page are direct Research Commons release-manifest members; desktop/mobile proof covers canonical, robots, JSON-LD, WCAG, and overflow.
  • Current evidence: dossier v1 answers an actor-level question across Getty, ERR, Legacy Explorer, Rose Valland, Lost Art, Proveana, the official RBS, Archives diplomatiques, and the MoMA Paul Rosenberg Archives. Human-operated browser searches retain visible-text hashes, expanded-query behavior, bounded negatives, aliases, and collision warnings. RBS entry *5034 / OBIP 37.954 matches Braque, the mandoline/fruits/bottle description, 97 × 130 cm, and Paul Rosenberg as named owner. The diplomatic correspondence supplies object-specific primary corroboration for Brussels recovery and intended restitution. The updated official JDP-0056 record embeds two image files now retained locally with SHA-256 receipts, providing relevant image-based registration evidence; image-specific republication rights remain uncleared. A full review of MoMA V.A.8 found no secure photograph-3492 match and positively disconfirmed its strongest apparent candidate as _L'intérieur au vase noir_. The dossier remains a draft because final physical handover and independent review/testing remain unproved.
  • Interface evidence: `/research/rosenberg` translates supported ordinary-language questions into visible fixed searches, shows uncertainty and all source outcomes before synthesis, rejects raw query syntax, and remains no-index. `/research/rosenberg/dossier` presents the frozen answer, primary evidence, claims matrix, source ledger, receipt hashes, disagreements, missing evidence, and limits for open checking; `/research/rosenberg/validate` presents the revision-bound independent-review and participant protocol. Genuine usability and reproducibility outcomes are still unmeasured.
  • External-validation protocol: the current dossier revision plus primary, manual, 1947 archive-bridge, and official image-registration receipt hashes; independent-review task; three-researcher counterbalanced timing study; five-non-specialist open-response study; human coding rule; and fail-closed thresholds are frozen in `artifacts/provenance-research/rosenberg-validation-protocol-v1.json`. The evaluator rejects synthetic, unconsented, duplicated, revision-mismatched, undersized, self-reviewed, slow, irreproducible, or poorly understood evidence. The validation workspace supplies a client-local independent-review return form that binds the review to the frozen digest, requires three distinct HTTPS sources, captures corrections and a decision, and requires the facilitator to attest that a private identity-to-code mapping exists. A separate qualified-researcher study randomizes baseline/interface order, times both conditions, requires three citations and a substantive answer per condition, retains the facilitator-confirmed washout, and exports a richer revision-bound session. `pnpm research:validation:assemble` now validates and projects returned files into an evaluator bundle while retaining SHA-256 receipts for every review, researcher response, raw reader response, and human-coding decision; it rejects duplicates and refuses overwrite. None of these workflows uploads responses or establishes real-world eligibility by itself. No real rows have been entered yet.
  • Participant collection: `/research/rosenberg/study` is a no-index, client-local five-minute comprehension flow. It collects explicit consent and eligibility, generates a pseudonymous code, opens the frozen report, and downloads two open responses bound to the dossier digest without uploading, retaining, or scoring them. The evaluator now retains both human coders' decisions and requires a distinct adjudicator for disagreement. Desktop and 375 px browser checks pass with no horizontal overflow; five genuine returned sessions remain required.
  • Current-cycle recommendation packet: impact is limited to the public research-study journey and Rosenberg validation contract. Highest risks are self-selected participants, response-file loss before facilitator handoff, and private attribution mappings not being retained with review custody. A versioned unsent recruitment packet names three current official professional routing channels, exact professional/non-specialist/coder copy, uncompensated and privacy disclosures, non-endorsement language, and send-evidence rules. The approval-ready copy is now personalized to the canonical public operator Joseph Chirum and `art@sunandrainworks.com`; it discloses that email return metadata identifies the sender to the facilitator. Its exact dossier, protocol, candidate, and copy digests plus requested scope are frozen with `approval: null`. `pnpm research:recruitment:approval` now reports only `Accountable human approval is required before recruitment outreach`; nothing has been sent. After genuine approval, recruit one qualified independent reviewer, three eligible research practitioners, five eligible adults, two coders, and an adjudicator while retaining provider receipts and recruitment channels. Acceptance remains one attributable independent review, three counterbalanced researcher records, five unique revision-matched reader files, complete two-coder decisions, and a passing fail-closed evaluator.
  • Novelty research: the reusable `research:novelty` workbench independently scores object-match signals and chronology conflicts while refusing verified identity, novelty, or publication. A newly discovered collision—two distinct 1938 Braques with identical 81 × 100 cm dimensions on adjacent Mangin references—corrected RBS *5036 versus the Lasker work from 85/99 to 60/99 (`possible`) and removed the qualified chronology conflict. The MoMA finding aid narrows decisive inspection to III.A.2.2.42, III.D.1.c, III.D.1.h, V.A.2, V.A.12, II.QQ.1, and II.QQ.8; a complete reproduction request is prepared, but MoMA reference intake is closed until September 8, 2026. The official French restitution workbook also exposes a 1945 Braque/Rosenberg row superficially overlapping the 1947 POP record JDP-0056. The correct primary target is now `209SUP/1`, dossier 45.15, views 199–1173—not the rejected E–I volume `209SUP/392`. Its inventory distinguishes the 1945 judicial-method and 1947 Belgium restitution groups, so the dates alone are not a contradiction. Internal pages 33–38 and 45–48, OBIP 37.954, Mangin pages 33–34, and transaction records remain required.
  • Primary-document advance: `209SUP/1` internal page 33 (viewer media 1154) identifies a Braque _Nature Morte_, 97 × 130 cm, found with a Brussels dealer and claimed by Rosenberg; page 48 (media 1173) identifies _Mandoline, fruits et bouteille_, reports Brussels recovery, and describes restitution as intended after formalities. The retained hashes and continuous dossier sequence create an object-specific primary bridge to RBS *5034 and JDP-0056. Two official JDP-0056 images now close the relevant-image registration gap, but a signed September receipt is still required to prove final physical handover and photograph 3492 remains necessary to establish the cited archive-photo lineage.
  • [x] Enforce five explicit reader-value promises and three-headline/two-opening package preparation.
  • [x] Reject synthetic, specialist, unconsented, and undersized package-test evidence.
  • [x] Put the story path before assurance detail without removing citations or boundaries.
  • [x] Instrument headline impression, open, 25/50/90% depth, continuation, share, save, and learned-something signals.
  • [x] Add a privacy-safe production evidence importer with denominators, monotonic-funnel checks, explicit thresholds, and independent-review/operator-release gates.
  • [ ] Complete five real non-specialist Rosenberg package sessions; three headlines and two openings are prepared and correctly remain ineligible for drafting.
  • [x] Reconcile Rosenberg identities and dated roles without collapsing Galerie Paul Rosenberg (Group) into Paul Rosenberg (Person); retain a sourced `operated-by` bridge and fresh live receipts for all three objects.
  • [ ] Complete contradiction, historical-context, visual-rights, editorial, subject-matter, and operator reviews.
  • [ ] Observe a valid production window and earn independent 95/100 with every rubric dimension at least 9/10.

See connection-reader-value-and-audience.md(product/connection-reader-value-and-audience.md).

Instrumentation and fixtures never count as A+ outcome evidence.

---

Status (evaluation baseline retained; operational checkpoint August 20, 2026)

<!-- BEGIN:PROJECT_STATS -->

<!-- Generated by `pnpm docs:stats`; do not edit by hand. -->

| Generated project stats | Current value |

|---|---|

| Next.js | `16.2.12` |

| React | `19.2.4` |

| App page files | root homepage + `76` non-root page files (`77` total) |

| API route handlers | `175` `app/api` route handlers |

<!-- END:PROJECT_STATS -->

Executive Assessment

Meta Museum is a strong, unusually complete Linked Art product and a credible

controlled-beta system. It is not yet defensible as institution-grade or

repeatable SaaS. The gap is mostly evidence quality, operational history, and

commercial proof rather than missing core product breadth.

| Category | Score | Evidence-based assessment | Immediate implication |

| ------------------------------- | ----------------------------------------------------------: | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------ |

| Product experience | 6.5/10 for the public; 8.2/10 for museum and research users | The source-backed professional journeys are coherent, but public discovery still exposes workspace language, operational controls, and specialist navigation before establishing a reason to explore or return. | Separate the public museum from the professional workspace and build one curiosity-led discovery loop before adding more platform breadth. |

| Linked Art and data semantics | 9.0/10 | Canonical JSON-LD, event-centric modeling, rights, provenance, equivalent identities, HAL relations, IIIF, and provider boundaries are deeply implemented and test-backed. | Sustain conformance; do not add another provider until readiness work clears. |

| Evaluator and trust story | 8.4/10 | `/projects`, `/evidence`, `/datasets`, `/docs`, and machine-readable APIs make the work inspectable. The public evidence ledger currently mixes contradictory artifact baselines. | Repair proof consistency before adding more evaluator copy. |

| Accessibility | 9.3/10 | `pnpm a11y:check` passed 18 routes with 0 violations on July 12. Auth redirects for protected routes were handled as expected. | Keep zero severe violations and add accessibility checks to any hero or navigation polish. |

| Engineering architecture | 8.0/10 | Boundaries, contracts, strict TypeScript, storage abstractions, tests, and evidence automation are mature. The route and script surface is large and operationally expensive. | Prefer consolidation and owner clarity over new surface area. |

| Reproducible local quality gate | 6.5/10 | Direct ESLint passed and the local review-goals gate passed, but the audit began with a stale closeout guard, direct TypeScript failed in generated `.next/dev/types/validator.ts`, and a direct full test run did not finish within 10 minutes. | Re-establish one clean, repeatable canonical gate before feature work. |

| Documentation and governance | 8.5/10 | The docs surface is rich, agent-readable, and guarded for drift. The previous roadmap had become a completion ledger rather than a decision document. | Keep this file current and concise; send completed detail to history. |

| Operational readiness | 6.0/10 | Production preflight is 20/20 and launch review is 7/8, but long-window SLO, uptime, KPI, and artifact-coherence proof remain red. | Clear deploy-environment proof, then let time-based evidence accumulate without manufacturing results. |

| SaaS and business readiness | 5.5/10 | The managed pilot offer and tenant-aware technical foundation are credible. There is no invoice-backed pilot, repeatable onboarding, retention, or margin proof. | Sell and execute one bounded concierge pilot before building self-serve billing. |

Overall decision: **8.1/10 as a local product and portfolio case study; 5.8/10

for strict public-production readiness.** The project is safe to keep demoing and

iterating. Broad institution-grade or profitable-SaaS claims remain blocked.

The July 21 live review confirmed that the primary public journeys render without

browser console errors. The same review found that `pnpm typecheck:diagnostic`

failed in test contracts and fixtures. Those type errors are now repaired:

`pnpm typecheck:diagnostic`, 59 focused tests, `pnpm lint`, and `pnpm build` pass.

The full serial `pnpm test` run completed `1,677/1,677` tests in `157.5s` on

July 28 after stale deployment-preflight and documentation-contract expectations

were aligned with canonical evidence. `pnpm review:goals:check`

now reports `22` production-proof blockers (`0` local-refresh, `0`

deployment-environment, and `22` real-world-evidence). Completing the canonical

local test run is no longer a P0 blocker.

The July 28 timeout investigation confirmed `1,704/1,704` tests pass in 154

seconds with a compact reporter. The canonical test script now uses that

reporter to prevent captured verbose output from stalling the command channel.

August 5 quality-gate refresh: P0.2 is green again after the 24-hour closeout

guard expired. Lint passed in `19.3s`, diagnostic TypeScript in `5.4s`, the full

serial suite in `150.4s`, and the Next.js production build in `50s`. The repair

keeps intentionally untracked `data/source-repos` mirrors out of tracked-text

drift scans and anchors synthetic SLO samples to each evaluation timestamp so

the gate remains reproducible after the original June fixture window. The three

new standalone evaluator briefings are preserved in the same clean worktree.

The architecture evaluation briefing now marks the canonical-gate risk closed

and carries these measured results instead of its earlier stale warning.

Readiness Boundary

reports `status: external-evidence-required`, `local gate status: passed`, and

`strict 10/10 gate status: failed`.

launch review passes `7/8`, with the launch-review Era C dependency still red.

paid-pilot proof, retention, and gross-margin evidence.

  • [ ] Strict 10/10 readiness is not green yet: `pnpm review:goals:check`
  • [x] The July 10 production deployment preflight passes `20/20`; the latest
  • [x] The latest strict handoff reports `0` local-refresh, `0` deployment-environment, and `22` real-world-evidence blockers, or `22` production-proof blockers in total.
  • [x] The evidence ledger, readiness dashboard, Era C, long-term, launch-review, and review-goals artifacts now share the same canonical blocker, SLO, uptime, and ActivityStreams values.
  • [ ] Remaining strict proof includes 30-day SLO/uptime evidence, production KPI exports, durable ActivityStreams syndication including a genuine `Delete`,
  • [x] Current launch status is governed by `pnpm review:goals:check`.
  • [x] `pnpm review:goals:check` must be green before broad public SaaS claims.

The allowed claim is: **local gates pass and strong external proof exists; strict

10/10 production readiness remains blocked by external evidence.**

Evaluation Findings That Change Priority

  1. Public evidence consistency is repaired and regression-guarded. The

canonical handoff, README, roadmap, and live ledger now agree on `21 = 0 + 3

generated ActivityStreams evidence, so its pending Delete row cannot repeat

confirmed`Create`or`Update` types as missing.

  1. The local gate is not yet reproducible from one clean command sequence.

The closeout guard, a malformed generated Next dev validator, and a long-running

direct test invocation obscure whether a fresh checkout is truly green.

  1. The product breadth is sufficient. Fourteen provider lanes, 34 pages, 146

API routes, public docs, validation, reconciliation, ActivityStreams, IIIF,

agent review, and tenant-aware pilot controls are enough for the next learning

cycle. More breadth would dilute the evidence and customer work.

  1. The next product-polish gain is quality, not more sections. The current

hero can enlarge a low-resolution source image, while Core Web Vitals and

primary-journey performance do not yet have a concise public baseline.

  1. The public product needs a distinct reason to return. Cross-provider search

is useful, but it does not yet turn the underlying data advantage into stories,

surprising discoveries, personal collections, or connections that no single

museum site can show.

  • 18`strict blockers. The ledger also reconciles partner confirmations with

---

Priority Plan

Priority is determined by trust and dependency, not by implementation novelty.

Do not start a lower tier while an actionable higher-tier exit condition is red.

Time-bound external evidence may continue accumulating in parallel.

P0 - Restore A Trustworthy Baseline (Now, 0-7 Days)

| ID | Outcome | Owner | Exit criteria | Verification |

| ---- | ----------------------------------------------------- | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |

| P0.1 | Keep public readiness evidence internally consistent. | Platform + Evidence | `/api/evidence/ledger`, `/evidence`, `/readiness`, review-goals, launch review, and the long-term artifacts agree on SLO sample/day counts, uptime, ActivityStreams observed/missing types, and current blocker scope. Fallbacks disclose missing artifacts instead of substituting incompatible values. | `tests/services/readiness-evidence-consistency.test.ts`; focused evidence-ledger, readiness, review-goals, and documentation-drift tests; `pnpm evidence:ledger:probe:check`. |

| P0.2 | Restore one clean local quality gate. | Platform | From a fresh generated state, `pnpm session:closeout:check`, `pnpm lint`, `pnpm test`, `pnpm build`, and `pnpm typecheck:diagnostic` all complete successfully. The generated `app/api/vanda/search/route.js` validator fragment is valid after regeneration, and test duration is recorded. | Run the five canonical commands and attach elapsed time plus the first failing test if any. |

| P0.3 | Clear deployment-environment proof drift. | Operations | Rerun production preflight, launch evidence, launch review, and public Era C evidence with production environment present; reduce the deployment-environment lane from `4` blockers to `0` without changing the real-world claim boundary. | `pnpm launch:preflight:production`; `pnpm launch:evidence:production`; `pnpm launch:review:production`; `pnpm era-c:exit-gate:public`; `pnpm review:goals:check`. |

| P0.4 | Preserve nightly k6 evidence artifacts. | Platform + Evidence | The Docker fallback writes `artifacts/performance/k6-slo-summary.json` as the host runner user, deletes stale summaries before each run, and fails when no fresh summary is produced. | `pnpm exec tsx --test tests/scripts/k6-slo-runner.test.ts`; confirm the next Era C workflow ingests one fresh SLO sample. |

Production database SSL drift was repaired on July 27: Vercel now uses the

existing `sslmode=verify-full` connection value, the production artifact was

redeployed, and `/api/ai/query` returned `200` through the public alias.

P0 exit gate: all local commands are reproducibly green, public evidence has no

cross-surface contradictions, and deployment-environment blockers are zero.

P1 - Convert Reliability And Demand Into Proof (Next, 1-6 Weeks)

| ID | Outcome | Owner | Exit criteria | Verification |

| ---- | --------------------------------------- | ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------ |

| P1.1 | Complete the 30-day reliability window. | Operations | One canonical source reports 30 distinct UTC days of complete passing deployed SLO samples, including the cold-record scenario, and at least 99.9% public-read uptime with failed rows aged out of the retained window. | Scheduled probes plus `pnpm longterm:evidence:public` and `pnpm era-c:exit-gate:public`. |

| P1.2 | Produce real SOTA KPI evidence. | Data + Curation | Production/Postgres or warehouse exports meet the reconciliation auto-approve and reviewed-precision thresholds; every capture row identifies its production source. | `pnpm monitoring:kpi-evidence:production`; `pnpm era-c:exit-gate:public`. |

| P1.3 | Close one invoice-backed managed pilot. | Founder + Product | One real buyer has a signed scope and invoice reference, one collection is activated within seven days, and support, required KPI, retention, and gross-margin rows are captured without placeholders. | `pnpm pilot:buyer-pack`; `pnpm pilot:activation`; `pnpm pilot:support`; `pnpm pilot:kpi`; `pnpm pilot:evidence --check`. |

The architecture evaluator now participates in the commercial pre-revenue claim control: it must disclose the absent invoice-backed pilot, remain `External evidence required`, and point to `pnpm pilot:buyer-pack` plus the tenant-scoped `pnpm pilot:evidence --check` acceptance gate. This improves the handoff but does not count outreach or local tooling as revenue proof.

August 5 ownership and operator-experience pass: the compact public mobile header and product-specific agent sign-in context are implemented and regression-tested. The operations-risk report now covers `35/35` page routes across `7` ownership/consolidation lanes alongside `61/61` API families and fully classified package-script namespaces. `pnpm ops:profile` makes the Next.js + Postgres portable baseline explicit and keeps Python services, AG2, Solr, GraphDB, and publication workers optional until their readiness gates justify enablement. The evaluator marks mobile and authentication closed while keeping broad surface area and specialist topology honestly managed rather than eliminated.

| P1.4 | Preserve honest ActivityStreams adoption. | Platform + Partnerships | Keep `3/3` real external consumers, `3/3` verified durable callbacks, and zero rejected subscriptions fresh. Add `Delete` only after a genuine upstream `404`/`410` tombstone and partner read; never synthesize it to satisfy the gate. | `pnpm providers:coverage:seed`; `pnpm activity:tombstone:scan`; `pnpm activity:subscriptions:guard`; `pnpm activity:syndication:evidence`. |

| P1.5 | Raise visible product quality and establish a performance baseline. | Product + Frontend | Home, Explore, one artwork detail, Projects, and Pilot pass mobile/desktop visual review; no hero image is rendered above a defensible intrinsic size; no text overlaps; LCP <= 2.5 s, CLS <= 0.1, and INP <= 200 ms on the agreed production profile. | Refresh public-trust screenshots, run a production performance audit, rerun `pnpm a11y:check`, and retain the metrics artifact. |

P1.5 performance checkpoint (July 28): all ten production cold-load traces pass

the agreed lab budgets. Explore mobile is the limiting LCP at `2,286 ms`, Pilot

desktop is the largest CLS at `0.0665`, and the representative interaction trace

is `28 ms`. The route matrix and profile are retained in

`docs/ops/frontend-performance-baseline.md`(ops/frontend-performance-baseline.md).

P1.5 remains open pending the mobile/desktop visual review and fresh axe run.

Step 6 remediation (July 28): the fresh axe run passes `18/18`. The two Home

blockers are fixed locally: the source band is now a normal full-width sibling

of the padded content container, and V&A IIIF services are promoted to a

1,200 px image derivative before thumbnail fallbacks. The repeated local

production-build matrix passes all ten mobile/desktop traces with 0.00 CLS, a

maximum 1,628 ms LCP, and 29 ms INP; the 18-route axe gate also remains green.

Repeat this matrix against the deployed revision to close P1.5.

Linked Art 1.1 watch checkpoint (July 28): the standalone

`linked-art-1-1-agenda-impact-tracker.html`(../public/linked-art-1-1-agenda-impact-tracker.html)

maps all 26 August 5 agenda issues to current support, expected impact, required

fixtures or schema changes, and pending community decisions. Update its decision

column and the canonical Linked Art reference ledger after the meeting before

changing validators or production mappings. Issue #637 now has an internal,

provenance-bearing

confirmed-negative reconciliation contract. It keeps curator-confirmed

non-matches separate from unresolved candidates and withholds Linked Art

projection until the community settles the property name and assertion pattern.

Issue #362 now has a provisional internal response-profile contract. Its

server-defined brief projection preserves canonical identity, marks itself

incomplete, and links deterministically to the full record; no new public

profile parameter is enabled before the community decision.

The pre-meeting implementation evidence packet now combines issues #362, #637,

and #780 with executable references and decision questions. Use it during the

August 5 discussion, then replace its pending questions with resolution links

before promoting any candidate behavior.

Linked Art 1.1 meeting checkpoint (August 19): the saved reference checkout now

tracks the upstream `v1.1` branch at `fded7e7`, 31 commits beyond the prior

`3ed503b` master snapshot. The standalone

`linked-art-1-1-august-19-2026-meeting-briefing.html`(../public/linked-art-1-1-august-19-2026-meeting-briefing.html)

records the supplied logistics and 21 agenda issues, current milestone counts,

post-August-5 label changes, issue #637's same-day reference-placement question,

and tested decision prompts. Agenda proposals and Meta Museum implementation

evidence remain explicitly non-normative until the community records outcomes.

An August 20 post-meeting refresh verifies that the reference branch is still at

`fded7e7` while the milestone has grown from 43 to 45 open issues. It records the

explicit #366 agreement to place the relationship on the affected thing; keeps

#524 and #637 gated with their clarified file-level and concept-versus-individual

boundaries; adds #804 and #806 to the watch list; and cites the Getty auction

example now recorded in #493.

The machine-readable meeting decision ledger covers all 26 agenda issues.

Issue #366 is now the sole resolved row. Agenda proposals are recorded separately

from outcomes; resolved rows require a

matching Linked Art issue URL, target release, and explicit local action before

they can drive post-meeting changes.

The Step 3 conformance pass now covers nine focused patterns under

`tests/fixtures/linked-art-1.1/`: qualified `AttributeAssignment` ambiguity,

inscribed `Name` evidence, `Name.created_by`, prototype-level provenance for an

unenumerated `Set`, member-side Addition and Removal, Person Joining and

Leaving, and auction selling/purchase separation.

Endpoint-family inspection now preserves the active terms-ontology inverse links

`added_member_by` and `removed_member_by`. The lifecycle fixture declares its

extension context explicitly and stays non-normative until the 1.1 meeting

decision is recorded.

The compatibility pass is now executable through

`LINKED_ART_1_1_COMPATIBILITY_BOUNDARY` and

`tests/quality/linked-art-1-1-compatibility-boundary.test.ts`. The audit keeps

six representative pending property placements rejected, leaves proposed and

deferred classes outside the endpoint map, and documents the post-meeting

promotion procedure in

`docs/linked-art/1.1-compatibility-audit.md`(linked-art/1.1-compatibility-audit.md).

P1 exit gate: 30-day reliability and production KPI rows pass, one real paid

pilot reaches first value, and strict ActivityStreams evidence is either complete

or explicitly waiting on a genuine upstream tombstone with all other rows fresh.

July 27 checkpoint: the three accepted durable callback rows are restored and

`pnpm activity:subscriptions:guard` passes `3/3`. Syndication remains honestly

blocked only on a real `Delete` activity read.

The nightly Actions environment now supplies all three production consumer IDs

to `activity:adoption:matrix`. Its production verification passed `12/12` feed

probes and resolved `3/3` declared consumers; `Delete` remains the sole missing

observed activity type.

A write-enabled July 27 tombstone scan checked 68 canonical upstream targets

across 14 provider lanes with zero errors and found no genuine `404`/`410`.

Accordingly, no `Delete` was minted; the scheduled scan must continue until a

real upstream removal can be observed and read by the three consumers.

The nightly workflow now runs the durable callback guard exactly once through

`activity:syndication:evidence`; the redundant standalone guard step was removed

without weakening its failure behavior.

Collection and readiness now have separate Actions semantics. The nightly

evidence workflow succeeds when probes and artifact generation work even if the

recorded status is red. The following `Era C Readiness Gate` workflow reports

those known external-evidence blockers without producing a failed scheduled job;

a manual dispatch remains fail-fast and owns strict Era C thresholds, durable

callback enforcement, and production launch review.

The Actions matrix now uses `actions/setup-node@v7`; execution-policy tests own

that major consistently, and runtime file metadata no longer depends on an

overload-derived Node type that can become optional in newer type packages.

Public Discovery Product Track (Next, staged behind the P0 gate)

Goal: turn Meta Museum's cross-provider data advantage into a welcoming public

museum built around curiosity, storytelling, and repeat visits. The professional

workspace remains available, but it must no longer dominate the anonymous public

journey. The defining promise is **connections no single museum website can

show**.

This track does not authorize a new provider, runtime dependency, or fully

autonomous publishing path. Each phase ships behind the existing rights,

provenance, accessibility, performance, citation, and accountable project-

operator release controls. Outside specialists improve assurance but are not a

prerequisite for ordinary source-backed research and bounded public experiments.

Connections use three explicit assurance tiers:

  1. Independent research — agents may discover, reproduce, rank, and draft

candidates without outside reviewers. Results remain internal and make no

novelty, historical-causation, rights-clearance, or scholarly-validation claim.

  1. Operator-reviewed public experiment — the project operator may release a

narrowly factual Connection when every material claim resolves to museum

sources, fact/inference labels and uncertainty are visible, media rights are

safe, sensitive claims are absent, and the page says it has not received

external expert review. This is the default independent operating lane.

  1. Externally validated — distinct qualified editorial, rights, and subject-

matter reviewers approve retained evidence. This tier is required before

claiming scholarly novelty, expert validation, sensitive provenance or

identity conclusions, externally cleared rights, or commercial licensing.

The project may advance through tiers 1 and 2 independently. Missing external

review blocks only tier 3 claims; it must remain visible as missing evidence and

must never be filled by an agent or fixture identity.

Independent-lane implementation checkpoint (August 9): the typed Connection

contract, public page, and machine record now expose assurance tier, external-

review status, and a revision-bound operator-release record. The current Flowers

journey is visibly labeled independent research and not externally reviewed.

Promotion to an operator-reviewed experiment requires a non-synthetic `OP-...`

human operator code, timestamp, release notes, and completed citation, rights-

boundary, sensitivity, and disclosure checks; agents cannot satisfy the record.

Live-source checkpoint (August 9): the reviewed independent-lane manifest made

four representative channel calls plus four candidate-source calls and retained

timestamped, hashed receipts for a Met API

record, bounded Getty SPARQL result, Getty IIIF manifest, and Getty Linked Art/LOD

record. All four returned 200 with expected media types; the capture report has

zero failures and zero ingestion rejections. Native payloads are receipt-only

until channel-specific mappings are reviewed, so capture evidence does not imply

local Linked Art conformance.

Independent discovery checkpoint (August 9): a four-record Met/AIC normalized

input passes deterministic pattern analysis and produces a medium-confidence Van

Gogh entity-reconciliation candidate plus bounded provenance-coverage notices.

Agent-assisted research ranks two public-story candidates: matching museum-

supplied year and maker metadata for _Wheat Field with Cypresses_ and _The

Bedroom_, and matching supplied wave labels for _Rough Waves_ and _Under the Wave

off Kanagawa_. The shortlist retains live hashes, facts, inferences, uncertainty,

rights boundaries, and refused novelty/causation/sensitive/licensing claims. Both

remain independent research with zero publication actions pending operator review.

Draft-contract checkpoint (August 9): both shortlisted candidates now satisfy the

full `CuratedConnectionJourney` contract, including synchronized story/claim

references, research method and rejected hypotheses, assurance disclosure,

pending revision-bound operator release, source-controlled rights, corrections,

licensing, and measurement. They remain in a separate draft collection; tests

prove neither slug is available from the public resolver or feed. The operator

packet at `docs/product/independent-operator-release.md` is the next gate.

Source-monitor checkpoint (August 9): `pnpm connections:source-monitor --

--check` compares semantic claim fields for all four Met/AIC sources while

excluding volatile response timestamps. The first live run is clean for 4/4

sources with no publication hold. Any field change, fetch failure, unsafe

redirect, invalid JSON, or oversized response opens a hold and requires a new

journey revision plus operator decision before release.

| Phase | Outcome | Initial scope | Exit criteria | Verification |

| ----- | ----------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |

| PD0 | Establish the public baseline and content boundary. | Record first-time task completion, artwork-to-artwork continuation, return visits, public-domain downloads, and share events. Define the minimum quality bar for public records: usable image, intelligible title/date, source, attribution, and resolved reuse message. | Baseline artifact exists; public and professional audiences, routes, vocabulary, and analytics events are explicitly separated; records below the quality bar cannot enter featured feeds. | Public-route inventory, analytics event contract, quality-filter fixtures, privacy review, and five non-specialist usability sessions. |

| PD1 | Separate the public museum from the professional workspace. | Public navigation becomes Home, Explore, Stories, Connections, My Collection, and About. Evidence, APIs, agents, imports, annotations, org status, and operational controls move behind one clearly labeled “For museums and researchers” entry point. Rewrite the homepage in plain language around art and discovery. | A first-time visitor can explain the product and reach an artwork without encountering workspace status or specialist implementation language; professional routes remain directly reachable and unchanged in capability. | Mobile and desktop visual review, keyboard pass, `pnpm a11y:check`, public-navigation tests, and moderated five-second comprehension checks. |

| PD2 | Ship the minimum delightful discovery loop. | Add `Surprise me`, a rights-safe Artwork of the Day with a stable dated URL, and public artwork pages led by image, essential facts, “Why this is interesting,” related works, and previous/next discovery. Collapse technical metadata and researcher feedback below the public story. | The complete loop works: entry point → artwork → short story → related discovery → another artwork. Featured records never have broken media or ambiguous reuse messaging. | Deterministic selection tests, record-quality tests, mobile/desktop E2E, share-preview checks, and measured artwork-to-artwork continuation. |

| PD3 | Launch Connections as the signature feature. | Discover non-obvious event, provenance, exhibition, authority, material, temporal, and graph relationships; reject bare shared-word, maker, year, place, material, or classification matches. Present the strongest findings as visually led, layered articles with accessible stories and expandable evidence. | At least three operator-reviewed public experiments span three or more institutions collectively. Each passes frozen rubric v1 at 95/100 or higher with no dimension below 9/10, demonstrates knowledge unavailable from one source page, contains no unsupported causal or sensitive claims, and preserves human-controlled release. | Triviality-filter and rubric tests; live hashed source receipts; candidate contradiction checks; independent final score; citation, rights, accessibility, adversarial, and journey E2E checks; five-session comprehension and production engagement evidence. |

| PD4 | Add source-backed Stories and guided exploration. | Publish short image-led stories using reusable formats such as “One artwork, three details,” “Same year, different worlds,” “A disputed identity,” and “How an object changed hands.” Add exploration by subject, place, century, and color; add mood only as clearly labeled interpretation. Hide empty maps and timelines. | A minimum viable editorial cadence is sustainable; every factual claim resolves to a source; guided filters return useful results; empty analytical surfaces do not appear publicly. | Story-schema and citation tests, filter-quality samples, editorial review log, structured-data/share-card validation, and completion-rate measurement. |

| PD5 | Make trustworthy reuse and participation useful. | Add public-domain image download with source, rights, attribution, and metadata. Add “What changed?” with plain-language record version comparisons. Allow local-first saved collections, then optional account sync and read-only sharing after demand is observed. | Downloads package correct rights context; record changes identify source, time, and change origin; a visitor can save and share a coherent collection without being forced to sign in first. | Rights/download fixtures, version-diff tests, local-storage and account-migration tests, privacy review, and collection share E2E. |

| PD6 | Prove retention before expanding. | Evaluate artwork continuation, Surprise Me use, story completion, connection opens, downloads, collections created/shared, and 7-day return visits. Improve the strongest loop; retire or revise weak entry points. | Two consecutive measurement windows show a credible repeat-use signal and no regression in accessibility, performance, rights, or citation quality. Any further personalization or recommendation work has a measured hypothesis. | Analytics review, usability replay, public performance/a11y matrix, editorial quality audit, and a written continue/change/stop decision. |

Recommended first release: PD0 + PD1 + PD2 + three PD3 journeys. It must

demonstrate one complete public loop:

Interesting entry point → beautiful artwork → understandable source-backed
story → unexpected cross-museum connection → another discovery → save or share.

Public discovery stop conditions: pause expansion if featured-record quality

cannot be guaranteed, if connection evidence cannot support the displayed claim,

if rights context is separated from a download, if public pages regress the

agreed accessibility or performance budgets, or if measured use shows no

improvement after two iterations. Missing outside reviewers is not an independent-

research stop condition; it prevents promotion to the externally validated tier.

PD1 implementation checkpoint (August 6): the primary public navigation is now

Home, Explore, Stories, Connections, My Collection, About, and one quiet “For

museums and researchers” entry. Anonymous Explore and artwork journeys no longer

render organization status or professional workspace chrome, and public Explore

suppresses import prompts, roadmap language, and provider implementation notes.

The homepage now leads with cross-museum discovery in plain language; dedicated

Stories, Connections, My Collection, and professional-workspace landing pages

make every navigation destination intentional. Automated navigation, homepage,

workspace-boundary, and Explore acceptance tests are green. The local production

build passes, the 18-route accessibility matrix reports zero severe violations,

and a 375 px browser review finds no horizontal overflow. A fresh deployed visual

review and non-specialist comprehension sessions remain PD1 evidence tasks rather

than reasons to reopen its implementation scope.

PD2 implementation checkpoint (August 6): `/surprise` selects from the same

image-backed, publication-eligible local artwork pool as the homepage and sends

visitors directly into a public artwork journey. `/today` resolves to a stable

UTC-dated `/today/YYYY-MM-DD` page whose selection is deterministic for that date.

The homepage exposes both entry points. Anonymous artwork pages now lead with

“Why this is interesting,” keep facts and source detail in an expandable section,

offer previous, next, Surprise Me, and related-artwork paths, and withhold

researcher annotations and operational relationship tools. Signed-in researchers

retain the complete professional view. Selection and surface acceptance tests are

green. The production build passes, the 18-route accessibility matrix reports

zero severe violations, Surprise Me resolves into an eligible local artwork, and

homepage, daily, and artwork routes show no horizontal overflow at 375 px. A

deployed review remains necessary before promoting this local checkpoint to

production proof.

August 7 deployed checkpoint: PD1/PD2 is live on the production alias. The full

serial test suite, lint, diagnostic typecheck, local and Vercel builds,

production 18-route axe audit, 66-check crawler preview, 20/20 deployment

preflight, public Explore smoke, repeated public-trust screenshot baseline, and

zero-advisory dependency audit pass. Ten retained Lighthouse captures report

100 accessibility and 0 CLS; the desktop routes pass the LCP budget, while the

stricter Lighthouse mobile profile reports 3.7-5.9 second LCP and keeps P1.5

open for remediation and an agreed-profile recapture. The refreshed adoption

matrix passes 12/12 operator-run endpoint probes for all three named consumer

IDs and remains blocked on genuine `Delete`. Those probes do not substitute for

fresh reads made by the external consumers themselves: the retained declared

consumer reads are outside the 30-day adoption window, so Era C correctly

reports `0/3`. Production preflight is 20/20 with zero deployment-environment

failures; the remaining launch-review blockers are time-bound SLO samples and

real-world adoption/KPI evidence.

The concrete production preflight is zero-failure on deployment

`dpl_2WH2w1hrM4zftM84kn3ouU3xu4N4`. The broader `review:goals:local` roll-up now

also reports zero deployment-environment blockers: aggregate launch-review and

Era C wrappers inherit real-world-evidence scope, while concrete preflight,

auth, smoke, IIIF, and k6 failures remain deployment-scoped when present. The

remaining 21 strict blockers are explicitly time-bound or human/external

evidence rather than deployment configuration failures.

August 7 PD0/PD3 checkpoint: the five public outcome events now have a typed,

consent-gated contract that excludes direct identifiers and professional

routes. First artwork completion, artwork continuation, 24-hour return,

rights-qualified download selection, and successful share are instrumented.

A dated baseline artifact and five-session non-specialist

protocol are present, while production observation and the five human sessions

remain open. PD3 has exactly one curated journey—Flowers across two centuries—

spanning Getty and Met records with citations, a high-confidence metadata

label, an explicit no-influence boundary, contract tests, and an editorial

decision packet awaiting external sign-off. That packet now represents the

optional externally validated tier: it separates editorial, rights, and subject-

matter decisions, and `pnpm connections:first-review` rejects direct identifiers,

synthetic evidence, duplicate reviewers, and incomplete approvals. Its absence

does not block independent research or operator-reviewed public experiments, but

the journey must disclose that external expert review and novelty validation are

absent. Before adding journeys two and three, implement the operator-release

record and public assurance-tier label, then keep claims inside the tier-2 boundary.

August 9 A+ quality reset: the prior Flowers, Van Gogh, and Waves concepts score

38, 40, and 43 under frozen rubric v1 and fail the triviality gate. They are no

longer eligible merely because their citations are reliable. PD3 now prioritizes

20 live-source candidates across at least three institutions, requires two

independent substantive signals plus knowledge unavailable from one source page,

and promotes only candidates reaching 95/100 with no dimension below 9/10.

The first deeper-discovery pass added a fail-closed 20-candidate portfolio gate

and expanded live capture from 8/8 to 14/14 jobs across four institutions. Exact

provenance-role detection and API/SPARQL/IIIF/Linked Art research surfaced Paul

Rosenberg and Wildenstein three-museum leads. Rosenberg now has typed actors, a

sourced person-gallery bridge, dated events, fresh receipts, and zero authority

contradictions; it remains unqualified pending real package observation,

human editorial review, and audience evidence. The automated visual-rights

preflight is fail-closed: one receipt-backed Getty image may display, while the

AIC surrogate awaits image-level terms and the copyrighted Braque is withheld;

accountable rights sign-off remains human.

The real-session handoff now includes a blinded five-session field guide and a

fail-closed importer; it records no direct identifiers and cannot create or

substitute participant evidence. Ties and continuation below 60% now block the

draft, preventing the team from overriding observed package preferences.

A no-index, client-only test surface now randomizes option order and downloads

candidate-bound responses without server collection; the CLI directly merges

those envelopes and rejects files from another candidate or schema version.

The consent step now validates both attestations at form submission rather than

depending on controlled-checkbox state, closing an interaction defect found in

live local-browser review. Static/type checks pass; that browser run remained

inconclusive because its React tree did not hydrate anywhere on the page.

Five synthetic specialist lanes then scored the unfinished Rosenberg work from

2/10 (comprehension evidence design) to 8/10 (non-obviousness/reproducibility).

The consolidated report is `docs/product/rosenberg-five-agent-specialist-review.md`.

Immediate changes removed unsupported promotional/causal wording, diversified

the frozen package frames, withheld the AIC surrogate pending image-level terms,

and added no-JS/download failure boundaries. Agent scores remain non-qualifying.

August 8 cultural-intelligence checkpoint: the first journey now produces three

synchronized representations from one typed contract: a public visual story, a

research dossier exposing fact/inference labels, method, uncertainties, rejected

hypotheses, novelty status, and rights boundary, plus an evaluation-only JSON

record with the same claim/citation graph and human-review state. Consent-gated

story-completion and reuse-interest signals are instrumented outside the five PD0

outcomes. Institutional usefulness, agent/editorial minutes, independent novelty

verification, five usability sessions, and attributable external approval remain

evidence needed for externally validated or commercially validated claims; they

do not block the independent operating lane.

The expansion gate is now executable through `pnpm connections:evidence`: its

privacy-safe artifact requires observed completion plus sharing/reuse, qualified

editorial sign-off, institutional usefulness, agent/editorial labor and cost,

independent novelty review, a real price response, and the valid synchronized

machine record. It reports tested gross value before labor separately from labor

minutes and cannot call that result profit. No real evidence has been imported,

so promotion to externally validated or commercially validated status remains

unauthorized. Independent research and operator-reviewed public experiments

remain authorized within their stated tier.

Internal hidden-pattern work can now proceed without violating that public gate:

`pnpm connections:patterns` emits a review-only collection-intelligence report

from normalized records plus explicit source rows. Deterministic candidates cover

equivalent-record conflicts, possible entity reconciliation, shared materials,

owner/custodian or set references, structured-provenance coverage gaps, and

geographic contrasts. Every lead carries citations, confidence, and a refusal

boundary; uncited records are rejected and unsupported demographic or market

conclusions are listed as refused analyses. No second public journey was added.

The review-only report now also consumes explicit Linked Art event evidence:

matching exhibition identifiers, `used_specific_object` groupings, structured

acquisition/transfer parts, dated event places, and shared activity actors. Its

event graph preserves source-record IDs on every edge. Geographic sequences are

movement candidates rather than transport claims; ownership histories do not

assert completeness, authenticity, custody, or legal title.

Additional explicit-evidence candidates now cover alternative maker assignments,

reversed event timespans, repeated technique identifiers, separate `represents`

and `about` iconographic concepts, and unidentified depicted people. Boundaries

prevent authorship resolution, invented corrected dates, workshop/influence

claims, collapsed depiction semantics, or demographic/underrepresentation

inference from these candidates.

The machine layer now also exposes `/api/cultural-intelligence` as a lifecycle-

aware collection feed. Each item preserves revision, production provenance,

review history, corrections, and separate editorial/licensing decisions. The

feed currently reports one evaluation item and zero licensable items; approval

metadata must be complete and all corrections resolved before eligibility can

change. Underlying source-record and media rights remain explicitly separate.

Five derivative formats now compile from the same versioned claim graph:

newsletter, daily feed, narrated visual essay, classroom package, and licensed

article. Evidence sections preserve claim/citation IDs and all formats repeat the

rights boundary. The derivative endpoint returns only an HTTP 409 release

manifest—not internal copy—while the first record lacks editorial and licensing

approval. Format availability is therefore implemented but audience demand,

quality, labor, accessibility, and price remain unproven external evidence.

Value-based pricing now has an executable evidence ladder through

`pnpm connections:pricing`: hypothesis, tested-no-signal, market signal, one

invoice-validated delivery, and repeatable price evidence. Repeatability requires

three scoped offers and two paid, accepted, value-confirmed, positive-contribution

deliveries across buyer segments. Labor is fully costed at an attributable rate;

interest is never revenue. No real offer artifact exists yet, so pricing remains

unvalidated.

The `pd0:evidence` intake command now validates a real GA4 export and moderated

session records, rejects direct identifiers or invented counts, and derives

completion only from five sessions plus attributable product-owner approval.

Its package namespace, script, artifact directory, and product-governance owner

are registered in the executable evidence-ownership and operations-risk controls.

Responsive browser review found the initial action block below the full image on

mobile; it now precedes the image and remains overflow-free at 390 px and 1440 px.

The Vercel ignore contract excludes local provider source mirrors, the `.tools`

binary cache, and the upstream `linked.art` checkout except its required schema

subtree. However, deployment `dpl_GiNFmNpo2UZs4uGEm4Y3B54Bya3b` still archived

79,077 files (669.1 MB), so CLI archive filtering remains an open packaging

optimization; the deploy itself completed and passed its runtime build.

P2 - Productize Only After The Pilot Loop Works (Later, 6-12 Weeks)

| ID | Outcome | Trigger | Exit criteria |

| ---- | --------------------------------------------------------- | ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |

| P2.1 | Guided organization onboarding and first-value dashboard. | One invoice-backed pilot completes activation and its friction is documented. | A new managed org can be provisioned with a sample or customer dataset in under 15 minutes; progress and first value are visible without engineering inspection. |

| P2.2 | Repeatable subscriptions and usage visibility. | Pricing, support load, and gross margin are validated on at least one pilot. | Checkout or invoice-backed subscription sync, webhook/audit evidence, quotas, usage, billing state, cancellation reason, and customer portal are supportable. |

| P2.3 | Institution procurement package. | A buyer starts security/legal review. | Deployment-specific subprocessors, DPA/legal artifacts, access review, incident drill, retention controls, backup/restore proof, status reporting, and SLA/SLO packet are buyer-reviewable. |

| P2.4 | Production agent bridge decision. | A named operator accepts the review workload and risk boundary. | AG2 bridge has explicit sign-off, eval evidence, rollback, auditability, and human-publication approval. A2A/AG-UI remain deferred. |

---

Readiness Scorecards

Launch Readiness

| Lane | Score | Decision | Next evidence |

| --------------------------------------- | ---------: | ------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |

| Internal development and portfolio demo | 9.0/10 | Safe to use and present with the strict-readiness caveat. | Reproduce the canonical local gate and resolve the public evidence inconsistencies. |

| Controlled public beta | 8.4/10 | Technically credible on Vercel + Neon, but the formal beta gate remains evidence-red. | Clear P0, then maintain narrow acceptance criteria while the 30-day window accumulates. |

| General public production | 6.5/10 | Do not claim complete readiness. | Passing long-window SLO/uptime, production KPI, and coherent launch evidence. |

| Institution-grade / strict 10/10 | 5.5-6.0/10 | Blocked by external and time-based proof. | `pnpm review:goals:check` passes with no production-proof blockers. |

SaaS Readiness

| Lane | Score | Decision | Next evidence |

| ------------------------- | -----: | ------------------------------------------------------------ | ---------------------------------------------------------------------------------------- |

| Technical SaaS foundation | 7.0/10 | Strong enough for concierge pilots. | Prove onboarding, support load, usage, and tenant operations with one buyer. |

| Paid pilot readiness | 8.2/10 | The offer and operator path are ready; revenue proof is not. | Reply or qualified follow-up, signed scope, invoice-backed entitlement, real activation. |

| Self-serve SaaS readiness | 3.0/10 | Deferred. | Start only after the pilot validates pricing and activation friction. |

| Profitable SaaS business | 4/10 | Credible wedge, but repeatable revenue is not proven yet. | Retention, support minutes, infrastructure cost, conversion, and gross-margin evidence. |

The primary wedge remains the Managed Linked Art Launch Pilot for small and

mid-size museums, archives, galleries, digital-humanities labs, and artist estates

that need standards-compliant collection publication without a semantic-web team.

Manual invoicing is correct for the first 1-3 pilots. Creator-side provenance and

self-serve billing stay deferred until the B2B pilot loop produces evidence.

---

Evidence Workstreams

| Workstream | Current state | Completion condition |

| -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |

| Local quality | Review-goals local status passes and direct ESLint passes; full canonical reproducibility was not demonstrated in the July 12 audit. | P0.2 is green from a clean generated state. |

| Deployment | Preflight `20/20`, both Render probes, all public smokes, and deployed k6 pass; launch review is `7/8`. The strict handoff has zero deployment-environment blockers because its remaining launch and Era C wrappers depend only on real-world evidence. | Preserve the green concrete deployment matrix while the real-world evidence windows mature. |

| Long-window SLO and uptime | The latest strict handoff reports `11` retained deployed SLO samples across `10/30` distinct UTC days, with `11` passing samples and no failed or incomplete rows in the active report window; the Era C artifact still reports only `15/30` samples toward its exit gate. | Continue distinct-day collection until the complete 30-day threshold is genuinely met. |

| ActivityStreams | Operator endpoint probes pass for three named IDs, but the retained real external consumer evidence is stale and therefore counts as `0/3`; genuine `Delete` evidence is pending. | Collect three fresh real external consumers covering `Create`, `Update`, and `Delete`, with verified callbacks. |

| Production KPI | Local enrichment is promising; production reconciliation distribution and reviewed precision are incomplete. | P1.2 passes the SOTA KPI acceptance rows from named production sources. |

| Managed pilot | The offer, runbook, entitlement, activation, support, and evidence tooling exist. The latest no-pricing buyer pack is specific to the recorded Te Papa outreach and has a real account, organization, and owner, but no paid-pilot tenant or invoice-backed entitlement exists. | Obtain a real buyer reply plus signed scope or invoice reference, provision the tenant, then use P1.3 tooling to record activation, retention, support load, and margin evidence. |

The buyer-review surface now includes a standalone ten-record demonstration at

`/museum-linked-art-pilot-demonstration.html`. It uses traceable public API records

to show source preservation, event-centric Linked Art JSON-LD, rights review

boundaries, validation findings, and museum questions. It proves a review pattern,

not a completed customer engagement or permission to reuse source images.

A one-page buyer brief at `/managed-linked-art-pilot-brief.html` now packages the

problem, five-day process, required inputs, deliverables, privacy and security

boundaries, and post-pilot decision into a printable pre-call handout. It links to

the ten-record demonstration and preserves the same evaluation-only claim boundary.

The first-call workflow now has a timed guide at

`/museum-pilot-discovery-call-guide.html`: a one-minute permission-based opening,

five fit questions, a boundary recap, three explicit decision paths, and a

follow-up record. The call qualifies a bounded pilot before any product tour and

links directly to the buyer brief and demonstration when supporting proof is useful.

The post-call handoff now has a public-data request at

`/museum-pilot-data-request-template.html`. It supplies a copy-ready museum message,

accepts CSV, JSON, XML, LIDO, or a public API, distinguishes minimum from optional

fields, excludes credentials and restricted material, and records the reviewer,

publication boundary, transfer method, receipt evidence, and agreed deletion date.

The conversion and delivery packet now adds a counsel-review sample agreement,

four-level introductory pricing, and a reusable results report. The public pilot

offer uses the same `$0` evaluation, `$3,500` fixed paid pilot, implementation from

`$12,000`, and ongoing service from `$1,250` monthly hypothesis, so buyer surfaces

no longer conflict. These remain unvalidated prices until invoice-backed delivery,

acceptance, retention, and gross-margin evidence exists.

The refined demonstration presents one primary path on desktop and mobile:

collection record, Linked Art mapping, validation, then reviewable result. Source

evidence is collapsed beneath the interaction, and the final state names open

museum decisions and the human publication gate.

The current outreach ledger records the eight user-confirmed August 5 submissions,

their real recipient or form channel, zero assumed replies, and August 12 follow-up

dates. `docs/sales/museum-outreach-pipeline.md` contains eight unsent follow-up

drafts plus a second official-source-researched group of eight prospects that must

remain `research_only` until first-round feedback is reviewed and sending is

authorized. `docs/ops/paid-pilot-commercial-evidence-process.md` closes the

invoice-to-margin capture design without treating placeholders as proof.

Standing Evidence Controls

`generatedAt` separate from source `sourceUpdatedAt` and `sourceUpdatedDoc`

checksum metadata.

`wikidataexplorer-metamuseum-prod` and the other real consumers distinct,

retain durable callback evidence, and preserve the zero rejected subscriptions

state without allowing placeholders to satisfy strict proof.

`pnpm longterm:evidence:public` output remain strict gates; a frontend Core Web

Vitals baseline is added in P1.5 rather than inferred from API SLO evidence.

remain separate scopes. No aggregate badge may silently promote one scope into

another.

  • Public docs metadata freshness: `/api/docs/manifest` must keep response
  • ActivityStreams onboarding ledger: partner rows must keep
  • Performance evidence: the cold-record budget, 30-day SLO depth, and
  • Claim boundary: local success, deployment success, and real-world success

---

Product And Engineering Guardrails

  1. Linked Art JSON-LD remains canonical; UI DTOs are projections at boundaries.
  1. Preserve rights, source attribution, provenance, multi-value arrays, event

semantics, carrier/content/surrogate separation, and opaque URI handling.

  1. Adapters do not import each other; provider parsing stays in adapters;

cross-provider mapping stays in `src/utils/artwork-builder.ts`; contracts remain

leaf modules.

  1. AIDD + TDD remains mandatory for behavior changes. Standards-critical work

cites reference rounds and fixture anchors before implementation.

  1. Public publication and agent-generated claims require citations, refusal paths,

audit evidence, and human approval.

  1. Cultural-intelligence candidates use an executable deterministic routing policy:

weak candidates archive without an agent, strong ambiguity permits at most one

bounded evidence-packet pass, detector reports derive their own scoring inputs,

and every traceable review-queue publication route remains human-gated.

  1. Probabilistic cultural-intelligence signals are versioned, thresholded,

fixture-calibrated, source-backed review candidates with zero LLM calls; they

never establish identity, influence, movement, meaning, or historical truth.

  1. Cultural-intelligence ingestion validates source envelopes and Linked Art,

retains immutable hashed snapshots, normalizes only a comparison projection,

and recognizes exact authority IDs without network or LLM calls.

  1. Review-ledger runs are append-only and attributable; only fully approved

entries generate synchronized JSON-LD, API, dossier, timeline, and accessible

page artifacts, all still requiring an operator to publish.

  1. Scheduled cultural capture evaluates explicit cadences and enforces HTTPS

allowlists, redirect/media/size bounds, hashed receipts, and separate

replay/live evidence before ingestion.

  1. The integrated production-like replay proves four of five candidates route

without an LLM (80%), one bounded pass is allowed but not executed, all nine

routine stages are deterministic, and external outcome evidence remains null.

  1. Cultural graphs preserve cited entity/activity/concept/equivalence edges and

their assertion certainty; contradictory rights become review candidates,

while unknown rights remain a separate reuse blocker.

  1. Completion readiness is machine-audited: technical controls and fixture proof

remain distinct from attributable live captures, human decisions, production

outcomes, observed costs, institutional usefulness, and commercial signals.

  1. Provenance rules detect broken dated transfer chains and ownership before

production; multiple-agent review is executable only for consequential

conflicts with measured value above cost and two independent typed results.

  1. Keep Next.js, React, TypeScript, custom CSS, Postgres/JSONB, Solr, GraphDB, and

canonical ID decisions locked as documented in CLAUDE.md(../CLAUDE.md).

  1. No new runtime dependency, provider, service, database, or architecture era is

started while P0 is red without explicit approval.

Deliberately Deferred

economics and activation are real.

by measured scale or customer evidence.

  • New provider integrations beyond the current 14 production lanes.
  • Self-serve signup, checkout, billing portal, and growth automation before pilot
  • Synthetic ActivityStreams `Delete` evidence.
  • Broad production agent autonomy or public publishing without operator sign-off.
  • A microservice, triple-store, vector-store, or framework expansion not justified

Active 30-day revenue experiment

The supporter, direct-sponsor, and institutional-pilot funnel is implemented

locally with transparent public offers, a one-page sponsor packet, consent-aware

revision-bound/deduplicated browser events, five first-party-sourced sponsor

candidates, an exact unsent message, human approval/send-evidence gates, and a

fail-closed revenue-report command. It does not yet

claim a live checkout, approved outreach, inquiries, contracts, invoices,

payments, accepted delivery, or revenue. The next gate is accountable human

approval of the production revision, introductory pricing, target list, and

outbound copy, followed by a predeclared exact 30-day window and honest outcome

report. See ops/30-day-revenue-experiment.md(ops/30-day-revenue-experiment.md).

The complete local WCAG browser audit passes, while the public-trust visual

smoke retains one expected privacy-baseline failure for human review. Direct

production probes currently return 404 for support, sponsor, and sponsor packet,

and show the previous pilot copy; release and the 30-day clock therefore remain

unstarted. A declaration command now prevents the clock from starting without a

released git SHA, production analytics identity, exact dates, and human approval.

Rosenberg primary-document checkpoint

The MoMA V.A.8 returned-paintings folder has now received a complete bounded

review, not merely a finding-aid citation. Pages 2-65 were inspected; the

apparently promising Braque sequence on pages 4-8 was disconfirmed by its own

notarized declaration as _L'intérieur au vase noir_. High-resolution review of

the grouped catalogue sheets likewise produced no secure photograph-3492 or

target-title match. This useful negative is retained with a file hash and rights

boundary in `rosenberg-primary-document-receipts-v1.json`. It narrows the next

research action to photograph 3492, RA1, or a handover receipt and prevents the

project from presenting V.A.8 as object-specific evidence. The case still awaits

independent human review and does not claim scholarly novelty.

Research-outcome instrumentation checkpoint

The Rosenberg report now has consent-gated, revision-bound measurement for

report view, actual historical-trail 25/50/90% depth, successful share,

browser-local save, “learned something new,” and study start. Actions are

session-deduplicated, carry no direct identifiers, and saving remains available

without analytics consent. Desktop interaction and a 375 px runtime check pass

with no horizontal overflow. Production probes remain decisive: the report,

reader study, support, and sponsor routes return 404 after locale routing; only

the older pilot page returns 200. Therefore current audience, comprehension,

support, and sponsor outcomes remain zero/unobserved rather than inferred from

local instrumentation. An accountable approved release is the next dependency.

Release-decision checkpoint: `value-release-candidate-v1.json` consolidates the

exact dossier, receipt, protocol, report, study, engagement, support, sponsor,

pilot, and production-probe hashes. Tests recompute every digest and the

candidate fails closed with eight incomplete human attestations, no operator

code, no approved commit, and `releaseEligible: false`. The companion

`product/value-release-decision.md` gives the operator one bounded decision and

makes clear that approval does not authorize novelty, legal, image-rights,

outreach, audience, or revenue claims. No further local feature is required to

begin the experiment; genuine approval and deployment are now the limiting

inputs.

---

Cadence And Ownership

Research Commons organic acquisition — local release candidate complete

The private AI question-collection approval workspace now converts the first evidence-operations blocker into an exact, attributable human decision. Four explicit approvals, a public attributable HTTPS evidence record, a pseudonymous accountable code, and a substantive note are canonically SHA-256 signed in-browser and downloaded locally. Browser validation now reuses the authoritative public-HTTPS predicate and rejects local, private, and example/test/sandbox/demo/fixture hosts before download; the question-publication and open-release deployment workspaces share the same rule. The workspace has no persistence or activation authority; approval verification, deployment, and runtime enablement remain separate. This improves operator usability but earns no external A+ point by itself.

The acquisition workstream now also exposes a private exact-contract editorial workspace covering all six question-led guides, their claims, sources, limitations, FAQ data, indexing effects, and excluded claims. Its downloaded approval is canonically signed and mutation-sensitive, while publication, indexing, sitemap inclusion, campaign activation, and deployment remain separate human actions. It removes review friction without counting editorial approval as traffic, validation, or income.

Both approval workspaces now avoid approval-only choice architecture: an accountable reviewer can sign a rejection or changes-required decision with all approval dimensions false. Rejections remain integrity-verifiable while every activation and publication checker stays fail-closed.

The external-validation workstream now has a device-local JSON preflight for returned validation and evidence-envelope files. It checks the exact pseudonymous three-researcher/one-expert mapping, facilitator verification, consent, independence, synthetic exclusions, direct-identifier keys, receipt binding, corrected-v2 packet, novelty decision, and published limitations. The privacy-safe receipt excludes file content, and the authoritative CLI repeats the shared inspection before frozen-study projection.

The conventional-baseline manifest now closes the post-outcome-freeze loophole: it requires an accountable `REG-…` preregistration code, attributable HTTPS record, SHA-256 receipt, and a preregistration timestamp that predates all observed sessions. Digest-valid paired results cannot pass if the study was only frozen after outcomes were visible.

The production AI evaluator now rejects any bundle whose dataset freeze or independent label approval does not strictly predate every production trace. This prevents expected claims and refusal labels from being retrofitted after observing model behavior, while retaining the existing 50-case, 10-answer/10-refusal, two-reviewer, citation, latency, and cost gates.

The production AI evidence chain now shares the strict public-HTTPS validator used

by other A+ lanes. Question-export attestations, manifest preparation, frozen dataset

sources, label approvals, run receipts, citations, and independent reviews reject

localhost, private/link-local addresses, `.local`/`.invalid`, and

example/test/sandbox/demo/fixture hosts. Focused service and subprocess fixtures use

non-fixture public-shaped hosts, while explicit regression cases prove that rehashed

non-public evidence cannot advance assembly or scoring.

External impact now derives the reviewed-release timestamp from the signed validation artifact and requires every citation, substantive reuse, or accepted correction to occur strictly afterward and no later than evaluation time. Events can no longer be retroactively attached to a release that did not yet exist or projected from future timestamps.

Six question-led research methods, canonical Article/FAQ metadata, internal

qualification paths, aggregate cockpit events, and a digest-bound publication

gate are implemented locally. The routes fail closed as `noindex,nofollow` and

remain absent from the sitemap. The next milestone is independent editorial and

research review of the exact release; only then may a human approve indexing.

Production impressions, qualified researcher participation, external reuse,

settled income, and the remaining A+ requirements remain unobserved.

Consented organic campaign — local controls complete, production proof pending

The exact double-opt-in audience, confirmation purpose, five-message copy,

privacy boundary, and aggregate measurement contract now produce a digest-bound

readiness artifact. Durable failed-delivery receipts stop automatic retry, and

receipt-backed permanent bounces or complaints can suppress by hash without raw

recipient data; temporary failures stay held for operator review. The gate still lacks six production delivery receipts and

attributable human approval, so lead capture and sending remain unauthorized.

Independent validation projection — local controls complete, external study pending

The Rosenberg comparison can now project into the A+ investigation, reviewer,

and frozen-baseline fields only after the underlying study passes and exact

external evidence proves three facilitator-verified researchers, one independent

subject expert, a distinct materially corrected v2 packet, a novelty decision,

published limitations, and unique receipt hashes. Direct identifiers,

self-reported-only qualifications, unchanged dossier digests, and synthetic

substitutes are rejected. The current readiness artifact has four blockers;

there are still no genuine returned sessions or expert decision to credit.

The artifact now has a canonical integrity digest and automatically supplies

the corrected investigation plus reviewer set to the A+ collector only when it

is ready and untampered. Measured improvement still comes exclusively from the

separate row-level conventional-baseline importer, preventing a weaker

projection from overriding that evidence lane.

Production AI evaluation — evaluator complete, production evidence pending

A separate A+ evaluator now refuses to treat the 120-prompt local golden run as

production proof. It requires at least 50 unique fresh production traces bound

to a pre-output, independently approved label manifest, with at least ten

answerable and ten refusal cases, two independent reviewers, 25 reviewed

citations, one model/prompt revision, provider or instrumented cost, and unique

trace/review receipts. Precision, recall, citation and refusal accuracy, p95

latency, and mean cost are calculated from rows. The current artifact is blocked

on the absent frozen manifest and production trace bundle.

| Cadence | Required action |

| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |

| Every behavior change | Red-green-refactor tests, focused smoke/evidence, README + roadmap update, and `pnpm session:closeout`. |

| Every 72 hours during active shipping | Canonical local gate plus fresh P0 evidence checks. Stop expansion when a required gate is missing. |

| Weekly | Refresh long-window evidence, inspect failed-sample age-out dates, review pilot pipeline, and update only changed roadmap decisions. |

| Monthly or before a buyer review | Refresh production preflight, launch evidence/review, procurement packet, access/DR evidence, and the strict handoff. |

The roadmap records current decisions and measurable outcomes, not every merged

change. Completed implementation detail belongs in

progress/era-history.md(progress/era-history.md), specialized docs, generated

artifacts, and git history.

Research release quality checkpoint (2026-08-14)

The A+ program now has a single-digest technical release gate covering tests,

lint, build, dependency security, docs validation, and diff hygiene. The point

remains unearned until a fresh successful artifact exists against the current

canonical Research Commons manifest; this does not relax any external-evidence

or production-income requirement.

The release receipt is now semantically replayable rather than outer-hash-only. Its

verifier recomputes the digest from the exact canonical file set, requires the fixed

six commands, validates command chronology and exact durations, and reconstructs

blockers/status. Rehashed blocked-to-passed, command-substitution, and manifest-

omission edits fail before scoring.

The initial gate exposed high-severity advisory 1139346 in Lighthouse's

`extract-zip` dependency. The redundant three-route Lighthouse check was

replaced by the retained 23-route axe/Playwright WCAG A/AA gate, and patched

transitive versions of PostCSS, Sharp, OpenTelemetry Core, and UUID are pinned.

The dependency audit now reports zero advisories without a baseline waiver. The

canonical release manifest includes the package manifest, lockfile, CI,

accessibility command, security command, and security baseline.

The repository and CI now share pinned pnpm 10.34.5 resolution semantics, so

workspace security overrides cannot be silently ignored by an older client.

Organic researcher contribution checkpoint (2026-08-14)

An indexable `/research/contribute` hub now gives researchers and experts a

single canonical path into packet reproduction, material correction,

counterbalanced workflow comparison, and method challenge. It requires no

account, uploads no response, exposes `ResearchProject` structured data, and is

linked from the sitemap, footer, and Research Commons. Controlled study tools

remain `noindex`, while the contribution guide makes privacy, custody,

qualification, correction, and human-decision boundaries explicit. This

improves discoverability but earns no external-validation point until genuine

receipt-bound evidence is returned and independently qualified.

Machine-readable Research Commons checkpoint (2026-08-14)

The indexable Commons now advertises a public, digest-protected JSON-LD catalog

and directly fetchable packet JSON Schema whose declared `$id` resolves. The catalog enumerates only public

methodology, contribution, correction, review, and schema resources plus

no-account actions; it deliberately excludes draft questions, unapproved

findings, private packets, and participant data. Both resources have explicit

media types, public caching, CORS, sitemap entries, and visible links.

Technical recommendation packet: the change adds two static research routes

and one catalog service without changing evidence or publication decisions.

Ranked risks are (1) leaking draft research through discovery, controlled by an

allowlist catalog and exclusion tests; (2) catalog drift, controlled by its

canonical digest and release manifest; and (3) schema divergence, controlled by

serving the exact release-bound packet schema. Next actions are to validate

external consumption (research operations; acceptance: attributable fetch or

reuse receipt), monitor crawler discovery (acquisition operator; acceptance:

fresh provider export), and keep catalog entries tied to publication state

(research lead; acceptance: no unapproved resource in the catalog). Validate

with `node --import tsx --test tests/services/research-commons-catalog.test.ts

tests/pages/research-commons-page.test.ts`, `pnpm research:commons:browser-proof`,

and `pnpm research:release-quality`. The next-cycle hypothesis is that an

external researcher can discover the packet contract without traversing a

private API; falsify with a clean-session fetch of both advertised alternates.

The catalog now also exposes a deterministic, schema-complete synthetic sample

packet. It uses a fictional object and non-executing provider, carries explicit

permanent exclusions from every qualifying evidence lane, and detects any

post-generation mutation through the same canonical digest contract used by

device-local packets. This closes the reproducible-sample requirement without

manufacturing a research result.

Technical recommendation packet: ownership remains within the packet and public

catalog services. Ranked risks are (1) a sample being mistaken for evidence,

controlled by permanent machine-readable exclusions and fictional content; (2)

silent mutation, controlled by canonical SHA-256 verification; and (3) sample

drift from the public schema, controlled by required-field and release-manifest

tests. Next actions are external clean-session reproduction (research

operations; acceptance: independently recomputed matching digest), schema-client

import (developer advocate; acceptance: generated client accepts the packet),

and catalog discovery measurement (acquisition operator; acceptance: fresh

attributable provider receipt). Validate with `node --import tsx --test

tests/services/research-commons-sample.test.ts

tests/services/research-commons-catalog.test.ts`,

`pnpm research:commons:browser-proof`, and `pnpm research:release-quality`. The

next-cycle hypothesis is that a third party can reproduce the digest without

project code; falsify by following only the public sample instructions.

Research correction propagation checkpoint (2026-08-14)

Corrections now use a public version-1 envelope schema, two independent digests,

raw-file receipts, allowlisted field targets, pseudonymous human submission,

HTTPS evidence, and explicit consent. `research:correction-propagate` preserves

the original packet and creates a distinct candidate revision with the proposed

change and an append-only pending decision. It rejects synthetic samples,

direct identifiers, digest drift, unsupported targets, and conclusive claims;

it cannot accept, publish, qualify, contact, or convert a proposal into impact.

Technical recommendation packet: ownership sits in the correction-propagation

service and file-receipt command. Ranked risks are (1) silent historical

overwrite, controlled by retaining the exact original and a previous-digest

link; (2) automated acceptance, controlled by a fixed pending human decision;

and (3) correction laundering into impact evidence, controlled by synthetic,

qualification, release, and acceptance separation. Next actions are independent

correction review (subject expert; acceptance: reproduced evidence and explicit

decision), released-revision binding (research lead; acceptance: distinct

approved digest), and impact import only after external acceptance (evidence

operator; acceptance: attributable receipt). Validate with `node --import tsx

--test tests/services/research-correction-propagation.test.ts`, `pnpm

research:correction-propagate`, and `pnpm research:release-quality`. The

next-cycle hypothesis is that any one-field correction can be propagated without

mutating its source; falsify by comparing the retained source bytes and digest.

The indexable contribution hub now contains a device-local correction-envelope

builder that shares the operator validator's browser-safe core. It requests no

name, email, account, upload, or artwork file; validates one allowlisted field,

HTTPS evidence, pseudonymous contributor code, human submission, consent, and

AI disclosure; computes the canonical digest with Web Crypto; and downloads the

file under researcher custody. Desktop and mobile Chromium prove the complete

download path, accessibility, and no horizontal overflow. That proof also

exposed and fixed a pre-existing nested-main landmark defect on the contribution

page.

Technical recommendation packet: ownership is split cleanly between the pure

envelope core, client builder, and Node-only propagation finalizer. Ranked risks

are (1) browser/server validation drift, controlled by the shared pure core; (2)

unnecessary personal-data capture, controlled by the field allowlist and direct

identifier rejection; and (3) proposal/acceptance confusion, controlled by

pending-only wording and disabled publication. Next actions are genuine expert

use (research operations; acceptance: independently returned envelope),

accountable review (subject expert; acceptance: explicit accept/reject receipt),

and candidate propagation (operator; acceptance: distinct retained digest).

Validate with `node --import tsx --test

tests/services/research-correction-envelope.test.ts

tests/services/research-correction-propagation.test.ts`, `pnpm

research:commons:browser-proof`, and `pnpm research:release-quality`. The

next-cycle hypothesis is that an expert can return a valid correction without

sharing contact or artwork data; falsify through a consented usability session.

Release diagnostic observability checkpoint (2026-08-14)

The canonical release gate no longer reduces an intermittent test, lint, or

build failure to an opaque exit code. Failed receipts retain at most five

prefix-filtered, redacted, individually hashed diagnostic lines plus an exact

gate-specific next action. Successful output remains digest-only. The gate does

not retry automatically, and its test instruction states that a passing retry

does not prove a flaky failure resolved.

Technical recommendation packet: ownership remains in the release-quality

service and runner. Ranked risks are (1) leaking command output, controlled by

prefix selection, redaction, length/count limits, and pass-output exclusion;

(2) hiding intermittent failures through retries, controlled by no automatic

retry and explicit operator wording; and (3) tampered diagnostics, controlled by

per-line and artifact digests. Next actions are to use the serial spec reporter

on recurrence (quality owner; acceptance: named failing test), repair or

formally quarantine its root cause (module owner; acceptance: repeated clean

canonical runs), and retain failure/pass history (release operator; acceptance:

both receipts remain attributable). Validate with `node --import tsx --test

tests/services/research-release-quality.test.ts`, `pnpm lint`, and `pnpm

research:release-quality`. The next-cycle hypothesis is that the next command

failure yields a safe named diagnostic; falsify with the unit-test error fixture.

Conventional-baseline integrity checkpoint (2026-08-14)

Measured improvement can no longer enter A+ scoring as unlinked aggregate

numbers. A dedicated importer now freezes cases, workflows, protocol and

thresholds before observation; derives results from unique paired session rows;

requires both counterbalanced orders and receipt hashes; and cross-links each

participant to the qualified-researcher evidence set. No genuine bundle is

present, so the external requirement remains unearned.

Production AI evidence assembly checkpoint (2026-08-14)

The production evaluator now has a privacy-safe staging path. Frozen labels,

production observations, and independent citation reviews remain separate until

an exact receipt-bound join succeeds. Retained rows exclude prompts and output

prose, include observed latency and cost receipts, and prevent the run operator

from self-review. No production observations exist yet, so the metric remains

unearned and no model execution is authorized.

The operator path now includes `research:ai-production-run-kit`, which converts

only a valid independently approved 50-case manifest into case-complete pending

observation and balanced two-lane review worksheets. It retains no prompts or

outputs and invents no human reviewer data. The run kit, assembler, and final

evaluator now share a canonical parsed-manifest digest, eliminating an

interoperability defect where harmless JSON formatting could prevent a genuine

bundle from aligning. The default configuration remains blocked because no real

approved manifest exists.

External-impact integrity checkpoint (2026-08-14)

External impact can no longer pass from aggregate counts. A dedicated importer

requires unique attributable public events, independent substantive

verification receipts, and an exact match to the reviewed investigation ID,

version, and packet digest. Meta Museum-owned and synthetic hosts are excluded.

No genuine external citation, reuse, or accepted correction is retained, so the

requirement remains unearned.

Research acquisition integrity checkpoint (2026-08-14)

A+ acquisition now requires one receipt-bound join across provider SEO, GA4

funnel, exact approved campaign controls, and pseudonymous human-verified

qualified actions tied to the current Research Commons release. Clicks, opens,

downloads, signups, and AI activity remain funnel signals only. No genuine

joined provider and qualified-action evidence is retained, so the requirement

remains unearned.

The contribution hub now closes the aggregate measurement gap with four

consent-dependent events for reproduction, correction, workflow comparison,

and method challenge. Normalized GA4 import and the operator cockpit preserve

each count and the aggregate contribution-path total. A+ acquisition now

requires at least one such fresh path signal alongside the separate

human-verified qualified-action ledger, so click activity remains explicitly

ineligible as scholarly validation.

Unified research evidence operations checkpoint (2026-08-14)

`research:evidence-operations` now reads the A+ scorecard and all six retained

external-evidence lanes, verifies available artifact integrity, calculates

requirement-specific freshness, and emits deterministic 30/60/90-day windows.

Seven missing criteria are grouped into six exact command workstreams because

external investigation review and reviewer qualification share one validation

bundle. Missing, blocked, invalid, stale, and ready-but-unaccepted states remain

distinct, and every action explicitly denies external side-effect authority.

The operations artifact now retains the readiness requirements, lane snapshots, and

stage-specific command overrides and semantically reconstructs freshness, reporting

windows, grouped workstreams, priorities, invocation contracts, alerts, and authority.

Rehashed command, priority, or authority substitutions fail before cockpit display.

The four older fallback writers now use the shared canonical integrity

finalizer, eliminating false `invalid` alerts while keeping every absent-input

lane blocked and ineligible.

The production-AI workstream now resolves its command from verified stage

evidence. It now begins by converting a receipt-bound 50-question export into

an integrity-protected, text-free labeling worksheet with blank labels and no

approval. A valid candidate advances to the separately human-approved run-kit

stage; a valid ready run kit advances to evidence assembly; and a valid ready

assembly advances to production evaluation. Missing, blocked, or digest-invalid

stages return the operator to the earliest safe command.

Technical recommendation packet: the behavior change is confined to evidence

operations and preserves every external-action boundary. Ranked risks are (1)

retaining sensitive question text, controlled by hash-only output; (2)

advancing from a modified intermediate artifact, controlled by canonical

verification; and (3) mistaking preparation for approval or qualifying

evidence, controlled by null labels and the unchanged A+ evaluator. Next actions

are to prepare a genuine 50-question candidate export (research operator;

acceptance: verified text-free worksheet), obtain independent label approval

(research lead; acceptance: receipt-bound frozen manifest), assemble

receipt-bound observations and reviews (evaluation operator; acceptance: ready

verified assembly), and run the unchanged production scorer

(research lead; acceptance: row-derived passing metrics). Validate with

`node --import tsx --test tests/services/research-evidence-operations.test.ts`,

`pnpm research:evidence-operations`, and `pnpm research:release-quality`. The

next-cycle hypothesis is that a completed verified run kit changes the first

command to evidence assembly; falsify by inserting a digest mismatch and

confirming the queue returns to run-kit preparation.

Evidence command-contract drift checkpoint (2026-08-14)

Evidence operations now have an executable maintenance contract across all nine

stage and external-evidence commands. The focused test resolves each documented

command through `package.json`, verifies its exact implementation target,

confirms that source parses every machine-readable input-contract flag, and

requires the implementation in the canonical research-release manifest. This

turns documentation/CLI drift into a release failure instead of an operator

surprise. The focused suite passes 6/6; this adds no external action authority

and does not change the evidence-backed A+ score.

Private evidence-queue cockpit checkpoint (2026-08-14)

The editor-only organic-income cockpit now consumes the signed research-evidence

operations artifact instead of leaving the prioritized queue accessible only as

generated JSON/Markdown. Authorized operators see each safe invocation plus

its structured private/non-private and human-attested/machine-derived inputs.

Missing, digest-modified, and malformed artifacts expose zero commands and a

clear fail-closed explanation. The page, loader, and their tests are now bound

into the canonical research release. Focused page, loader, and release-manifest

tests pass 5/5; no model, recruitment, publication, email, or payment authority

is introduced.

Production-question export contract checkpoint (2026-08-14)

The first production-evaluation handoff no longer trusts a TypeScript cast over

operator JSON. A release-bound Draft 2020-12 schema and matching runtime

validator enforce the exact root and case fields, reject additional or direct-

identifier fields before creating an output directory, and use row-number-only

diagnostics so malformed identifiers and private values are never reflected.

The CLI still hashes valid questions into blank, unapproved labeling rows and

performs no model call. Service, subprocess, schema, non-reflection, and release

mutation coverage passes 9/9 focused tests.

Canonical production-question schema discovery checkpoint (2026-08-14)

The production-question schema no longer claims an unserved, noncanonical URL.

Its `$id` now resolves through `/schemas/production-question-export/v1` on the

canonical `www` origin with `application/schema+json`, public CORS, and bounded

cache controls. The Research Commons page, signed JSON-LD catalog, and sitemap

all discover the same route, and the route/schema are release-bound. Ten focused

route, catalog, page, validator, and manifest tests pass. Only the data-free

contract is public; production questions remain private.

Canonical-host equality is asserted from the served schema response rather than

inferred from filenames.

Attested production-query transformer checkpoint (2026-08-14)

The first genuine-data handoff can now be produced from an explicitly exported

Meta Museum `ai-query-log.json` without hand-editing or implicit live-storage

access. `research:ai-question-export` requires a separate human production

attestation bound to the exact raw digest, accepts only fresh successful rows,

deduplicates normalized questions, rejects contact data, pseudonymizes raw query

IDs, and writes the private strict-schema export only at 50 eligible cases.

Aggregate-only stdout and no-output-on-blocked behavior keep questions out of

operator logs. Eleven focused service, subprocess, schema, and release tests

pass; the transformer cannot establish production genuineness or approve labels.

Production-question aggregate readiness checkpoint (2026-08-14)

`research:ai-question-attestation` now prepares a digest-bound, text-free

candidate before any human production claim. It shares the transformer's exact

freshness, success, length, contact-data, and deduplication analysis; retains

only aggregate exclusion counts plus one privacy-safe outcome code per row; and leaves environment, source, operator, and

collection approval null. The consent-hardened current local log yields zero

eligible rows from the legacy input because none satisfies the exact current export

shape and consent boundary. The candidate remains

blocked and makes no production claim. The current replay found 2,497 legacy rows,

all malformed for the exact current evaluation-export contract and therefore zero

eligible. Evidence operations verify its digest but correctly keep consented

collection approval ahead of this downstream prerequisite. Candidate, subprocess,

resolver, command-contract, and manifest-chain tests pass.

Verification reconstructs counts, blockers, readiness, and the blank attestation

without retaining questions, raw IDs, timestamps, or contact values. Rehashed count,

blocker, or blocked-to-ready substitutions fail.

The AI-query request contract now has an optional strict evaluation envelope:

`consent: true`, consent version `research-ai-evaluation-v1`, explicit human

submission, and an allowlisted collection lane. Partial, false, or extended

claims fail request validation. Ordinary future queries retain only a question

SHA-256 and null text; exact text is retained only for the consented lane. The

aggregate readiness analyzer excludes every ordinary or legacy row as

`unconsented`, while the separate exact-digest human production attestation

remains mandatory. API, logger, analyzer, transformer, and operations coverage

passes 25/25 focused tests, and these trusted inputs are release-bound.

Approval-gated AI question surface checkpoint (2026-08-14)

`/research/evaluate-ai` now provides the technically complete path for future

diverse consented questions without requesting identity, but remains no-index

and inactive. A stable contract digest covers the exact consent copy, purpose,

data fields, exclusions, and 35-day retention. Activation requires both an

attributable checked-in approval of every dimension and a separate runtime

switch; either alone fails closed. The inactive page renders no form, and the

API independently returns 403 for crafted collection requests. Contact-bearing

questions fail before execution, ordinary logs are hash-only, and the writer

nulls expired consented text while preserving receipts. Twenty-six focused

page, gate, API, logger, analyzer, transformer, and release tests pass. No

approval, deployment, publication, recruitment, or collection occurred.

Research browser-proof scoring checkpoint (2026-08-14)

A+ readiness now verifies the three required Research Commons journeys by

exact title across both desktop and mobile Chromium, while also requiring the

report totals to reconcile and unexpected/flaky counts to remain zero. This

replaces the obsolete assumption that a healthy report always contains exactly

four executions. The score therefore stays strict when required coverage is

lost and remains stable when additive browser coverage is introduced. Focused

regression tests prove both cases; the current artifact-backed readiness score is 3/11,

with all eight remaining requirements dependent on attributable external

production, validation, acquisition, impact, or income evidence.

A+ scorer trust-boundary checkpoint (2026-08-14)

The canonical research release now binds the A+ evaluator implementation,

operator CLI, blank evidence configuration, service and CLI tests, plus a

manifest regression test. The regression test mutates each trusted scoring

input independently and proves the release digest changes. This prevents an

easier threshold, altered requirement, or changed evidence loader from being

applied under an older release identity. As designed, expanding this boundary

temporarily invalidates the prior release-quality receipt until the canonical

gate regenerates evidence for the new digest.

Production-question intake usability checkpoint (2026-08-14)

The first external-evidence handoff now accepts a genuine production-question

export directly through `--input`, computes its exact raw-file receipt and a

truthful preparation timestamp, and emits only the integrity artifact plus a

question-hash-only blank labeling worksheet. Controlled runs may override the

timestamp and output directory, while configuration mode remains available for

scheduled operation. An end-to-end subprocess test proves that 50 cases become

a ready worksheet and that neither stdout nor either generated artifact

contains the retained question text. No model call, label inference, approval,

contact, or publication authority is added.

The private status now retains a hash/length-only replay projection with valid

pseudonymous case IDs and redacts invalid IDs. Its verifier reconstructs source

eligibility, uniqueness, normalized-length gates, blockers, candidate count, status,

and every blank labeling row; rehashed readiness or worksheet edits fail.

Approved-manifest run-kit usability checkpoint (2026-08-14)

The second production-evaluation handoff now accepts an independently approved

manifest and pseudonymous run-operator code directly through CLI arguments. It

derives the canonical dataset identity and truthful preparation timestamp,

rejects the label approver as run operator, and writes only blank case-complete

observation and balanced independent-review worksheets. A subprocess test

proves the separately exported worksheets omit question hashes, approval references,

and the approver code. The private status artifact retains the exact approved

manifest and reconstructs the dataset digest, approval chronology, operator

separation, blockers, status, case-complete observations, and balanced review lanes;

rehashed blocked-to-ready or row-removal edits fail. The command still performs no

model call and running it without genuine approved inputs remains blocked.

Production evidence assembly usability checkpoint (2026-08-14)

The third production-evaluation handoff now accepts the approved manifest,

completed production observations, and separate independent reviews directly.

It derives the canonical manifest identity and both raw file receipts, removing

six copy-prone configuration values. The operator must still provide the

attributable HTTPS production evidence URL, its retained receipt, and a

pseudonymous run code; the assembler cannot infer those facts. End-to-end proof

assembles 50 aligned traces and reviews, while absent real inputs still produce

only an integrity-protected blocked status and no evaluation bundle.

The assembly is now self-contained and semantically replayable: its private artifact

retains only the sanitized manifest, observations, reviews, release/operator fields,

and external receipt inputs, then reconstructs all row joins, chronology, reviewer

separation, citations, costs, blockers, counts, status, and bundle. Rehashed blocked-

to-ready, row-removal, or bundle-substitution edits fail before scoring.

Production AI scoring usability checkpoint (2026-08-14)

The final production scorer now accepts the approved manifest and ready bundle

directly, derives both receipts, revalidates frozen-label alignment, and emits

the same integrity-protected evaluation artifact without configuration edits.

An end-to-end subprocess test proves a 50-case, two-reviewer bundle yields only

row-derived precision, recall, citation, refusal, latency, and cost metrics.

The scorer still performs no model call, ignores aggregate assertions, and

fails closed when either genuine input is absent or any threshold is missed.

Together, all four production-evaluation handoffs now have direct CLI paths.

External validation bundle-binding checkpoint (2026-08-14)

The highest-leverage external lane now accepts a returned validation bundle

and private evidence envelope directly, keeping qualifications, attribution,

and the versioned correction decision out of tracked configuration. The signed

readiness artifact retains both the frozen protocol digest and the exact raw

validation-bundle digest before it can project the corrected investigation and

four qualified reviewers. End-to-end proof covers three researchers, one

subject expert, five independently coded readers, a distinct corrected v2,

explicit novelty/limitations decisions, and measured comparison outcomes.

Absent genuine evidence still produces four blockers and imports nothing.

Conventional baseline direct-reconciliation checkpoint (2026-08-14)

The separate measured-improvement lane now accepts a frozen manifest and

paired-session bundle directly and derives their canonical/raw receipts. It

still requires the attributable study source, measurement time, both

counterbalanced orders, unique consented independent researchers, three cited

sources per condition, facilitator receipts, 20% median improvement, and 80%

reproduction. End-to-end proof yields 50% median improvement and 100%

reproduction from rows and preserves the three participant codes that the A+

scorer must reconcile against the validation lane's qualified reviewers. No

aggregate metric or validation projection can bypass this independent check.

External impact cross-lane binding checkpoint (2026-08-14)

External citation/reuse/correction evidence now imports against the signed

validation readiness artifact rather than three retyped investigation fields.

The importer derives the exact corrected investigation ID, version, and packet

digest plus the raw event-file receipt; A+ scoring independently repeats the

release match. Public-URL validation now rejects localhost, private and

link-local IPv4, and private/loopback IPv6 in addition to internal and

synthetic hosts. End-to-end proof binds a verified citation to the corrected v2

packet, while the default lane remains blocked with no event to credit.

Acquisition release-binding checkpoint (2026-08-14)

Attributable acquisition now requires the integrity-verified organic provider

report itself—not only the qualified action ledger—to match the current

research release. Direct organic-report, campaign, and qualified-action inputs

derive raw receipts that remain inside the signed artifact. Shared public-URL

validation excludes localhost, private/link-local IPv4, and private/loopback

IPv6 verification references. End-to-end proof joins provider SEO/funnel data,

the exact approval-gated campaign controls, and one consented human-qualified

organic action; clicks, page opens, bots, and activity alone remain ineligible.

Net-income receipt and freshness checkpoint (2026-08-14)

The economics lane now accepts settlement and observed-cost exports directly,

derives both raw-file receipts, and retains them inside the signed artifact.

Live rows cannot cite Stripe test-mode dashboard paths; cost evidence must be

fresh, public-network, observed, order-attributable, and receipt-bound. The

row-level proof computes $29.00 gross less $1.14 fees and $2.00 observed cost

as $25.86 net while excluding a $99.99 sandbox settlement. The checked-in lane

still has no live settlement and therefore remains blocked; no checkout or

payment capability was activated.

Evidence-operations invocation checkpoint (2026-08-14)

The six-workstream operations artifact now exposes both a stable command ID and

the complete non-executing direct-input invocation for each current stage. The

production-AI template changes as verified artifacts advance; validation,

baseline, impact, acquisition, and economics templates enumerate their exact

bundle, evidence, provider, action, settlement, and cost inputs. The Markdown

handoff includes the same templates, while authority flags remain false and no

secret, external file, model run, contact, publication, or payment action is

performed.

Agent-safe evidence input contracts checkpoint (2026-08-14)

Every stage-aware invocation now carries structured contracts for all required

flags: evidence kind, description, private/non-private handling, and whether a

human attestation is mandatory or the value may be machine-derived. All nine

possible stage commands have non-empty, duplicate-free contracts. This lets a

bounded agent assemble local tasks and route sensitive files correctly without

inventing reviewers, receipts, production observations, approvals, or external

authority; the signed operations artifact retains the complete contracts.

Research net-income integrity checkpoint (2026-08-14)

A+ economics now derives positive net income from current-release live Stripe

settlements and separate observed order-cost receipts. Currency, offer revision,

refunds, fees, disputes, attribution, and sandbox exclusion remain explicit;

gross receipts and forecasts cannot substitute. No genuine live settlement is

retained, so the income requirement remains unearned.

AI question-collection prerequisite checkpoint (2026-08-14)

Production-evaluation operations now detect that the consented question lane

is inactive before recommending log attestation. The new

`research:ai-question-collection-readiness` command emits the exact stable

collection contract, its digest, a blank approval record, explicit operator

sequence, runtime-switch state, blockers, and a tamper-evident packet receipt.

It distinguishes awaiting human approval, ready for separate runtime

activation, and active states without approving research, deploying, enabling

collection, or contacting anyone. This removes a dead-end loop while retaining

the human approval and deployment boundaries.

Question-led SEO approval-contract checkpoint (2026-08-14)

The six substantive research-question guides no longer depend on an impossible

self-referential full-release approval digest. A stable content contract now

binds every title, question, section, source, FAQ, canonical publication effect,

and excluded claim. The publication-readiness command emits a blank human

review record and tamper-evident packet; indexing requires attributable HTTPS

review evidence and explicit claims/source, usefulness, rights/privacy, and

desktop/mobile accessibility decisions. Evidence operations selects this gate

before provider acquisition imports while the pages remain noindex and absent

from the sitemap. Approved pages also stop displaying the contradictory

"publication is not authorized" notice.

The packet now semantically replays the exact contract, explicit signed decision,

preparation chronology, derived blockers and status, operator sequence, and false

authority boundary. Its default blank is generated from the current contract;

checked-in configuration supplies no runtime approval. Non-public evidence URLs,

future approvals, and rehashed awaiting-to-approved substitutions fail closed.

Open-source research discovery checkpoint (2026-08-14)

Researchers can now discover one synchronized software identity through a

canonical, indexable `/research/software` landing page, repository-native

`CITATION.cff`, CodeMeta, and public JSON-LD at `/api/research/software`. The

human page gives direct reproduction, correction, expert-review, repository,

methodology, schema, dataset, and citation paths while sharing the machine

record's explicit claim boundary. The Research Commons and footer link it.

GitHub issue routing disables unstructured blank reports and directs people to

bounded research or private-support paths. Release tests reject repository or

version drift and verify public cache, CORS, OPTIONS, and artifact integrity.

The metadata intentionally asserts no DOI, deposit, named authorship, external

reuse, endorsement, validation, demand, impact, revenue, or income.

Approved-research SEO candidate checkpoint (2026-08-14)

The publication layer now has a deterministic bridge from an exact approved

report to a reviewable SEO landing-page candidate. It requires the reviewed-

findings channel, publication-approved status and flag, an attributable dated

researcher decision, two HTTPS-cited claims, limitations, and the current

research release. Output binds the full report digest and carries canonical

metadata, Article JSON-LD, citations, missing evidence, and internal links while

publication, deployment, indexing, and outreach authority remain false. The

current Rosenberg report correctly produces only a blocked artifact; a verified

approved-report envelope can be supplied without changing application code.

The candidate now retains the exact report and full publication envelope and

semantically rebuilds all copy, citations, JSON-LD, links, and robots state during

verification; rehashed landing-page edits cannot pass.

Review-to-publication envelope checkpoint (2026-08-14)

The report layer now binds the missing transition from genuine external review

to an SEO-eligible approved report. A ready validation artifact must prove the

materially corrected, limitations-bound non-synthetic revision. Separate human

publication approval must bind both exact report and validation digests and

affirm citations, limitations, rights, and novelty framing. The resulting

approved-report envelope changes report status/channel only inside a signed

handoff while retaining false publication and deployment authority. SEO input

now rejects approved-looking bare JSON and accepts only this verified envelope.

The current run remains blocked with no approved report emitted.

The envelope is now self-contained: it retains the unapproved source report and

complete validation artifact and replays correction/version requirements, approval

chronology, reviewer separation, editorial decisions, and the approved projection.

The preparation packet is self-contained as well and canonically reconstructs blocked

or ready state from the retained report, validation, explicit approval, and generation

time. Approval evidence must be public attributable HTTPS and cannot postdate packet

generation; rehashed status, blocker, or projection substitutions fail closed.

Dependency-aware approval-register checkpoint (2026-08-15)

Six previously fragmented approval and evidence lanes now feed one canonical,

tamper-evident register. It orders the core research path before optional email

and commercial activation, models prerequisites explicitly, and supplies the

private income cockpit with the first safe preparation command only after

digest verification. Every authority flag remains false. The current first

action is consented AI question-collection approval; external validation and

the post-review publication decision remain genuine external requirements.

Campaign and commercial readiness artifacts now sign their full persisted

records. The register therefore distinguishes integrity-valid blocked work from

tampered or unsigned artifacts without weakening any activation gate.

The register itself now retains those six source artifacts and reruns every

lane-specific verifier before reconstructing states, dependencies, and first action;

caller-supplied integrity booleans are ignored. Ready reviewed publication is sourced

from the full replayable approval envelope, while the preparation packet is used only

for blocked state. The private artifact can therefore reject rehashed routing edits

and forged ready summaries rather than trusting a one-time script check.

AI question-collection readiness now reconstructs its retained contract, explicit

signed approval, preparation chronology, runtime switch, blockers, and status during

verification. Its inactive template is generated from the current contract rather

than trusting stale null configuration fields. Public attributable approval URLs and

approval-at-or-before-preparation are mandatory; rehashed activation-state edits fail.

Organic expert-return checkpoint (2026-08-15)

The indexable contribution journey now closes the previously disconnected

handoff from device-local review to the repository's structured independent-

review form. The form captures a pseudonymous reviewer code, exact packet

digest, at least three checked public HTTPS sources, corrections, alternative

hypotheses, novelty, limitations, and AI-assistance disclosure while excluding

identity evidence and private files. The frozen Rosenberg workspace links the

same route but remains noindex. Public contribution is explicitly not accepted

as facilitator-verified qualification, A+ validation, publication approval, or

compensation.

It now also removes GitHub and prior facilitator contact as prerequisites for

creating a review. The no-account browser form collects only a packet digest,

pseudonymous code, reproduction, three public-source findings, corrections,

alternative hypotheses, provisional novelty, limitations, AI disclosure, and

four explicit declarations; it rejects direct identifiers and placeholder

hosts and downloads an integrity-signed envelope with every consequential

authority false. The page, sitemap, JSON-LD catalog, and public JSON Schema make

the pathway discoverable to humans and agents. Any A+ use still requires

separate facilitator verification of identity mapping, qualification,

conflicts, independence, consent, and custody.

The handoff now has its own consent-gated aggregate event and cockpit metric,

kept separate from contribution-path totals to prevent double counting. No

reviewer code, packet digest, account identifier, or external-form outcome is

collected, and an open never counts as a returned or qualified review.

The private cockpit now derives a fail-closed expert-review funnel state from

the verified aggregate report and signed approval register. It prefers genuine

returned validation evidence over click activity and limits operator actions to

evidence refresh, handoff inspection, approved-channel inspection, or distinct

human review—never visitor identification or contact.

The reproducible browser proof builds the current worktree before starting its

production server, then covers seven Research Commons journeys across desktop

Chromium and Pixel 5. The expert-return coverage exercises the no-account,

device-local review envelope and verifies its exact packet binding, authority

denials, optional public return route, privacy boundaries, accessibility, and

horizontal overflow. All 14/14 executions pass with no unexpected or flaky runs.

The release-handoff regression asserts all sixteen open targets, including the

expert-review schema, and fails if implementation and capture coverage diverge.

The no-account expert-review journey now continues through a non-overwriting,

offline envelope importer. It replay-verifies the returned device-local artifact

and routes it to the existing facilitator decision while requiring an exact

envelope-receipt match and retaining zero qualification or validation authority.

Facilitator decisions now also carry the exact candidate artifact SHA-256; the

verifier rejects missing bindings and reuse against a different returned review.

Operators can now generate a replay-verified, non-overwriting decision worksheet

from any supported return candidate. It copies the exact digest but leaves all

human evidence and checks blank, reducing transcription risk without automating

qualification or approval.

An editor-protected, no-index facilitator workspace now verifies the blank

worksheet entirely on-device, displays its exact candidate binding, requires

all applicable human checks, and downloads the completed decision without an

upload. The server verifier also rejects contact details in the retained

qualification basis, while the CLI remains the authoritative replay boundary.

The browser now also reconstructs and semantically checks GitHub and device-local

candidate projections, source receipts, nested signatures, and authority denials;

a rehashed candidate that grants itself qualification fails before form unlock.

Post-publication acquisition checkpoint (2026-08-15)

The attributable-acquisition join now requires a fourth, digest-bound evidence

lane: accountable human approval followed by evidenced public publication of

the exact Research Commons release. Qualified researcher or expert actions must

occur after that publication and inside the provider window, while every SEO

and analytics export must be captured only after the window closes. This

prevents internal tests, launch preparation, pre-publication visits, premature

exports, clicks, and opens from being credited as researcher acquisition. The

command remains blocked by default and grants no publication, campaign,

contact, qualification, or income authority.

The action-level attribution contract now closes the remaining self-assertion

gap. Every qualified row needs a unique consented first-touch receipt from a

Google or Bing organic visit to an approved research path, explicit internal-

traffic exclusion, independent-external participant status, and a human

attribution verifier distinct from the eligibility verifier. Missing evidence

fails closed; aggregate traffic and an `organic-search` string cannot establish

acquisition on their own.

The same ledger contract is now an open Draft 2020-12 JSON Schema served with

public caching and CORS, linked from the Research Commons, advertised in its

signed JSON-LD catalog, and discoverable through the sitemap. It encodes the

pseudonymous action, first-touch, independence, consent, and verification shape

while leaving cross-row uniqueness, distinct-verifier, chronology, freshness,

public-network, and digest checks to the authoritative importer. This reduces

real evidence-return friction without publishing any participant evidence.

Raw action JSON now crosses a strict runtime structural boundary before the

evaluator reads nested fields. Missing and additional fields, invalid enums or

constants, malformed pseudonymous codes, and incorrect hashes return a

canonical blocked artifact rather than throwing. Contract tests compare the

public schema's complete action, attribution, and qualification required-key

sets against a passing runtime fixture so documentation drift fails CI.

A+ projection-injection checkpoint (2026-08-15)

The authoritative A+ CLI now accepts only version and boundary metadata from

configuration. Attempted scoring projections are listed as rejected and never

evaluated; every scoreable external or technical projection must instead come

through its dedicated signed-artifact verifier. An adversarial CLI test injects

a perfect-looking AI evaluation and proves the requirement remains missing.

Hand-authored configuration can no longer substitute for production evaluation,

review, baseline, impact, acquisition, income, browser, or release-quality

artifacts.

The persisted A+ scorecard now retains its verified evidence projection and replays

all eleven requirements at the recorded evaluation time. Grade, summary, evidence

copy, next actions, rejected configuration fields, and authority are derived rather

than trusted. Evidence operations verifies this artifact and falls back to zero

passed requirements when it is missing or modified, so a rehashed grade cannot

misroute the private operator queue.

Release-bound production evaluation checkpoint (2026-08-15)

Production AI evidence now forms one release-consistent chain. The privacy-safe

assembler derives the canonical Research Commons release digest into its bundle;

the metric evaluator independently recomputes and checks that digest; and the

A+ scorer requires the projected AI digest to equal the open-research digest.

An older valid 50-case evaluation can no longer be combined with a newer code,

prompt, workflow, schema, or documentation release. Focused assembly, evaluator,

CLI, scorer, chronology, and mismatch tests pass without executing a model.

Release-bound external validation checkpoint (2026-08-15)

Independent validation now carries the same release identity. The readiness CLI

derives the canonical Research Commons digest into the signed corrected-v2

investigation projection, and the A+ scorer requires it to match the current

open-research release. Packet-level correction, reviewer qualification, and

artifact integrity remain mandatory, but they can no longer validate a changed

workflow merely because an older packet still verifies. Direct validation,

signed import, CLI, A+ mismatch, and publication-handoff tests retain coverage.

Preregistered release-bound baseline checkpoint (2026-08-15)

Measured workflow improvement now names the tested implementation before any

session occurs. The conventional-comparison manifest freezes the exact Research

Commons release digest with its protocol, cases, workflows, thresholds, and

preregistration receipt. The importer independently checks that digest against

the current release, retains it in the signed projection, and A+ requires it to

match the open-research release. An older speed/reproduction result cannot be

credited after the workflow changes, even when its study artifact still verifies.

---

History And References

Evidence correction: open does not mean locally implemented

The Research Commons open-release point now requires exact-release production

deployment approval and successful post-deployment response receipts for every

public resource. Local file existence cannot score. The importer is non-executing,

so publication remains a human operation and the point stays unearned until genuine

production evidence is supplied.

The deterministic evidence-operations report now covers all eight currently

missing evidence lanes: open-release production proof plus the seven external

evaluation, review, baseline, impact, acquisition, and income requirements. Its

first safe action is an importer invocation, never an automatic deployment.

Bounded agent work orders

The device-local Commons journey now converts its narrative agent plan into eight

standard AgentTask envelopes bound to the packet digest. Dependencies, candidate

ceilings, model-call ceilings, rights warnings, and human-trigger requirements are

machine-readable. Retrieval starts blocked and no task authorizes execution,

conclusions, publication, contact, or payment. The work-order schema is cataloged,

sitemapped, CORS-readable, and part of exact-release public probe evidence.

The client now verifies the signed artifact before download by recomputing the

canonical packet and work-order digests and reconstructing the expected DAG from the

packet. Rehashed dependency removal, inflated candidate/model budgets, source-plan

substitution, or cross-packet reuse fails even when generic AgentTask fields remain valid.

The public schema now closes task/input/budget/authority objects and fixes the hard

ceilings and false authorization flags rather than documenting them only in prose.

Open-release artifacts now retain their complete normalized deployment and probe

bundle. The verifier recomputes the evaluation and compares canonical outputs, so

an outer integrity hash no longer substitutes for semantic revalidation. The A+

next action names the exact evidence importer and required post-deployment receipts.

External citation/reuse/correction artifacts now retain their full normalized

event ledger and exact reviewed-investigation binding. Integrity verification

replays the evaluator at the retained evaluation time, preventing rehashed changes

to projected citation, reuse, correction, source, or event-ID results.

Economics artifacts now retain their row-level Stripe settlement and observed

cost ledgers. Verification replays live eligibility, freshness, sandbox exclusion,

refunds, fees, attributable costs, and net income, preventing rehashed profit edits.

Acquisition artifacts now retain the complete provider, campaign, publication,

first-touch, and qualification join. Semantic replay prevents rehashed traffic,

contribution-open, or qualified-review inflation from entering the A+ scorecard.

Measured-improvement artifacts now retain the frozen manifest and complete paired

session bundle. Semantic replay recomputes medians and reproduction while enforcing

session-before-measurement/evaluation chronology, preventing rehashed performance claims.

Production AI evaluation artifacts now retain the approved frozen label manifest

and complete trace bundle. Verification rechecks alignment and recomputes all

quality, latency, cost, coverage, and reviewer-independence metrics row by row.

External validation now enforces attestation chronology: qualifications and the

material-correction decision cannot postdate measurement, and correction cannot

predate independent review or returned-evidence assembly.

External-validation artifacts now retain the complete identifier-free returned

evidence and semantically replay the A+ projection at every downstream ingestion

boundary. Rehashed edits to reviewers, correction, investigation version, or study

results therefore fail before publication, impact, or A+ scoring can consume them.

The open-release handoff now starts with `pnpm research:open-release-handoff`, a

non-network preparation command that binds the current release, fixed sixteen-route

allowlist, blank human decision, exact sequence, and false authority flags in a

non-overwriting, semantically replayable packet. Evidence operations selects this

safe first action while no integrity-valid capture exists. Only after attributable

approval and production deployment may the bounded receipt-capture command probe the

fixed sixteen-route

allowlist covering the human Commons, schemas, JSON Feed, RSS, and JSON-LD software

metadata; refuses redirects, failures, unexpected content, time inversions, and

oversized bodies, and writes non-overwriting exact-body hashes for the independent

evidence importer. The operations queue exposes the two-stage capture/import path

without granting deployment or publication authority. Its stage resolver now replays

the configured capture through the authoritative evaluator: only an integrity-valid,

exact-current-release bundle advances from capture to import, preventing both repeated

network work and premature progression from malformed receipts.

The open-release queue now also closes the duplicate-handoff loop. It searches retained

non-overwriting handoffs, accepts only a canonically verified packet for the exact

current release prepared no more than 35 days ago and not in the future, and then points

to the approval-gated capture invocation. A stale, forged, malformed, or previous-release

packet returns to preparation. This improves operator usability without treating the

blank handoff as approval or granting deployment, network, publication, or indexing

authority.

The approval boundary no longer depends on manually setting `humanApproved: true`.

An editor-only, noindex, device-local workspace builds the exact current handoff,

requires an attributable APR-coded approve/reject decision with five explicit safety

checks and a public receipt digest, and downloads a signed artifact without persistence

or execution. Receipt capture now replays that decision against the current release and

requires its decision time to match the deployment configuration. A forged decision,

rejection, stale release, timestamp substitution, or standalone boolean cannot reach

the network capture stage.

The post-deployment handoff is now non-manual as well. A no-network, non-overwriting

assembler consumes the signed decision and production deployment receipt, replays exact

release and chronology, validates attributable public receipts and distinct APR/OPS

roles, and writes an ignored private capture config with false execution authority.

Capture re-verifies it before networking and automatically emits the completed importer

config beside the bounded probe receipts. Evidence operations can consume an explicitly

named private config to advance deterministically from config assembly to capture and

then import without editing timestamps or relative paths.

The private deployment workspace now offers the same join as a device-local browser

flow. It accepts the two JSON files without upload, checks both canonical signatures,

the current release and fixed target set, public evidence hosts, APR/OPS roles,

production/synthetic flags, and 35-day chronology, then downloads a configuration that

the capture CLI must replay before networking. Browser proof feeds its output to the

authoritative server verifier on desktop and mobile. This removes shell-only friction

without creating deployment, capture, publication, indexing, or A+ authority.

Open-release receipts are now self-contained rather than hash-only: every bounded

public response body is retained as canonical base64, exact hashes are recomputed at

ingestion, and media-specific JSON/JSON Feed/RSS/HTML/Markdown structure is reparsed.

This closes the forged-hash and status-200 error-page gap while retaining only public,

allowlisted content and enforcing the existing per-resource size bound.

Correction propagation now emits a signed, self-contained artifact rather than an

unsigned derived packet. It retains both exact input files within a 2 MB-per-file

bound, recomputes raw receipts, validates the original/envelope digests, reruns the

allowlisted field edit, and compares the entire revision canonically. Rehashed

publication authorization, revision, correction, or source-file substitutions fail;

the separately written candidate still carries no acceptance or publication authority.

The expert-validation flywheel now has a deterministic return preflight between the

public structured issue form and private facilitator review. The offline command binds

one return to its packet and exact repository issue, enforces public-source, privacy,

consent, pseudonym-role, and AI-disclosure rules, writes only to a new path, and signs

the complete retained input. Its output is explicitly non-qualifying until a separate

accountable facilitator verifies identity, expertise, independence, conflicts, and

consent; no review status can authorize contact, publication, payment, or an A+ claim.

Facilitator qualification now forms a second signed link rather than an editable

configuration row. The verification command retains the complete candidate and a

distinct post-capture facilitator decision, requires an attributable receipt and

eight explicit checks, and deterministically projects one pseudonymous attestation.

Validation readiness rejects inline qualifications and requires the three researcher

plus one subject-expert artifacts to replay and bind the exact dossier packet before

the external-review projection can advance.

Material correction and novelty are no longer editable validation projections. A

new non-overwriting decision artifact replays the signed correction propagation,

binds the same verified subject expert and original packet, requires a later human

acceptance with six explicit checks, and derives the distinct v2 packet identity.

Validation readiness now rejects inline correction decisions and consumes only this

replayable artifact, while publication, title, authenticity, legal status, contact,

and payment authority remain false.

The measured-improvement lane now closes the cross-study pseudonym-reuse gap. A

same-study join replays the frozen comparison and external validation, requiring

their release, protocol, packet, measurement, participant rows, timing, durations,

reproduction, and citations to match. A+ retains the join and validation sources,

so recomputing a scorecard hash cannot substitute copied codes or aggregate metrics.

The scorecard trust root now covers every external lane, not only validation. Its

private provenance retains the complete verified public release, production AI

evaluation, validation, same-study baseline join, impact, acquisition, and economics

artifacts. Verification replays every source and requires exact projection equality;

missing, unused, or semantically modified provenance invalidates the scorecard even

when an attacker recomputes its outer digest.

The same trust root now covers the three technical projections. Workflow replays the

full responsive browser report; automation retains exact package bytes tied to the

release manifest; quality retains and replays the six-command release artifact. This

removes the final summary-only path by which forged technical booleans could complete

an otherwise source-valid A+ scorecard.

The public Commons now routes each visitor by intent across seven low-friction paths:

read, reproduce, build, contribute, private assessment, self-serve kit, and bounded

human work. Its machine-readable catalog advertises the same discovery and commercial

options. Typed consent-gated aggregate events measure route choice without identity or

artwork data, while the evidence boundary prevents navigation from becoming qualified

interest, scholarly validation, demand, a sale, settlement, revenue, or income.

Commons intent measurement now continues through the provider-export boundary rather

than ending at browser markers. Strict ingestion retains the seven route counts,

rejects unknown or duplicate rows, derives their aggregate, and exposes both levels in

the private cockpit. This closes the instrumentation-to-operator gap while preserving

the separate qualified-action, settlement, and net-income evidence lanes.

A repository-public JSON Schema now specifies the complete normalized organic evidence

bundle, and the private cockpit links directly to it. A drift test compares the schema

event enum with the runtime allowlist, eliminating TypeScript/test inference as the

operator contract while retaining runtime rejection as the authoritative gate.

The authoritative importer now closes the remaining structural and chronology

gaps. Exact keys are required at every bundle level, exports must be captured

after their reporting window, lead statuses and settlement IDs are unique, and

payment state/mode/source/receipt values are validated before any projection.

Payment rows also require explicit USD currency, and a non-overwriting offline

settlement-ledger normalizer now rejects identifying/unknown fields, duplicate

IDs, invalid receipts, impossible amounts, and chronology inversions before

bundle assembly. Normalization does not authenticate provider evidence.

A non-writing organic evidence preflight now checks that full runtime contract,

exact release binding, and 30/60/90-day window before an operator can replace a

report. Its output is deliberately limited to digests and row counts, and the

writing command reruns the same gate to prevent preflight/import drift.

The legacy standalone report writer can no longer bypass that gate. Both

writers require the period, exact current release, and non-future provider

timestamps before their first write, with parity locked into the release

manifest and regression suite.

Raw-provider automation now begins with a bounded GA4 adapter. A filtered

two-column aggregate export becomes a non-overwriting normalized fragment

offline, with the runtime event allowlist shared directly with ingestion and

strict rejection of dimensions, duplicates, unknown events, invalid counts,

future capture times, and non-Analytics evidence hosts.

A deterministic offline assembler now joins all five normalized provider

fragments, injects the current release and declared reporting window, and runs

the complete preflight before a non-overwriting write. This removes manual

bundle editing while retaining provider authentication and evidence acceptance

outside the assembler. Every fragment now independently binds the exact window,

so valid evidence from different periods fails assembly rather than silently

mixing. A paired Search Console/Bing CSV normalizer also validates exact daily

click/impression rows, provider hosts, indexed-page counts, dates, and ratios

offline without provider or evidence authority.

The signed acquisition-evidence regression fixtures now replay those same

window-bound provider sources, closing the legacy-report bypass at the next

evidence layer.

All five fragments and signed source receipts now retain the raw-input SHA-256,

enabling replay against privately retained exports without placing sensitive

lead or settlement rows in normalized or public artifacts.

A lane-aware, no-write fragment replay command now regenerates every search,

GA4, lead, or settlement fragment from retained raw input, fails on any drift,

and emits only privacy-safe digests, row count, and false authority.

Temporal attribution is now enforced below replay: each lead state requires its

status-specific event in the half-open window, and every payment carries a

window-validated recognition instant through normalization, import, and report

calculation. Period metrics can no longer silently use lifetime status rows or

out-of-period settlements.

Period accounting now separates settled-order gross from refund/dispute

movements. Refund-only periods preserve negative net income, fees, and costs

without increasing gross, order count, or purchase conversion.

Five separate replay steps are now composed by an all-or-nothing custody

preparer. It validates every raw/fragment pair and exact-release bundle before

creating a new directory containing only the bundle and privacy-safe custody

manifest; drift creates no partial package and grants no authority.

as separate external steps.

The first external-evidence handoff no longer requires Release Operations to

hand-author deployment-receipt JSON. The private open-release approval workspace

downloads an exact-release-bound worksheet whose evidentiary fields remain null

and explicitly require provider proof. Both the browser and authoritative CLI

assemblers reject the incomplete worksheet, so this usability bridge cannot be

mistaken for deployment or an earned open-Research-Commons readiness point.

Desktop Chromium and Pixel 5 browser proof now download that worksheet, bind it to

the visibly rendered release digest, verify every provider fact is still null, and

pass axe plus horizontal-overflow checks. This is reproducible usability evidence,

not production-deployment evidence.

The federated provenance query UI and API were deployed to the canonical

production domain on August 18. Vercel build `metamuseum-mwcrf3ypb` completed

all 216 routes and the canonical supported/refusal probes passed. Compile-time

public research inputs now live under `src/data/research`, separating deployable

product data from ignored generated evidence. The deployment clears the runtime

availability blocker; a new exact-release quality digest, signed decision, and

provider receipt remain required before the open-release A+ point can be earned.

Offline normalization now covers the sensitive lead lane as well. Known

encrypted store rows become only three consent-state counts; plaintext or

unknown fields, duplicate IDs, unsupported states, time inversions, and output

overwrite fail closed, and no identifier-like value enters the fragment.

The public-page quality gate now includes an explicit regression contract for

research contribution controls at narrow mobile widths. The 390px Linux

Chromium failure is reproduced and corrected with a zero-overflow measurement;

GitHub Actions run `32589651885` and canonical smoke run `32589705508` passed.

The next public-surface audit increment replaces the former partial route list

with a 45-route manifest and dynamic coverage for every published Connections

and Stories page. The resulting 150-case desktop, zoom-equivalent, and mobile

matrix found two previously invisible defects: `/docs` escaped its viewport and

the Connections chapter numerals missed contrast requirements. Both now pass.

Homepage rotation also uses verified open-image records to repair sparse stored

provider groups, increasing observed representation from two sources to five in

the running app. Full fourteen-provider image representation remains open and

must be earned with item-level rights evidence and reliable display URLs.

The next provider-qualified increment adds Smithsonian American Art Museum to

the fallback rotation through a verified CC0 object and media record, bringing

the clean-deployment fallback to six museums. The quality gate is also now

reproducible on developer machines: regression checks use an isolated

seed-backed record store, while `--managed-storage` remains available for an

explicit operational drift audit of imported records. This prevents mutable

local collection data from masquerading as a CI code regression.

The first Connections editorial redesign is complete locally. Its public index

now previews three decisive turns, while the flagship detail page follows six

question-led steps from shared catalog subject through documented museum facts

to a bounded interpretive reveal. Every turn references declared claims, anchor

navigation is keyboard-addressable, and the route passes Axe and overflow checks

at desktop, zoom-equivalent, and mobile sizes. Expanding beyond one published

journey remains gated on equally strong source evidence and review discipline.

The public-route audit now includes Surprise Me. The former redirect could eject

visitors to an arbitrary provider page, including slow or incomplete third-party

states, while dropping Meta Museum's explanatory and rights context. The route

is now a force-dynamic internal artwork page selected from one verified lead per

provider, with explicit attribution, rights and source links, a non-ranking

boundary, and onward discovery choices. Deterministic tests prove that all

fourteen curated fallback providers are reachable and that later records from a

repeated provider cannot displace its verified lead. Canonical desktop and

mobile browser proof remains required for the exact deployed revision.

The static public-page audit manifest now names `/surprise` because the route is

intentionally absent from the canonical sitemap; this closes the browser-gate

coverage gap without presenting random responses as indexable editorial pages.

The next public-route increment replaces the sparse Research Reports list with

an editorial research desk. The sole Rosenberg report remains an under-review,

non-indexable lead, but its public index now exposes the motivating question,

source-backed significance, four evidence counters, unresolved evidence, exact

HTML and JSON routes, and a three-stage lead/review/approval ladder. The empty

reviewed-findings lane is deliberately prominent, and participation links route

to evidence contribution, the validation protocol, and the published method.

No reviewed finding, scholarly novelty, completed restitution, adoption, or

revenue is claimed. Full browser and canonical deployment proof remains open

for the exact revision.

The following public-route increment addresses the Linked Art Inspector. The

deployed tool was functional but supplied almost no orientation beyond its

editor: it did not state what was checked, distinguish inspection from

certification, or offer a route from the result to the model, datasets, or

records. The revised surface frames three explicit validation areas, a

read-only default, a three-step inspect/import workflow, and four inspectable

onward paths. Import authorization is unchanged and stays fail-closed.

Technical recommendation packet: behavior ownership remains split between the

server-rendered route framing and the existing client workbench. The three

highest observed risks are (1) a specialist editor appearing before a novice

can understand its purpose, (2) inspection being mistaken for certification or

publication, and (3) a successful result becoming a dead end. Curator

Experience owns the explanatory hierarchy and responsive layout; API Platform

owns inspect/import semantics and must preserve the permission gate; Quality

owns source-contract, API, browser, and accessibility regression checks.

Acceptance requires the exact public copy and four onward links, unchanged

import authorization, zero severe Axe findings, no horizontal overflow at all

three audited widths, and passing `pnpm test`, `pnpm lint`, and optimized

`pnpm build`. The next-cycle hypothesis is that another public tool route with

low context density will be the weakest remaining surface; falsify it by

re-running the canonical route inventory and direct browser comparison after

deployment.

The Getty provider workspace is the next repaired canonical route. Its former

layout exposed useful capabilities through repeated generic cards and raw

endpoint output, but did not connect those capabilities to an artwork or tell a

reader what the outputs could and could not establish. The revised experience

begins with a verified public-domain Getty image and source record, then follows

a numbered path from Linked Art and IIIF through read-only SPARQL to

ActivityStreams. Query results remain distinct from rights clearance,

attribution, provenance conclusions, and entity reconciliation; stream events

remain distinct from changes to, destruction of, or loss of a physical object.

Technical recommendation packet: the route remains a server component with URL

state, and the Getty adapter retains all provider parsing and network behavior.

The three highest risks were (1) technical capability without cultural context,

(2) raw query success being overread as a research or rights conclusion, and

(3) ActivityStreams deletion being mistaken for object destruction. Curator

Experience owns artwork-led hierarchy and responsive styling; Provider Platform

owns the adapter and read-only enforcement; Quality owns page, API, adapter,

browser, and accessibility evidence. Acceptance requires the verified featured

image/source/rights tuple, preserved shareable query parameters, explicit

semantic boundaries, no inline page styling, zero severe Axe findings, no

horizontal overflow, and successful full lint/test/build/smoke gates. The next

cycle should test whether the remaining lowest-context provider/tool route is

now `/research/query`, `/collections`, or another canonical surface rather than

assuming sparsity alone proves weakness.

The homepage visual-coverage gap is now closed in the release candidate. A

rights-qualified fallback round covers all fourteen production providers once

before any provider repeats, uses fourteen distinct image URLs, and keeps

collection-record provenance separate from image-source and license provenance.

Unknown or review-required rights are excluded. The carousel also drops failed

media during a session and disables automatic movement when reduced motion is

requested. Production completion still requires the full quality gate,

deployment, and canonical browser verification for this exact revision.

The first canonical exercise caught the AIC fallback IIIF endpoint returning

403; the release correction preserves the AIC object record while sourcing a

verified public-domain image of that exact painting from Wikimedia Commons.

The next low-context public-tool repair replaces `/graph`'s brief technical

header and pointer-dependent canvas with an artwork-led relationship atlas. A

verified public-domain Met sculpture demonstrates four documented node types,

then the live graph exposes every encoded edge through a text relationship

index linking both endpoints. Museum fact, catalog relationship, editorial

sequence, uncertainty, image rights, and inference boundaries remain visibly

distinct. The Cytoscape node-pivot behavior and tenant preview scope are

preserved.

Technical recommendation packet: the server route continues to own record

loading and graph projection, while the existing client component owns only

visual layout and selected-node state. The three highest risks were (1) a

specialist graph appearing without a cultural question, (2) pointer-only

navigation excluding keyboard and screen-reader users, and (3) shared concepts

being overread as influence, attribution, provenance, identity, or historical

contact. Curator Experience owns the artwork-led example and claim labels;

Graph Platform owns the derived nodes and edges; Quality owns page, component,

responsive, image, accessibility, and complete-link evidence. Acceptance

requires the verified image/source/rights tuple, preserved live pivot behavior,

all graph edges available as text routes, no page-authored inline styles, zero

severe Axe findings, no horizontal overflow, and successful full

lint/test/build/smoke gates.

The reconciliation review route is the next public-method repair. It now asks

the reader-facing question of when two records describe the same event, explains

the evidence reviewers compare, and states that confidence is a routing signal

rather than proof of identity, attribution, ownership, provenance, or catalog

authority. The four service-configured threshold bands, live decision counts,

human-review links, and optional AI-tiebreaker status remain intact. A zero-item

queue is explicitly a dataset status rather than evidence of completeness and

offers direct paths to entities, source datasets, and the canonical B6.1 method.

The route is read-only and continues to preserve rather than merge or rewrite

source records.

Technical recommendation packet: Curation continues to own threshold decisions

and reversible review, Data Platform owns candidate generation and provider

evidence, and Public Experience owns explanation and navigation. The three

highest risks were (1) a score being mistaken for proof, (2) an empty queue being

mistaken for a clean corpus, and (3) an optional model tiebreaker being mistaken

for publication or merge authority. Acceptance requires unchanged tested

threshold behavior, an explicit read-only and non-proof boundary, useful empty

state routes, one main landmark, no page-authored inline styles, responsive and

accessible browser evidence, and complete quality/deployment gates.

The route audit also closed an authorization mismatch on

`/curator/annotations`: anonymous visitors could render reviewer identity and

decision controls even though every review submission correctly failed closed at

the editor-only API. The workbench now requires an editor for `GET`, while the

read-only reconciliation methods route deliberately remains public. Acceptance

requires role-unit coverage, an anonymous browser redirect to the custom sign-in

surface, authenticated editor reachability, and unchanged API write protection.

The next public-route increment closes an irreversible-action defect on My

Collection. “Clear collection” no longer deletes every browser-local save on the

first click: it opens a labelled inline alert, focuses the non-destructive choice,

states that clearing cannot be undone, and requires a separate confirmation.

Collection data remains local to the browser and no account, sync, upload, or

sharing behavior is introduced. Acceptance requires reducer coverage for every

state transition, keyboard-operable controls, distinct destructive styling, zero

severe Axe findings or mobile overflow, and the complete build/deployment gates.

The visual-comparison route is the next repaired public research surface. It

now asks what changes when two museum images meet at the same scale, explains a

three-step close-looking method, and keeps the current pair's museum records,

rights language, and attribution adjacent to the viewer. Public selection

round-robins only the image-ready providers measured in the current record

store, retains one preferred rendition per artwork, and defaults to two

different works. The current seed/store truth is one image-ready provider;

fourteen governed connectors are documented separately and are not represented

as fourteen-provider image coverage. Optional rectangles are off by default and

labelled as Meta Museum demonstration guides, never museum annotations.

Technical recommendation packet: Public Experience owns the question-led

hierarchy and plain-language controls; Data Platform owns provider-aware

selection and rendition deduplication; Quality owns localized navigation,

desktop/mobile interaction, image loading, and accessibility evidence. The

three highest risks were (1) two renditions of one artwork masquerading as a

comparison, (2) connector count being mistaken for image-ready coverage, and

(3) visual similarity being overread as attribution, influence, identity, date,

or provenance. The optimized-server audit also found and repaired a Next.js 16

locale rewrite loop by distinguishing the propagated internal locale pass from

a new public navigation. Acceptance requires distinct default artworks, one

choice per artwork, truthful provider counts, source and rights links for the

active pair, demo guides off by default, a localized 200 without redirect loops,

zero severe Axe findings or horizontal overflow, and the complete

test/lint/build/deployment gates. The next-cycle hypothesis is that a canonical

Connections or Stories surface now carries the weakest combination of cultural

question, source visibility, and useful onward action; falsify it through the

route inventory and direct comparative browser audit after this exact revision

deploys.

The Connections audit found that the flagship episode itself is visually and

structurally credible, but its surrounding index stretched one approved journey

without an editorially useful onward route. Both the index and flower journey

now present the three existing source-backed Stories as companion trails. Each

card begins with a different question, retains its lead artwork, museum sources,

and public-domain statement, and links directly to the corresponding Story.

Visible copy prevents readers from treating those Stories as additional

approved Connections. Public Experience owns the shared companion component;

Editorial owns the question framing and format distinction; Quality owns exact

link, source, rights, responsive, accessibility, and image checks. The three

highest risks are (1) breadth being implied by relabelling Stories as episodes,

(2) an onward-action grid becoming generic card filler, and (3) source names or

rights disappearing at the format boundary. Acceptance requires all three

canonical Story routes on both Connections surfaces, explicit format language,

distinct questions and images, retained source and rights text, zero severe Axe

findings or horizontal overflow, and complete test/lint/build/public-smoke

gates. The next-cycle hypothesis is that the three Story detail pages now have

the weakest onward path into the wider collection; falsify it by comparing

their terminal actions and completion behavior after this exact revision is

deployed.

The companion-Story audit then found an exact terminal-path defect across all

three details: each finished with the same generic index, Explore, checklist,

and paid-kit links, leaving no narrative handoff from the clue the reader had

just followed. The shared Story template now computes a deterministic next

Story and previews it with a distinct entry question, artwork, provider set,

rights statement, and direct link. A separate longer-trail band routes readers

into the six-turn flower Connection and retains its museum-fact, inference, and

no-influence framing. Both discovery routes precede the optional paid method,

which remains clearly identified and secondary. Public Experience owns the

cyclic continuation and responsive hierarchy; Editorial owns entry questions

and handoff language; Data/Trust owns museum attribution and image-rights

retention; Quality owns route, ordering, image, focus, mobile, and accessibility

evidence. The three highest risks were (1) a paid offer interrupting cultural

discovery, (2) a generic “read another” action hiding the next intellectual

question, and (3) a related-story image losing its source or reuse boundary.

Acceptance requires one unique next-story route per detail, a complete three-

story cycle, the longer Connection route, discovery before product, retained

source and rights text, working high-resolution images, no severe Axe findings

or overflow, and full test/lint/build/public-smoke gates. The next-cycle

hypothesis is that the Stories index now understates the three-story sequence

because it presents a lead plus two supplements rather than a connected reading

route; falsify it through a production hierarchy and navigation audit after

this exact revision deploys.

The next flagship-Connection audit found a different narrative break: the

opening comparison was visually strong, but every one of the six explanatory

chapters became text-only precisely when the reader was asked to inspect a

flower, crop, horizon, or format. Each public turn now declares one or more

artwork evidence views in the journey model, including an intentional crop,

plain-language label, source-visible caption, and exact museum artwork target.

The final reveal restores both paintings side by side so “time reconstructed”

and “space reconstructed” remain a comparison rather than an unsupported

influence claim. Public Experience owns the responsive evidence-view system;

Editorial owns the crop labels and bounded captions; Data/Trust owns artwork-ID,

source, and rights integrity; Quality owns model validation, link, image,

keyboard, mobile, and accessibility evidence. Key risks are decorative image

repetition, crops implying facts beyond the record, and draft material becoming

publishable without visual evidence. Acceptance requires evidence imagery in

all six public turns, seven valid declared-artwork views, a two-work reveal,

visible rights attribution, no overflow at desktop or 390 pixels, and the full

test/lint/build/accessibility/interaction/deployment gates. The next-cycle

hypothesis is that the Connections index still overstates breadth around one

approved episode; test whether a clearer “one released investigation” frame and

an editorial preview queue improve credibility without promoting drafts.

That index hypothesis was confirmed: “Episode 01” and “Tonight’s mystery” made

one released investigation feel like a thin entertainment series rather than a

deliberate research boundary. The index now names exactly one released

independent investigation and introduces two future topics only as questions

under source review. Each notebook entry explains intellectual significance,

lists the verification still required, and links to both official museum

records; neither unpublished slug is routable or linked, and no unchecked image

or conclusion appears. Editorial owns question quality and the distinction

between inquiry and finding; Data/Trust owns official-record targets and status

derivation from the two isolated drafts; Public Experience owns the ruled

notebook hierarchy and responsive behavior; Quality owns exact-count, no-draft-

link, source-link, focus, overflow, and route checks. Principal risks are a

research queue being mistaken for publication, static questions going stale as

drafts change, and non-clickable cards suggesting missing interactions.

Acceptance requires explicit “not published” labels, four functioning official

record links, zero public draft routes, one and only one released feature,

desktop/mobile accessibility, and all deployment gates. The next-cycle

hypothesis is that the Stories index, which still frames one lead and two

supplements, now understates the deterministic three-story reading sequence.

That hypothesis is now implemented locally: `/stories` presents all three

released stories as one numbered route with six artwork images, equal editorial

weight, explicit source and rights context, and two visible next-clue handoffs.

The route ends in the longer flower Connection rather than a generic card grid.

Acceptance requires focused and full tests, lint, optimized build, desktop and

mobile browser inspection, accessibility and interaction audits, exact-head CI,

canonical smoke, and direct production verification before this slice closes.

The August 23 completion audit verified all four notebook identities through

the museums' collection APIs and received image/jpeg 200 responses from all

three companion image proxies. Direct 1440px and 390px browser inspection found

the released feature and notebook contained with no horizontal overflow. Local

launch evidence in `artifacts/launch/connections-index-a11y-2026-08-23.json`

passes 207/207 route-viewport cases with zero violations, while

`artifacts/launch/connections-index-interactions-2026-08-23.json` passes 67

routes, 4,795 visible links, 906 unique internal targets, and 45 fragment targets

with zero failures. The focused 27-test editorial contract, full serial suite,

ESLint, and the 218-page production build pass; the build used an ephemeral

process-only `AUTH_SECRET` and wrote no secret to the repository.

The current public-quality slice also replaces the raw documentation catalog

with a curated six-entry start path, searchable complete inventory, and readable

source view. The route inventory now enumerates 77 pages, 190 handlers, and 351

HTTP methods; the interaction audit covers 90 routes and 685 internal targets.

Duplicate nested main landmarks found by that expanded audit were repaired on

the docs hub and six Rosenberg research routes before release.

Production reader inspection additionally caught a duplicated source-title H1

that automated WCAG checks did not flag. The reader now removes only the leading

Markdown title from its rendered body and preserves the substantive section

anchors beneath one page-level heading.

The next public-discovery repair addresses the original cross-provider flower

search complaint. Ranking now removes presentation-identical records, excludes

zero-relevance filler when genuine query matches exist, places image-backed

works first, and interleaves equally relevant providers before repeating one

source. In all-provider mode, relevant no-image records move into a separate

linked source list with an explicit combined-provider count; they no longer

weaken the visual gallery or masquerade as missing artwork thumbnails.

Local browser evidence for `flowers&source=all&limit=18` now shows 18

image-backed gallery cards, eight separately labelled no-image source records,

six represented providers overall, and zero desktop/mobile overflow. Full tests,

lint, the optimized build, accessibility, and interaction gates pass locally;

canonical verification remains required after deployment.

The next route-quality slice replaces `/patterns`' raw twelve-card heuristic

dump. The public page now excludes administrative collection-page groupings and

generic or singular/plural duplicate labels, exposes six reviewable questions,

states the record count and non-inference boundary for every lead, and provides

twenty-four named artwork links. The record-quality queue is capped at eight

clearly labelled tasks, while the reviewed flower Connection demonstrates the

separate publication standard for a lead that survives source review.

Acceptance evidence now passes locally: focused selection/rendering tests, the

full serial suite, ESLint, the 218-page optimized build, 90-route interaction

integrity, and all desktop, 200%-equivalent, and mobile accessibility checks.

Exact-head CI, canonical smoke, and direct production inspection remain the

release boundary for this slice.

The Contact route now replaces an undifferentiated message panel with three

explicit, privacy-bounded paths: collection pilot, provenance review, and

support for the public work. Query input is resolved through a fixed allowlist,

only the approved subject reaches the visitor's own email app, and no contact

form or upload surface is introduced. The route explains what context belongs

in a first note, keeps confidential material out of scope, and makes clear that

an inquiry grants no engagement, publication, data-access, or payment

authority. Unit coverage locks the allowlist and rejects arbitrary subject

text; the complete organic funnel passes six desktop/mobile Chromium cases with

zero Axe or overflow failures. Canonical production inspection remains the

post-deployment boundary.

The first governed Calliope comparison run now covers both registered held-out

cases across five repetitions using Anthropic Haiku 4.5. Exact response usage

totals 8,395 input and 2,631 output tokens ($0.021550); every per-run cost stays

below policy. The maximum 13,130 ms latency breaches the 12,000 ms ceiling, and

independent blinded scoring remains incomplete, so the candidate retains no

promotion, deployment, or publication authority. Runtime responses now expose

only provider/model/token metadata needed for cost evidence and fail closed when

required usage is absent or malformed.

Calliope's first defect-repair cycle now passes the previously failed local

operational gates. Explicit malformed Linked Art rights statements refuse before

model spend, already-refused tasks cannot enter the provider branch, and the

Anthropic timeout is clamped below the 12-second value ceiling. The repeated

held-out rerun recorded five representative provider calls and five adversarial

pre-model refusals: maximum latency 3,920 ms, total exact cost $0.010505, and no

adversarial provider calls. Independent blinded quality and reviewer-time scores

remain the next gate; the result grants no promotion or runtime authority.

Calliope's v3 governed rerun now closes the missing reviewer-instrument gap. The

same representative and adversarial fixtures produced ten system observations:

five paid Haiku 4.5 object-label drafts and five repeated pre-provider rights

refusals. Exact usage was 5,360 input and 1,043 output tokens ($0.010575 total;

$0.002237 maximum per paid run); maximum system latency was 3,409 ms. Provider

prompts now carry a bounded source note, source URL, and rights summary with

explicit heritage no-inference boundaries, and only `end_turn` plus non-empty

content counts as model output. A private randomized packet, separate key, and

two-reviewer worksheet bind SHA-256 `b389e698…d50b94c`; public v3 trial and

readiness receipts contain aggregates only. No reviewer evidence was invented,

so Calliope remains constrained pending two genuine independent returns.

The OpenAI reconciliation tiebreaker safety cycle is locally complete. Exact

evidence-equivalent ties now abstain before provider spend; provider responses

may abstain explicitly, cannot select outside the supplied candidate set, and

must return valid model/token metadata. Requests are capped at five seconds and

latency plus usage are retained on the review decision. The focused 14-test

reconciliation suite passes, and the canonical ten-observation worksheet is

prepared. A real adjudicated comparison remains blocked on an `OPENAI_API_KEY`

and independent judgments, so the surface remains unproven rather than promoted.

The reconciliation trial is now a v2 orchestration test rather than ten direct

provider calls. Five evidence-equivalent observations abstain before spend; only

five distinguishable observations may call OpenAI. Its comparator is a strong

identifier-aware deterministic evidence rule that explains its selection, so a

model cannot claim value merely by matching an obvious authority identifier.

Janus runtime reconciliation and diagnostic tools also emit instrumented

receipts. With no OpenAI key or independent review, the decision remains to keep

the model disabled unless blinded reviewer effort improves materially.

Janus no longer trusts model-written reconciliation explanations. OpenAI may

return only a confined candidate or abstention plus supplied evidence-field

names; local postflight proves those fields distinguish the candidate and renders

the explanation deterministically. Paid postflight failures preserve returned

usage, cost, latency, and model identity instead of disappearing as free errors.

Voyage embeddings now have a credible conventional comparator: normalized

lexical unigram/bigram feature hashing replaces the former SHA-byte vector, so

token overlap produces meaningful cosine behavior. Provider calls declare

document-retrieval intent, disable silent truncation, enforce a 2.5-second cap,

validate response indices, finite consistent dimensions, model identity, and

usage, then expose aggregate latency/token telemetry. Deterministic fallback no

longer consumes AI-call quota. The balanced worksheet is prepared, but no

Voyage credential or independent relevance judgments exist; the grade remains

unproven.

The Voyage lane now has an enforceable pre-provider egress boundary. Model use

requires exact public-catalog, public-safe, and cultural-care-review literals;

missing declarations and obvious email, labeled telephone, private/non-public,

or culturally restricted content return a structured refusal before network

access or AI quota evaluation. API-rate gating remains first. Clearly reviewed

public catalog text still reaches the bounded, metered provider path, while local

deterministic embeddings remain

available without egress. No credential is configured, so this is safety/readiness

evidence only and does not change the 0/5 AI-value grade.

The embedding lane now also has a strict executable provider trial: two held-out

ranking cases (representative and adversarial false friend), five repetitions per

case, paired query/document telemetry, dated list-price accounting, and private

randomized A/B review materials with an aggregate-only public receipt. The trial

launcher was exercised, created no evidence artifacts, and stopped before

network access because `VOYAGE_API_KEY` is absent. Next is a governed provider run followed by two

independent cultural-heritage relevance/effort reviews; until then the grade

remains 0/5.

Voyage trial v2 repairs the evaluation itself. Opaque document IDs remove answer

leakage, private packets include the query and corpus needed for genuine judgment,

BM25 replaces hashed cosine as the deterministic comparator, and parallel

provider calls are measured by wall-clock pair latency. The receipt adds mean

reciprocal rank and ranking-stability gates; runtime rejects zero, oversized, or

query/document-incompatible vectors and retains paid postflight usage/cost. The

canonical launcher still stops before artifact creation because no Voyage key is

configured, so no provider advantage is claimed.

Voyage trial v3 expands the preregistered held-out corpus from two cases to six:

three semantic representative queries and three adversarial museum-assertion

false friends, each repeated five times. The runner derives all dataset, call,

safety, and tool totals from the manifest, enforces the registered $0.01 pair-cost

ceiling, and emits a strict replayable v3 receipt. Trial readiness no longer

credits the Voyage filename alone; it verifies the receipt against the exact v3

manifest before reporting provider evidence. Provider and independent-review

evidence remain absent, so the grade is unchanged.

Visual similarity now fails closed at its actual SigLIP boundary. Only public

HTTPS image URLs may leave the app, and provider use now requires exact public-

image classification, rights review, provider-fetch permission, and cultural-

care review before AI quota evaluation or network access. Route and service

callers share that contract. Model output must return an identified model and

one unique, finite, range-valid score for every supplied candidate, with no

invented URLs. Calls are capped at 3.5 seconds and return latency/model/candidate

telemetry plus an advisory-only contract that visual resemblance cannot establish

identity, attribution, influence, or provenance. The no-service fallback is now

truthfully named IIIF derivative/topology ranking rather than a visual heuristic.

The balanced worksheet is prepared, but no configured service or independent

expert judgments exist, so the model remains constrained and unproven.

The visual lane now has an executable rights-documented trial over public-domain

Art Institute of Chicago records and IIIF images: representative and adversarial

false-friend cases run five times each against normalized metadata overlap, with

candidate confinement, top-one stability, provider-reported-cost completeness,

latency, and private randomized expert review. Missing provider cost remains null

and blocks the gate. The launcher stopped before network access and created no

artifacts because `SIGLIP_SERVICE_URL` is absent; next is a governed service run

and two independent cultural-heritage relevance/effort reviews. Grade remains

0/5.

Visual trial v2 replaces the two-case design that gave the deterministic baseline

10/10 and therefore made the registered quality lift impossible. Six official AIC

series comparisons now produce thirty observations; normalized title, creator,

medium, and subject metadata scores 25/30, so SigLIP must be perfect to clear the

10-point lift threshold. Before model spend, the runner replays all 24 allowlisted

AIC public-domain records with bounded redirect-free requests and retains only an

aggregate digest. The cost gate now matches the audit's $0.02 ceiling, all counts

are manifest-derived, and readiness strictly replays v2 receipts rather than

trusting a filename. Live rights replay passed 24/24; no SigLIP evidence exists.

Clio's social-editor cycle now separates deterministic safety from generative

copy judgment. A reproducible house-style preflight flags hype, unsupported

certainty, endorsement language, and excessive punctuation; provider egress

requires a bounded public-safe evidence attestation. Anthropic responses must

end normally, identify the returned model, include exact token usage, obey the

ten-second cap, preserve retained copy exactly, materially change revisions,

and introduce neither unapproved URLs nor deterministic safety failures. Model

self-scores were removed. The real balanced Haiku 4.5 trial completed ten calls

for $0.016060 total with 3.899-second maximum latency, five retains, five

adversarial revisions, and zero local boundary failures. Its private blinded A/B

packet still needs independent quality and reviewer-effort scoring, so Clio

remains unproven.

Agent execution now produces a persistent evidence loop rather than a persona-

only response. Every Clio, Mercator, Janus, Themis, and Calliope run writes the

same `runTelemetry` receipt into its organization-scoped AgentTask: execution

kind, requested/executed provider state, returned model/tokens, known Haiku cost

or honest unknown pricing, latency, actual tools, source count, HTTP(S) Linked

Open Data record IDs, and citation URLs. Deterministic and provider-fallback runs

retain zero model cost. This upgrades latency/cost evidence from missing to

partial without changing any AI-value grade; held-out aggregates and genuine

reviewer time remain required.

The agent workbench now closes the previously missing human-measurement handoff:

each run exposes task and telemetry SHA-256 targets, starts a reviewer timer, and

offers a strict pseudonymous scorecard for decisions, quality, tool usefulness,

Linked Open Data grounding, and bounded correction reasons. The separate

append-only receipt rejects digest drift, duplicates, unmeasured time, free text,

and incomplete declarations; it is organization-scoped and explicitly cannot

change the collection, publish output, or promote an agent. This is measurement

infrastructure, not a completed observation, so the 0/5 grade remains unchanged.

OpenAI reconciliation now has an executable promotion-candidate lane rather than

an unpriced optional call. Exact returned token totals must reconcile, known

GPT-4.1 mini snapshots receive a dated official USD basis, and unknown models

remain unpriced. A strict held-out manifest drives five repetitions over one

adjudicated selection and one ambiguity/abstention case against the stable

highest-score-or-abstain baseline, producing a private randomized A/B packet and

aggregate-only false-authority receipt. The attempted run stopped before network

access because no `OPENAI_API_KEY` exists; provider and review evidence is pending.

The independent-review handoff is now executable rather than a labeled

baseline/model spreadsheet. `audit:ai-agent-value:blind-review:prepare` binds a

private A/B packet into a strict blank worksheet while withholding the

randomization key. `audit:ai-agent-value:blind-review:aggregate` requires two

distinct pseudonymous reviewers, exact packet and candidate-text replay, eight

complete output-quality dimensions, pair preference, and directly observed

review seconds before unblinding. Its public receipt contains only content

hashes and aggregates. The real Clio packet has a ten-row private worksheet and

privacy-safe readiness receipt; no human scores were fabricated, so the grade

remains 0/5 pending genuine returns.

Independent review now has a usable offline interface rather than requiring raw

JSON editing. The packet-bound form preserves blinding and strict response replay

while supplying accessible labeled controls, anchored scores, live progress,

automatic per-pair timing, declarations, and private JSON export under a no-

network content-security policy. Browser verification completed the workflow,

found and repaired long-digest overflow, and proved a 360px layout with no

horizontal overflow or console errors. Future Calliope, reconciliation,

embeddings, visual-similarity, and Clio trial launchers generate the form beside

their private worksheet; genuine two-reviewer returns remain the next gate.

Calliope's evidence loop now distinguishes executed tools from declared

capabilities. Instrumented receipts expose provider, grounding, rights, quality,

SEO, analysis, and fallback outcomes with privacy-safe hashes and bounded

metrics; the workbench shows their status and attribution. Other agents retain

derived tool attribution as an explicit evidence gap rather than inheriting a

false claim of equivalent instrumentation.

history, including the milestones consumed by `/api/roadmap`.

artifact handoff.

contract.

command ownership and preferred entry points.

standards rounds and fixture anchors.

architecture and SOTA acceptance criteria.

  • progress/era-history.md(progress/era-history.md): full Era A, B, and C slice
  • roadmap-to-10.md(roadmap-to-10.md): executable strict-readiness checklist and
  • risk-register.md(risk-register.md): open engineering and operating risks.
  • ops/review-goals.md(ops/review-goals.md): review-goals policy and command
  • ops/evidence-script-ownership.md(ops/evidence-script-ownership.md): evidence
  • linked-art/LinkedArtModel1.0-Reference.md(linked-art/LinkedArtModel1.0-Reference.md):
  • linked-art/LinkedArtSOTAWebApp.md(linked-art/LinkedArtSOTAWebApp.md): target

Visual-similarity trial v2 removes answer-bearing case names, item-label leaks,

and model/baseline score signatures from the private A/B review packet. The

metadata baseline normalizes simple plural variants, the aggregate manifest

digest binds the rights declaration, and paid SigLIP responses rejected after

receipt retain model/cost/latency telemetry through the API. No

`SIGLIP_SERVICE_URL` is configured, so no provider result or AI advantage is

claimed; next is a governed live run and two independent heritage reviews.

Mercator and Themis now report executed tools rather than persona-derived lists.

Mercator separately commits mapping-plan and MappingTemplate-validation results

with bounded column/error metrics; Themis commits rights summaries, provider

named-graph resolution, and final provenance review with known-rights and blocked

record counts. Janus and Calliope were already instrumented. This improves

reviewability without reclassifying deterministic agents as AI or claiming

benefit. Clio now closes the last runtime gap with separate instrumented receipts

for pattern analysis, conditional sparse-scope diagnosis, diagnostic synthesis,

and local drafting. All five agents distinguish observed execution from static

capability declarations; independent human value evidence is still required.

Clio social-editor trial v3 replaces the original straw-baseline packet with a

usable deterministic safe edit and gives both A/B sides identical original,

approved-source, and evidence context. Opaque case IDs and alternating placement

produce an exact 5/5 label balance. Ten real Haiku calls passed machine gates for

$0.015965 total with 4.603-second maximum latency; paid postflight failures now

retain model/tokens/latency. The old packet is superseded, and no benefit is

claimed until two independent reviewers score the current packet.

The shared independent-review aggregator now publishes privacy-safe per-dimension

model/baseline means and lifts, plus strict non-regression gates. A model cannot

pass by averaging a grounding or cultural-care loss against better prose scores.

Protected gates cover grounding, citation quality, calibration, robustness, and

cultural care; overall quality, pairwise preference, agreement, and review burden

remain visible, with all consequential authority still false.

Clio evidence-cited reader hook — provider delta ready, value unproven

The retired free-form social editor has been replaced experimentally by a

narrower `clio-evidence-hook` hypothesis. Haiku may reorder exact factual words

from bounded public-safe evidence into one or two concise sentences, but every

sentence must cite one or two known evidence IDs. Deterministic postflight rejects

unknown IDs, unsupported factual tokens, schema drift, unsafe evidence, and any

implied publication, attribution, or provenance authority. The comparator is a

strong deterministic concatenation of the same first two evidence statements.

V1 retained a $0.00052/1.199-second paid JSON-envelope failure; v2 retained a

$0.000515/1.128-second discourse-token mismatch. Although v3 completed, an

aggregate private-packet audit found that all five adversarial model hooks omitted

the required negated boundary while the deterministic baseline preserved it.

V3 is superseded and must not be assigned. Contract v2 marks fact versus boundary

evidence, requires every boundary ID, and requires explicit negation in every

boundary-citing sentence. V4 confirmed the repair; v5 adds public derived

grounding counts so readiness can replay 10 required boundary citations, 10

observed citations, and 10 preserved negations. V5 completed ten balanced calls;

all ten pairs differ from the deterministic baseline. Exact usage was 2,215 input

and 717 output tokens, $0.0058 total, $0.000617 maximum per run, and 1.388 seconds

maximum latency. The public receipt is semantically replayed

against both fixture hashes and recomputed pricing before readiness routes it to

review. No AI advantage is claimed: the private packet still requires two

independent blinded reviewers and must clear quality, grounding, citation,

calibration, robustness, cultural-care, cost, latency, variance, and reviewer-

effort gates.

The v5 assignment attempt exposed an orchestration mismatch: trial rows carried

extra private context fields rejected by the shared exact-schema parser. V6 moves

that context exclusively into the candidate envelope and emits the canonical

four-field packet row. Ten fresh calls retain 10/10/10 boundary counts, cost

$0.00581 total, and peak at 1.198 seconds. The v6 packet now has a privacy-safe arm-aware materiality receipt. It binds the

exact private packet, withheld randomization key, and public trial receipt by

SHA-256, then retains no candidate text or labels. All 10 pairs differ after

punctuation/case normalization, each case has one stable repeated model output,

mean token-set Jaccard is 0.942, and the model is 5.1% longer (143.5 versus 136.5

characters). This prevents cosmetic churn from being called reviewable while

making the unresolved tradeoff explicit; only independent reviewers can decide

whether the structural rewrite improves clarity enough to offset added length.

The shared assignment command accepts v6 and produced two distinct offline forms.

Their public receipt binds packet `1a89247b…b77cbcb` and both form digests while

retaining no reviewer codes, paths, text, labels, or identities. Readiness parses

that receipt against the v6 review-readiness receipt and routes to aggregation;

completed independent responses remain absent.

The canonical AI-value audit now consumes the already verified provider evidence

rather than displaying null machine metrics. Calliope v10 replays to 10 system

observations, $0.000979 mean cost, 747 ms mean latency, zero safety failures, and

passing machine gates. Clio evidence-hook v6 replays to 10 observations,

$0.000581, 953 ms, zero safety failures, and passing machine gates. Each is 60%

complete because blinded quality lift and reviewer-effort delta are still absent;

neither grade changes. Receipt absence yields no projection and receipt mutation

fails closed. Next remains two genuine independent returns per prepared surface.

AI evidence routing now uses `ai-agent-evidence-registry.ts` as a versioned leaf

contract for all six readiness surfaces. Readiness file discovery, assignment

bindings, provider verification, reconciliation source/retirement replay, and

audit machine-scorecard loading resolve canonical paths from that registry.

Focused drift tests prove unique surface/provider registrations and forbid copied

provider receipt literals in consumers. Canonical replay remains 4 provider

surfaces, 2 ready assignments, 2 retired behaviors, and 0 proven advantages.

The registry's five governed verification kinds now dispatch through one reusable

`ai-agent-evidence-verifiers.ts` table. Coverage validation fails on a missing,

duplicate, unknown, or unregistered verifier before any readiness projection.

Calliope, Clio hook, reconciliation/retirement, Voyage, and visual receipts retain

their existing exact parsers and fixtures. The CLI no longer owns family-specific

branches, while canonical replay remains 4/2/2/0 and grants no new authority.

Public assignment-receipt replay now joins the shared evidence verifier result.

Both exact assignment/readiness pairs verify canonically. Missing halves withhold

only assignment credit; a packet-hash mutation yields the bounded code

`ASSIGNMENT_RECEIPT_INVALID` while leaving provider evidence independently valid.

The CLI's separate assignment loop and binding export are removed. Readiness

remains 4 provider surfaces, 2 assignments, 2 retirements, and 0 proven benefits.

Unified readiness v3 now exposes privacy-safe `blockerCodes` per surface. The

canonical artifact contains six empty arrays. Mutation coverage proves an invalid

Clio assignment keeps provider evidence true, assignment readiness false, and

publishes only `ASSIGNMENT_RECEIPT_INVALID` plus generic prose—never packet hashes,

paths, parser text, candidate content, or reviewer metadata. The schema version

advances from v2 to v3; canonical counts and false authority remain unchanged.

Public review completion is now fail-closed: readiness semantically replays the

aggregate receipt and surface binding instead of trusting a filename. Forged

math, counts, gates, hashes, schema, or authority produce only

`REVIEW_RECEIPT_INVALID`; provider and assignment evidence remain independent.

The next external action is unchanged: obtain two genuine blinded returns for

Calliope v10 and Clio evidence-hook v6.

The public review verifier now binds completion to the full registered handoff:

verified assignment, exact readiness packet hash, surface, two-reviewer count, and

observation count. Readiness receipts themselves fail closed on scoring-dimension

or false-authority drift. This removes the remaining path for a structurally valid

but unrelated aggregate to receive completion credit.

Promotion evaluation now shares this verifier instead of calling the receipt

parser directly. The CLI rejects a wrong-packet aggregate before reading it as

subjective evidence or writing a candidate-comparison artifact. This closes the

downstream bypass while preserving private response-file handling.

Reviewer ergonomics now addresses the 10-pair navigation burden without reducing

evidence depth. Every pair announces complete/incomplete state, and a keyboard-

focusable control advances cyclically to unfinished pairs. New private Calliope

and Clio forms are bound by versioned v2 public assignment receipts. Browser

automation could not open the local `file://` artifact under its URL policy;

static accessibility contracts and end-to-end prepare/assign/aggregate tests pass.