- [ ] Prove or retire AI-agent value: inventory all 14 current agent/model-assisted surfaces, stop crediting eight deterministic heuristics/wrappers as AI advantage, and run repeated blinded comparisons for the six model-backed workflows against explicit deterministic baselines. The 28-case held-out representative/adversarial manifest and fail-closed comparison engine now cover repetition, complete registered case/run balance, duplicate rejection, leakage, quality lift, safety, cost, latency, within-case repeatability variance, and reviewer effort. Promotion additionally requires a parsed packet-bound receipt from at least two independent blinded reviewers covering every observation; worksheet booleans alone cannot qualify evidence, and grounding, citation quality, calibration, robustness, or cultural-care regression blocks promotion even when aggregate quality wins. The initial score remains 0 proven AI advantages; genuine reviewer evidence remains pending.
- [x] Make agent-value evidence operationally auditable: attach prompt sources, actual provider/model or deterministic engine, tools, state/queues, data access, consumers, authorization/privacy boundaries, cost instrumentation, failure modes, and real-or-missing eval coverage to all 14 inventory rows; add a privacy-minimized comparison CLI with immutable registered thresholds; and replace false LLM language in the deterministic Visual ETL Mapper while preserving its compatibility route.
- [x] Publish a machine-readable 17-dimension evidence ledger per surface and expose the deterministic-reliability versus AI-value boundary in the operator eval dashboard.
- [x] Correct usage accounting so deterministic chat, query, mapping, and visual-similarity fallback requests do not consume AI-call quota; retain AI accounting only for actual model-service invocation.
- [x] Expand comparative scoring to cover citation quality and operational quality explicitly, with non-compensating privacy/security, authorization, tool-discipline, and safety gates.
- [x] Generate directly runnable, privacy-safe comparison worksheets for all six model-backed surfaces: two held-out cases, five repeated runs, immutable thresholds, null-until-observed scores, and no raw prompts, outputs, or reviewer identities. Preparation and comparison runtime-validate the entire manifest for exact coverage, unique IDs, strict fields, safe relative fixture paths, and readable fixtures before writing evidence.
- [x] Make comparison receipts replayable without widening retained data: version 2 binds the exact scored worksheet and held-out manifest with SHA-256 plus manifest identity, retains only aggregate result/count/blocker fields, omits local paths and raw observation rows, and supports an explicit private-safe output destination.
- [x] Measure Calliope operational repeatability across real provider sessions: after v7/v8 privacy-safe failure receipts exposed brittle model-authored excerpts, `claim-evidence-v2` assigns exact evidence-line selection to deterministic tooling and confines Haiku to bounded claims/source IDs. The homogeneous v9/v10 stability receipt spans 10 paid calls, 10/10 pre-provider rights refusals, full excerpt/source coverage, minimum 0.667 lexical support, $0.00979 maximum trial cost, and 1,991 ms worst latency. It explicitly grants no user-benefit, deployment, publication, or audit-grade authority while blind human review remains incomplete.
- [x] Remove reviewer priming and assignment errors from the offline blind-review lane: reviewer-visible legends now expose only opaque pair ordinals rather than internal representative/adversarial case and run labels; a fail-closed command creates exactly two packet-bound forms with distinct fixed pseudonymous codes. Real desktop and 390px browser checks confirm one main landmark, labeled/unclipped controls, zero overflow, and zero console warnings/errors.
- [x] Make Calliope's independent review handoff executable and auditable: validate both reviewer codes before filenames are constructed, generate two exclusive offline forms for the exact v10 packet, retain a public assignment receipt containing only packet/form digests and counts with false authority, and route unified readiness v2 to private-response aggregation instead of recreating assignments. Genuine reviewer responses remain external and unclaimed.
- [x] Remove filename authority from Calliope provider evidence: readiness strictly replays v10 against both frozen fixture hashes, recomputed Anthropic token pricing, latency ordering, refusal safety, claim-evidence-v2 grounding, blockers, and false authority before exposing the reviewer aggregation route. Extra-key, count, fixture, cost, or authority drift fails closed.
- [x] Prevent configured credentials from silently enabling unproven models: Voyage embeddings and SigLIP visual similarity now remain deterministic unless an authenticated request explicitly sets `useModel: true`; explicit model requests fail closed with `MODEL_PROVIDER_UNAVAILABLE` when configuration is absent, and AI quota is charged only for genuine opt-in model execution. Public health reports model availability rather than claiming model enablement.
- [x] Align Clio with the same experimental authorization rule: the standalone Facebook review command now fails before file or provider access unless the operator supplies `--use-model`; the governed trial stays explicitly model-backed, while review output retains false publication authority and still requires exact-digest human approval.
- [x] Retire Clio's current model behavior when evidence shows no benefit: v4/v5 retain paid speculation-laundering failures; deterministic preflight now handles unsafe drafts without provider spend; v6/v7 bind 10 paid safe-draft reviews and 10 zero-cost adversarial edits; and a privacy-safe equivalence receipt proves all 20 blinded candidate pairs exactly match the deterministic baseline. Readiness now requires a new model hypothesis rather than pointless reviewer assignment.
- [x] Retire the current reconciliation model behavior when it underperforms: a provider-neutral adapter placed Haiku behind the same candidate-confinement and equivalence preflight as GPT-4.1 mini; two governed sessions retained 10 paid calls plus 10 no-call abstentions, but Haiku scored 10/20 against the identifier-aware deterministic baseline's 20/20 and had one earlier schema-inconsistent response. The aggregate retirement receipt blocks human review and further spend until a future hypothesis targets genuinely unresolved adjudicated cases.
- [x] Make reconciliation retirement reproducible rather than filename-driven: readiness strictly parses both historical source receipts against the frozen v2 cases, then recomputes every retirement aggregate, deterministic-dominance predicate, disposition, next hypothesis, and false authority. Provider or retirement path presence alone no longer changes routing.
- [x] Preserve reconciliation trial failures rather than losing paid evidence: trial-level aggregation now catches postflight, missing usage/cost, and provider failures; writes an exclusive aggregate-only receipt with partial progress, tokens, known-or-null cost, latency, and bounded category; excludes error text/prompts/responses/candidate evidence; and grants no retry, merge, deployment, publication, or audit authority.
- [x] Enforce reconciliation retirement at runtime and before new spend: the public server page always uses deterministic reconciliation and ignores the legacy model flag; the trial CLI no longer embeds the dominated v2 corpus and constructs no provider adapter until a strict v3 manifest passes a matching digest-bound, agreeing two-reviewer adjudication receipt, a deterministic-baseline-abstention eligibility check, and an evidence-equivalent adversarial check. Historical v2 parsing remains solely for receipt replay/tests.
- [x] Unify external-input orchestration across the six model surfaces: readiness v2 accepts alternative provider keys per surface, distinguishes assignment-ready from review-complete without retaining values or private paths, routes both Calliope and the evidence hook from semantically verified prepared forms to response aggregation, and routes the dominated Clio and reconciliation behaviors to hypothesis revision with false spend/deployment/publication/audit authority.
- [x] Prevent synthetic comparison output from masquerading as canonical proof: remove the test-generated promotion receipt, label every receipt `candidate-comparison`, retain false audit-grade/deployment/publication authority plus required external-evidence acceptance, and keep the canonical comparisons directory empty until genuine governed evidence lands.
- [x] Complete local authenticated browser and accessibility proof for the AI-value dashboard: anonymous users reach `/en/signin?callbackUrl=%2Fai-evals`; the existing non-production token-gated editor override renders `/en/ai-evals` through a loopback-only proxy while production remains fail-closed. Desktop and 375px-content checks preserve the zero-proven-value boundary, one main landmark, language/title, heading structure, unique IDs, named controls, image alternatives, zero console warnings/errors, and zero overflow/clipped elements after fixing long source and artifact path wrapping. Retain `artifacts/ai-agent-value/browser-proof-latest.json` as local UI evidence, not provider or deployment proof.
- [x] Replace the public documentation dump with a curated, searchable hub: six maintained starting points, complete title/path discovery, readable canonical documents, source dates rather than deployment freshness, stable section navigation, checksums, machine endpoints, and responsive no-overflow layouts.
- [x] Retain canonical production evidence for the homepage provider cycle: after pausing autoplay, fourteen manual advances produced fourteen distinct providers and fourteen distinct images from a 52-record rights-qualified pool with no duplicate URLs; preserve the point-in-time and no-partnership claim boundaries.
- [x] Establish a site-wide audit ledger and authoritative App Router inventory: classify all 77 pages and 190 route handlers (351 HTTP methods), eliminate every unassigned page, expand browser checks to all noindex/success/workflow surfaces, centralize all 15 anonymous protected-route expectations, and align private research approval/preflight pages with the proxy role policy.
- [x] Replace the homepage's generic gradient-and-card treatment with a museum-editorial visual system: warm paper, artwork-led composition, object-label metadata, serif display hierarchy, restrained rules, unboxed pathways, source visibility, responsive layouts without decorative glass effects, and WCAG-safe caption contrast.
- [x] Stop the homepage hero from restarting with the same artwork: generate a fresh deterministic order per browser session, preserve provider round-robin balance, retain every unique image, and disclose measured image-ready coverage against the fourteen governed production connectors without equating connector availability with reusable visual coverage.
- [x] Raise homepage visual coverage to all fourteen governed production collections with source-backed, high-resolution artwork candidates, explicit record/image/rights attribution, fail-closed rights qualification, provider-balanced ordering, deduplication, broken-image fallback, reduced-motion behavior, and measured live coverage.
- [x] Rebuild the Connections index and first public journey as artwork-led editorial experiences: frame the index as an episodic chain reaction, lead the journey with a visual mystery and two clickable verified public-domain works, follow a six-turn from-to chain with an explicit bridge in every chapter, use concrete looking prompts and a visual hinge to vary the rhythm, label museum facts versus editorial inferences and confidence row by row, end at a bounded reveal, keep provenance and rights visible, and preserve the machine-readable research layer without presenting a generic dashboard.
- [x] Refine the first Connection for reader choice and narrative momentum: add direct look/follow/audit entry paths, carry the clue forward between all six turns, subordinate repetitive claim rows inside labeled keyboard-accessible evidence disclosures, retain source links and fact/inference boundaries, and verify zero overflow or broken images at 1440px and 390px.
- [x] Stabilize the hosted accessibility gate by auditing the optimized production build rather than recompiling all 150 route/viewport cases through the development server; explicitly trust the CI localhost host and enforce both requirements with a workflow regression.
- [x] Bind the accessibility/overflow matrix to every canonical sitemap entry in addition to the explicit operational-route inventory, distinguish successful typed JSON/JSON-LD/schema resources from HTML documents, and pass 204/204 local route-viewport checks across 68 routes at desktop, 200%-equivalent, and mobile.
- [x] Rebuild Artwork of the Day as a source-backed daily editorial surface with a high-resolution image, compact fact ledger, explicit collection/image/rights verification links, attribution, stable dated navigation, and a provider-first UTC sequence proven to cover all 14 verified providers without image repetition across a 14-day window.
- [x] Browser-exercise the canonical homepage's paused provider-first round: 14/14 sources, 14 unique images, zero broken media, and visible rights context. Replace the 352px first-round Met thumbnail with a verified 1920px CC BY-SA derivative and lead every provider group with its manually verified display candidate while retaining unique stored records later in the rotation.
- [x] Add a canonical public-interaction integrity gate and clear its first defect set: enumerate the route inventory plus sitemap, render 66 routes, distinguish 58 HTML pages from eight machine resources, require one main landmark, inspect 2,784 visible links, probe 898 unique same-origin targets, validate 38 fragments, retry bounded rate limits, stop correctly at external redirect handoffs, and retain source-route diagnostics. Repair two nested landmarks, route protected browser traffic through the resilient custom sign-in page, disclose unavailable OAuth configuration instead of returning Auth.js 500s, and replace two non-resolving curated artwork links with authoritative Getty and Met pages. The optimized local candidate passes with zero interaction failures.
- [x] Replace the public professional-workspace card matrix with an intent-led institutional entry point: lead with collection, research, and pilot pathways; give every research/source route descriptive purpose and action copy; separate sign-in-required agent, organization, readiness, income-evidence, and worker controls inside an explicit operator boundary; preserve every registered professional destination; and verify one main landmark with zero horizontal overflow at desktop and 390px mobile.
- [x] Rebuild the public External Evidence Ledger as a human-readable proof surface rather than an internal card dashboard: lead with the current verdict and exact claim boundary, explain verified facts versus open gaps versus live reachability, present every row in a non-compensating claim ledger, move raw receipts into keyboard-native disclosures, separate current production probes from adoption and revenue proof, compact the tombstone history, and state what evidence would change the verdict without altering any underlying status. Add production-length timestamp, live-probe wrapping, scrollbar-safe full-bleed, and explicit small-screen gutter regressions after direct canonical inspection caught otherwise hidden horizontal-overflow defects in dynamic evidence values and `100vw` viewport math.
- [x] Rebuild the provenance-path assessment as a private editorial decision guide with visible possible outcomes, four-answer progress, a device-local recommendation receipt, explicit non-verdict boundary, preserved deterministic routing, and a desktop/mobile funnel test aligned with both live Stripe checkout and the intentionally disabled email course.
- [ ] Complete the route-by-route public experience audit, including every actionable control, content state, responsive breakpoint, image-quality boundary, source link, search path, accessibility gate, and conversion path; retain canonical production evidence for each corrected slice.
- [x] Add a verified, server-only Meta Museum Facebook Page connection and a supervised publisher that creates deterministic source-linked drafts, requires exact digest-bound human approval, keeps durable publishing disabled by default, sends credentials only in the authorization header, and writes token-free post receipts. The first release pairs human-sounding Sun & Rain Works attribution with a high-resolution 2400×1260 digital-cultural-heritage visual built around the Cleveland Museum of Art's faithful CC0 image of Van Gogh's Landscape with Wheelbarrow. Clio, the Greek-history-inspired social editor, performs a bounded evidence-only review, and the native Page photo flow binds the media URL into the approved digest while recording both photo and post IDs.
- [x] Operationalize Clio's August 22–28 launch plan as a validated seven-post campaign package with exact-plan hashing, source sets, candidate local times, per-post success signals, privacy-safe Facebook organic attribution, deterministic draft digests, and explicit false scheduling/publication authority. Runtime artifacts remain local; each image and post still requires rights review and exact human approval.
- [x] Finish the public My Collection experience with private browser-local saving, preserved museum sources and rights labels, relevance-ranked cross-provider discovery, a true total result limit, compact source guidance, and graceful provider-outage messaging that never exposes raw API failures.
- [x] Reframe `/records` as an honest public collection-network page: separate fourteen governed production connectors from the retained normalization sample, expose every source from the canonical registry, remove image-less/held records from the visual gallery, prefer full-resolution assets, and preserve complete records through the dataset export.
- [x] Replace the Stories placeholder with three complete, clickable, image-led editorial stories: Kōrin close looking, a bounded 1787–88 Houdon/Goya comparison, and uncertainty preserved in two Jacometto catalog titles. Every article separates observation from record facts, carries museum citations, materials and rights, emits Article JSON-LD, and appears in the sitemap.
- [x] Give Stories an editorial reading hierarchy instead of three interchangeable cards: one artwork-led opening, two deliberately secondary paths, visible source and rights context on every entry, and keyboard-addressable chapter routes within every article. Verify the index and a representative story at desktop and 390px with no broken images or horizontal overflow.
- [x] Apply the question-led Connections narrative system to every published Story: each turn now asks a concrete question, carries a visible clue-to-clue bridge, identifies its evidence class, and links directly to the museum records used in that turn. Rebuild the lead wave story as a five-turn Met/Cleveland trail whose surprise is a rejected direct-influence hypothesis: Cleveland's record points to Hokusai rather than Kōrin. Use verified 3400px CC0 Cleveland and 3811px Met Open Access images, preserve interpretation and rights boundaries, and verify both the index and story at desktop and 390px with zero broken media or overflow.
- [x] Add a consent-gated, UTM-bound Stories-to-Research-Kit assist with a free-checklist alternative; retain clicks as funnel evidence only, never revenue.
- [x] Rebuild the Research Kit offer as a conversion-focused editorial page: remove false card affordances, establish distinct purchase/terms/audience/contents/preview/boundary hierarchy, make all three sample tiles functional in-page links, verify desktop and 375px containment, and bind dedicated CSS into the commercial release digest.
- [x] Rebuild `/support` as a public-value editorial journey instead of two generic card grids: lead with the evidence promise, link to three inspectable live outcomes, distinguish research/review/infrastructure work, keep terms adjacent to the offer, provide a useful contact path while checkout is inactive, and preserve the exact-digest production activation boundary. Desktop and 390px checks show zero overflow and no nested landmarks.
- [x] Rebuild `/projects` as a public case study rather than an internal scorecard: lead with the cross-collection problem, correctly label the 25-second tour, provide a three-click live-product journey, explain reader outcomes before architecture, remove stale test/readiness scores, separate deployed capability from outside proof, add a collection-pilot path, and repair the broken architecture-overview link through the live docs renderer. Desktop and 390px checks show zero overflow and no nested landmarks.
- [x] Repair the pnpm/action-setup v6 workflow conflict by matching every workflow to packageManager pnpm 10.34.5 and enforce the shared pin with a repository-wide regression test.
- [ ] Renew the Meta Page token, re-run the read-only identity check, and obtain exact-digest human approval for the prepared Kōrin Rough Waves Facebook release packet; do not schedule or publish while OAuth code 190 persists.
- [x] Implement and verify a Stripe-hosted one-time supporter Checkout Session in the Sun & Rain Works sandbox.
- [x] Build the production supporter Checkout path behind an explicit disabled release flag, restricted live key, bounded one-time amounts, direct Stripe return verification, and a GET-only four-offer production conversion probe.
- [x] Preserve sponsor and institutional-pilot intent through allowlisted contact routes and prefilled email subjects; measure consented contact opens as interest only, require the contextual paths in the GET-only production probe, and keep clicks separate from received inquiries, commitments, or revenue.
- [ ] Obtain exact-digest approval for the supporter release, enable it only in Vercel Production, rerun the conversion probe to 4/4, and wait for genuine external payment plus settlement/fee evidence before claiming revenue.
- [x] Persist genuine Stripe-delivered sandbox events idempotently, reproduce immutable receipts after restart, and enforce zero verified sandbox revenue.
- [x] Build the approval-gated single-tenant Microsoft Graph outreach path with encrypted operational data, mailbox pinning, single-use human approvals, immutable send receipts, reply/opt-out suppression, and replay-safe subscription handling. Certificate-backed production reconnection, one controlled reply, and one explicit opt-out are verified against the operator-owned mailbox; suppression receipts are idempotent and unattended sending remains closed.
- [ ] Renew human approval against the current Research Kit release digest in `docs/ops/research-kit-current-release-review-2026-08-22.md`(ops/research-kit-current-release-review-2026-08-22.md), then probe the exact release; the live account, terms/refunds, durable webhook destination, and sensitive production configuration gates are complete, but no external sale is yet evidenced.
- [x] Repair the editor-only Org Scope Route Matrix layout: expand the desktop workspace to 96rem, contain the wide table rather than clipping the document, define stable readable columns and sticky headings, and provide semantic route cards for narrow screens with focused regression coverage.
This is the current, authoritative execution plan. It supersedes
development-roadmap.md(development-roadmap.md), which is retained as the legacy
pre-Next.js plan. Completed slice history lives in
progress/era-history.md(progress/era-history.md); the strict evidence checklist
lives in roadmap-to-10.md(roadmap-to-10.md); open engineering risks live in
risk-register.md(risk-register.md).
The architecture north star is
linked-art/LinkedArtSOTAWebApp.md(linked-art/LinkedArtSOTAWebApp.md). Provider,
schema, protocol, and validation work must map tests to fixture anchors in
linked-art/LinkedArtModel1.0-Reference.md(linked-art/LinkedArtModel1.0-Reference.md).
This document owns sequencing and stop/go decisions.
Immediate operational priorities — August 20, 2026
Public UI/UX A+ program: the canonical homepage baseline is now measured rather
than graded by impression alone (910 ms LCP, 0.00 CLS, Lighthouse accessibility
100, best practices 77, SEO 92). The first tested slice fixes localized canonical
metadata and the mobile-menu label mismatch; establishes a single primary hero
journey plus explicit audience paths; makes first-visit analytics choices compact
and even-handed; replaces generic conversion copy; reduces the footer from 25 to
18 links; and prevents oversized small-tablet artwork rows. The next gate is
shared page-header/card/form/state normalization across critical public routes,
followed by full quality gates and desktop/mobile production browser proof. The
baseline, boundaries, and acceptance matrix live in
the public UI/UX A+ program(product/ui-ux-a-plus.md).
The Research Commons public shell now follows that editorial system: a focused
question-led opening, four numbered research moves, an early device-local
workbench, three plainly differentiated help paths, and a separate standards
index replace the previous 17-link header and seven repeated cards. The full
18-journey desktop/mobile proof passes after aligning its Reports assertion with
the current evidence-desk heading; this changes presentation, not the existing
fail-closed research, publication, payment, or human-review boundaries.
The second tested slice separates buyer-facing pilot information from internal
sales operations: `/pilot` retains transparent pricing, scope, success metrics,
support limits, and a concrete conversation path, but no longer publishes the
activation ledger, prospect profiles, named outreach queue, or named target
accounts. Its 412-pixel mobile journey fell from roughly 24 viewports to eight
with no horizontal overflow or data table. The mobile footer now groups its
concise directory into two scannable columns while keeping the brand and public
promise full-width. Commercial-readiness controls now verify the public page's
new honest-status language without requiring internal operating evidence to be
rendered to buyers.
The shared launch accessibility gate now exercises 35 resolved routes at desktop,
a 200%-zoom-equivalent reflow width, and 390-pixel mobile (105 cases), and fails
on WCAG A/AA findings or horizontal overflow. The first expanded run found real mobile width defects in Insights,
the visual ETL mapper, and the Getty workspace; all three are now contained at
390 pixels. Production reruns remain part of the deployment gate. The browser
approval test now reads its generated capture configuration from the correct
download stream, restoring a warning-free lint gate and meaningful receipt
verification. Shared focus treatment now combines an explicit outline with the
existing halo, and reduced-motion users receive effectively instant transitions.
The retained browser gate also traverses the first six keyboard targets on six
critical public journeys and rejects missing, off-screen, or visually unmarked
focus.
Local production Core Web Vitals evidence now covers six critical routes plus a
stressed mobile interaction. Desktop LCP is 120–1,022 ms; stressed mobile LCP is
641–1,349 ms; menu INP is 104 ms; every measured CLS is 0.00. Explore's database
query dominates its 870–915 ms TTFB, so canonical production measurement remains
the decisive performance gate rather than extrapolating from localhost.
Auth host trust is now wired explicitly and fail-closed: only the literal
`AUTH_TRUST_HOST=true` enables forwarded-host trust. The retained performance run
was rebuilt with that opt-in and verified without localhost `UntrustedHost` noise.
The resulting full serial tests, warning-free lint, type diagnostics, and
216-route production build are green for the same local candidate.
The August 20 shipping rerun reconfirmed ESLint, the full serial suite, and the
production build using an ephemeral build-only `AUTH_SECRET`; no environment
secret was written to or staged from the repository.
The following public-journey slice normalizes Stories and Connections to the
shared page spacing, title/lede hierarchy, descriptive metadata, CTA language,
and readable feature width. Stories now labels its formats as previews rather
than implying published editorial content, Connections identifies its one
curated journey without inflating breadth, and Privacy no longer repeats the
analytics-event wording. The focused regressions, public trust suite, ESLint,
full serial suite, and 216-route build are green. Canonical production browser
proof and the five-session comprehension gate remain open.
The first canonical 105-case accessibility run passed every route except
`/iiif` at the reflow and mobile widths, where OpenSeadragon's HTML drawer
injected tile and navigator images without `alt` attributes. The deployed fix
labels each deep-zoom mount as a group and uses a scoped `MutationObserver` to
mark transient canvas fragments decorative. The subsequent canonical rerun
passed all 105 route-and-viewport cases with zero severe axe violations and zero
horizontal-overflow failures, closing this accessibility finding.
Canonical production mobile performance traces now pass the program thresholds
across Home, Explore, a representative artwork, Connections, Research Kit, and
Pilot: LCP is 615–2,489 ms, CLS is 0.00, and mobile-menu INP is 74 ms under Fast
4G and 4x CPU slowdown. The artwork detail is only 11 ms inside the LCP gate and
remains the watch item. Lighthouse SEO is now 100; its remaining best-practices
deduction is solely the documented third-party-cookie boundary on direct Met
artwork images, not a first-party failure.
Research Kit and support checkout returns now have explicit pending, invalid,
and sandbox-success presentation with assistive status/alert semantics and clear
recovery actions, without treating a redirect as payment evidence.
Commit `1706b595` is live on the canonical domains through Vercel deployment
`dpl_9RMjniY8bz56uADexLVHgTTYfkTm`; mobile probes pass all three state contracts
with zero horizontal overflow.
The privacy-safe PD0 evidence importer now encodes the exact A+ comprehension
gate: no more than 5,000 ms homepage exposure, separate offer and primary-action
results, five unique consented non-specialists, and at least four unassisted
joint passes. Generic task completion or `with-help` responses cannot satisfy it.
Linked Art community contribution: Meta Museum now has a conformance-tested
Ferdinand Bol/Rembrandt fixture that enriches the official Unmodeled
Relationships pattern with Getty's exact `ulan1102_student_of` predicate in
`assigned_property`. The case study and minimal upstream patch description are
prepared in
the ULAN relationship note(linked-art/ulan-associative-relationships.md) and
submission packet(ops/linked-art-ulan-upstream-contribution.md). The public
fork commit `84a7203` passes the upstream MkDocs build and pull request
open against `v1.1`. Community review and acceptance or merge—not submission or
Slack visibility—remain the evidence required for the A+ community-positioning
goal.
Revenue activation update (August 20): the retained Research Kit sandbox proof
still verifies payment, buyer-route delivery, and zero eligible test revenue.
The operator-owned live Stripe account is connected; a least-privilege `rk_live_`
key successfully accessed Checkout Sessions, and the enabled production webhook
targets `/api/stripe/webhook` with nine payment lifecycle events. Both secrets
are sensitive Vercel Production variables. The public offer and terms now state
immediate fulfillment, the 14-day refund policy, privacy boundary, and disabled
automatic-tax status. Owner approval is bound to release digest
`bfe54344d432de1a8a93a5b1a7219a01ea0b6ead749975dacd2babf97ac831eb` in
the launch approval(ops/research-kit-launch-approval-2026-08-20.md). The next
gate is exact-release deployment plus production probes; neither readiness nor
an operator rehearsal counts as revenue. A four-week attributable distribution packet is prepared for
August 21 through September 17; it authorizes no publication, contact, or spend.
The sponsor outreach gate remains closed with zero approved candidates and zero
messages sent.
Separately, the managed-pilot ledger records nine earlier first messages and no
replies; every recorded follow-up date has passed. The next pilot action is the
bounded, one-time, same-channel copy in
the August 20 follow-up packet(ops/pilot-follow-up-2026-08-20.md), subject to
accountable approval and suppression rules. Do not expand the cold list before
closing this follow-up cycle.
The shipped `main` release is healthy. Commit `3177c6c` built with Next.js
`16.2.12`, generated all 216 routes, and is live on the canonical production
domains through Vercel deployment `dpl_UZ46fBArKEAQAe4y35CcWpsrc28J`. The red
deployment in the dashboard is a separate Dependabot Preview build for
commit `6fa2828`; compilation and TypeScript passed, but page-data collection
correctly failed because `AUTH_SECRET` is scoped only to Production.
P0 — restore isolated preview verification
| Status | Owner | Action | Completion evidence |
|---|---|---|---|
| [ ] | Vercel operator | Generate a new high-entropy `AUTH_SECRET` for Preview only. Never copy or broaden the Production secret. | `vercel env ls` shows `AUTH_SECRET` for both environments without exposing either value. |
| [ ] | Vercel operator | Redeploy Dependabot commit `6fa2828` after the Preview secret exists. | The preview deployment reaches `Ready`; its log contains no missing-secret failure. |
| [ ] | Engineering | Test the dependency branch as an upgrade, not as an ordinary content preview: read the repository-bundled Next.js 16.3 guidance, then run diagnostic TypeScript, lint, the full serial suite, and a production-mode build. | All four gates pass on the exact dependency commit and the result is attached to the PR. |
| [ ] | Engineering | Review changes from Next.js `16.2.12` to `16.3.1`, React `19.2.4` to `19.2.8`, and the accompanying browser, database, telemetry, and type packages. | The PR records reviewed breaking/deprecation notes and either a merge decision or a specific rejection reason. |
| [ ] | Platform | Add a non-secret preview-environment readiness check so future preview failures identify missing required scopes before page collection. | A test or deployment check detects an absent Preview `AUTH_SECRET` while preserving the runtime fail-closed assertion. |
Stop/go rule: do not weaken `auth.ts`, introduce a shared development fallback
in production-mode builds, expose a secret in logs, or merge the dependency PR
merely because its preview becomes green. Merge only after the exact upgraded
branch passes the complete gate. This preview lane is operational maintenance;
it does not block or downgrade the healthy production release.
Next product evidence after P0
- Run the approved four-week provenance distribution experiment and retain
consented source/medium/campaign aggregates; do not interpret unavailable
analytics as zero.
- Complete the 50-case production AI evaluation and independent validation
program with three qualified researchers and one subject expert.
- Obtain real external reuse, qualified acquisition, and provider-confirmed
economics evidence before changing the current `3/11` A+ claim.
Autonomy A/A+ checkpoint — production feedback loops
The fail-closed A+ rubric and operator control surface are documented in
the supervised automation runbook(ops/supervised-automation-a-plus.md).
The revision-bound independent-review packet and strict response validator are
implemented; attributable external review remains the sole A+ evidence gap.
deployment. Ten public research, validation, feed, and revenue routes are
checked against `https://www.metamuseum.org`; every run writes a timestamped,
SHA-256-addressed receipt and uploads it even on failure.
receipts and a fail-closed GitHub issue alert. The monitor remains restricted
to configured HTTPS hosts, response bounds, and named semantic fields; drift
opens a publication hold and never updates claims automatically.
keys, dead-letter state, and operator escalation. Transitions use row locks and
rollback on invalid events; evidence completion always stops at an attributable
human-approval gate. Focused reducer/store/type tests pass. A production
database exercise and scheduled escalation check remain required before this
capability earns production-proven status.
Production proof v2 used Vercel's in-process environment runner, persisted a
row with monotonic timestamps, and stopped at `awaiting_human_approval` with
one visible escalation and no approval. The first run exposed a retrograde
timestamp defect; monotonic enforcement and attributable cancellation were
added, and that flawed row was cancelled rather than counted as proof.
recruitment approval; connect real commercial systems without fabricating
outcomes. Human approval remains mandatory for historical conclusions,
rights, publication, outreach, and financial commitments.
- [x] Trigger canonical-domain smoke tests after a successful Vercel production
- [x] Schedule the bounded connection source monitor weekly with immutable run
- [x] Implement Postgres-backed orchestration state, retry budgets, idempotency
- [ ] Complete genuine researcher/non-specialist validation after accountable
Autonomous first-profit loop
- [x] Define the Research A+ evidence gate as eleven non-compensating technical and external requirements. The evaluator emits `A+` only when every requirement has fresh, attributable, non-synthetic proof; otherwise it emits `ungraded`, preserves exact next actions, and authorizes no publication, outreach, email, or payment action. The initial retained baseline is honestly `0/11`, making the remaining external validation and outcome work explicit rather than awarding points for feature breadth.
- [x] Build the cohesive open Research Commons journey and bind its accessible desktop/mobile browser packet, public methodology/schema, contribution/correction/attribution contracts, and expert-review route to the A+ release evidence. The device-local question-to-packet workflow passes desktop/mobile accessibility and overflow proof, emits canonical digest-bound JSON, proposes only no-side-effect agent tasks, and links to versioned reports and independent review. The A+ command hashes the exact workflow, tests, browser receipt, schema, public policies, guides, and issue templates; local workflow, open-contract, and automation requirements now pass while external requirements remain missing.
- [x] Replace manual GitHub expert-review transcription with a deterministic, privacy-safe Issue Form importer. It discards account metadata, validates exact public fields, labels, source-and-finding rows, and declarations, retains and replays the bounded issue body, and hands the candidate to separate facilitator verification without inferring qualification or validation.
- [x] Replace manual public correction transcription with a deterministic, privacy-safe GitHub Issue Form importer. It discards account metadata, validates exact packet targets, public evidence, AI disclosure, and consent, retains and replays the bounded issue body, and emits an unaccepted correction envelope for independent human review without granting publication or impact authority.
- [x] Connect signed correction-import artifacts directly to receipt-bound correction propagation. The propagation verifier replays the retained GitHub body and nested envelope before deriving a distinct unpublished packet candidate, removing manual extraction while preserving independent acceptance and publication boundaries.
- [x] Add deterministic mixed GitHub researcher-intake batching. The offline router classifies up to 100 exported expert reviews and corrections by exact label contracts, rejects ambiguous or duplicate issues, strips account metadata through the specialized importers, writes non-overwriting routed artifacts, and replays the complete safe batch manifest before exposing distinct facilitator next actions.
- [x] Make mixed-intake handoffs directly executable without manual filenames or config editing. Every replayed item carries a deterministic artifact filename and exact command; direct correction propagation verifies retained packet/import file receipts, rejects overwrite and drift, replays the imported Issue Form, and emits only an unpublished candidate awaiting a distinct human decision.
- [x] Add privacy-minimizing GitHub Discussion category forms for bounded questions, reproduction notes, source discoveries, identity-collision warnings, research gaps, and reversible workflow proposals. The forms require public evidence, AI disclosure, limitations, privacy, and authority boundaries; because the repository Discussion URL currently returns 404, the activation guide forbids public linking until a maintainer enables matching categories and captures attributable rendered-form receipts.
- [x] Expose the structured public correction Issue Form beside the device-local correction builder. Desktop/mobile Chromium prove the exact off-site URL, new-tab isolation, aggregate form-open event, privacy and non-acceptance copy, accessibility, and no horizontal overflow without leaving the site; the same proof caught and corrected an unrelated 1.92:1 quiet-header-link contrast regression.
- [x] Complete correction-form funnel instrumentation end to end. The aggregate event is allowlisted, normalized from identifier-free provider exports, retained in the reproducible organic report, and displayed separately in the private cockpit; boundaries and tests prevent an open from becoming a returned or accepted correction, qualified action, validation, impact, revenue, or income.
- [ ] Run the production AI evaluation and independent validation program: at least 50 cases, three qualified researchers, one subject expert, a materially corrected versioned investigation, frozen-baseline improvement, and external citation/reuse/correction evidence.
- [x] Project the retained Rosenberg multi-source dossier into deterministic ranked claim-level results with evidence class, confidence, direct HTTPS citations, retrieval timestamps, uncertainty, rights boundaries, and an explicit no-live-search notice. Focused service and page tests protect query-bounded negative findings and the legal/restitution/publication refusal boundary. General multi-archive live search and external validation remain open.
- [x] Restore the canonical local compiler/lint/test/build chain after strict fixture typing drift. `pnpm typecheck:diagnostic`, `pnpm lint`, the full serial `pnpm test`, and `pnpm build` pass; the build uses the documented CI-only ephemeral auth secret and retains the production missing-secret failure contract.
- [x] Deliver the first public `/research/query` federation slice: ordinary-language input executes against the current Linked Art museum index, matches the reviewed Rosenberg retained archive package, ranks claim-level citations/uncertainty/rights, and exposes executed, retained, unsupported, and disabled lane status. API validation rejects malformed, oversized, contact-bearing, and raw-query-shaped input. The 16-journey desktop/mobile Research Commons proof passes axe and overflow checks. This is not yet general live multi-archive search; additional approved archive packages/connectors and production deployment proof remain open.
- [x] Reassess the French and BnF connector boundary against current official documentation. POP/Rose Valland remains manual-import-only because the official open-data inventory does not list that collection and no collection-specific machine/rate contract is documented; public application code is not treated as permission. BnF SRU is documented and Open-Licence eligible as a future bounded discovery connector, but remains unimplemented pending explicit approval to add a new external API.
- [x] Close the natural-language deployment-proof gap: the exact-release open Research Commons handoff, approval, bounded capture, replay verifier, tests, and operator documentation now require eighteen fixed probes—seventeen public resources including `/research/query`, plus one fixed privacy-safe `POST /api/research/query`. The POST receipt binds its request digest and must replay as ranked Linked Art claims with uncertainty, direct HTTPS citations, executed museum and retained archive lanes, and false consequential authority. A missing page, malformed answer, citation-free claim, or elevated authority cannot satisfy the technical open-release requirement.
- [x] Register `POST /api/research/query` in the organization-scope route matrix after the full release gate identified the omission. The route is classified as a provenance read over scoped museum records, with the immutable reviewed archive package unable to expose sibling-scope museum records; route coverage, API, and matrix tests pass.
- [x] Promote dossier disagreements into first-class federated results. `/api/research/query` now returns ranked `conflict` and `not-comparable` entries with exact comparison keys, source lists, direct citations, retrieval times, explicit unresolved boundaries, and false resolution/identity/legal-title/restitution/publication authority. Unsupported archive questions return no inherited contradictions. The public UI renders them separately from claims and refusals; the desktop/mobile 16-journey proof validates the unresolved boundary with zero axe or overflow defects. The production POST receipt now rejects missing or malformed contradiction evidence.
- [x] Bind deployment-smoke latency to the exact release rather than relying only on transport timeout. Every open-release probe retains a non-negative duration and fails capture/replay above 15 seconds; both fixed natural-language POSTs have a stricter 5-second ceiling. Invalid, missing, negative, non-finite, rehashed, or slow duration evidence is rejected. These production smokes remain explicitly insufficient for the separate 50-case p95/cost evaluation.
- [x] Production-prove the refusal path separately from the supported answer. The 19-probe exact-release contract adds a second digest-bound privacy-safe POST whose archive lane must be `unsupported`, whose refusal must state no reviewed package matches, and whose result must contain neither inherited contradictions nor archive claims. Capture/replay rejects a rehashed retained-match lane, contradiction, or archive claim. The public refusal journey passes on desktop/mobile, bringing the Research Commons proof to 18 journeys with zero axe or overflow defects.
- [x] Make multi-source breadth explicit and machine-verifiable. Every federated response now includes a normalized source-coverage ledger mapping source ID/URL, museum or Nazi-era lane, evidence classes, claim and contradiction ranks, and `indexed-result` versus `retained-reviewed-evidence` execution. All rows state `liveQueried: false`; unsupported archive questions expose no archive coverage. The production supported-answer receipt requires at least three distinct retained Nazi-era sources including primary-document and official-catalogue evidence, while replay rejects hidden breadth, duplicate IDs, live-query mislabeling, or a refusal response that leaks archive coverage. Desktop/mobile proof renders and verifies the ledger.
- [x] Build and browser-prove the approval-gated organic acquisition foundation: canonical robots and sitemap contracts, one flagship provenance workflow, ten search-intent guides, Article and Product/Offer JSON-LD, a no-email-required checklist, and a deterministic device-local assessment routing visitors to free, kit, or bounded-review paths. Encrypted optional double-opt-in lead storage, a digest-approved and suppression-safe nurture worker, bounded human-review inquiry scopes, deterministic PII-rejecting provider-export ingestion, reproducible net-income reporting, a retained production baseline, and exact activation plus 30/60/90-day plans are implemented. The editor-only Organic Income Operator Cockpit and private no-store API add release/source-digest validation, 35-day freshness alerts, assessment starts/completions/recommendations, unknown-not-zero semantics, sandbox exclusion, activation states, exact next actions, and deterministic period packets. Release-bound tests explicitly exercise assessment routing plus indexing, conversion, delivery, and settlement alert states. Desktop/mobile Chromium prove both public funnel and authenticated cockpit accessibility and no-overflow behavior. Production publication, provider verification, lead capture, email sending, and live checkout remain separately blocked on attributable approval and real receipts.
- [x] Package the organic foundation into a reproducible four-week provenance distribution experiment. Consented guide and funnel events now retain only sanitized source, medium, and campaign attribution; guide views are session-deduplicated; eight niche touchpoint links cover Linked Art, museum/provenance partners, research communities, newsletters, and owned surfaces. The content-addressed packet defines one primary funnel, supporting signals, minimum denominators, learning thresholds, and an explicit boundary that authorizes no external publishing, contact, promotion, or spend. See the experiment runbook(ops/provenance-distribution-experiment.md).
checkout probe, approved/rights-safe content selection, fraud/privacy/rights
stops, a USD 50 total cost ceiling, and a strict settled-external-revenue
definition. Synthetic payments, self-funding, clicks, unpaid invoices, and
receipt-free revenue fail closed.
excluding the unapproved Rosenberg report, research listings, Connections
previews, validation studies, and package-test pages.
merchant identity, price/cadence, cancellation, refund, and privacy terms.
The August 11 live probe confirms checkout is inactive, so no revenue can yet
be collected and the loop correctly reports `configure-production-checkout`.
payments, import immutable commercial receipts and attributable costs, and
continue until independently reproducible verified net profit is at least
USD 1. No profit is currently claimed.
exports from payment, CRM, contract, and invoice systems. Imports are bound to
the complete export SHA-256, emit deterministic revenue/contract IDs, reject
PII and duplicate external records, require operator attestation for qualified
conversations, require attributable prior financial approval for contracts
and invoices, and explicitly authorize no financial action. Actual provider
credentials, production exports, and revenue outcomes remain unproved.
- [x] Add the first direct-to-buyer product: a USD 29 Provenance Research Kit with a transparent offer page, fixed-price Stripe Checkout Session, product-specific metadata, consent-gated checkout-start analytics, and a versioned Markdown workbook. Fulfillment retrieves the Checkout Session server-side and refuses missing, unpaid, wrong-price, wrong-currency, or wrong-product sessions; production checkout remains inactive until a restricted live key and operator-approved commercial terms are configured.
- [x] Extend the commercial evidence model with distinct research-kit checkout/payment events, processor fees, gross and net kit revenue, a content-addressed sandbox proof command, and a release-digest-bound launch-readiness handoff. Sandbox proof uses the authenticated Stripe CLI, verifies the exact paid test-mode offer metadata, and exercises both the fulfillment handler and buyer-facing HTTP download route without trusting a stale local key. The handoff recomputes retained proof integrity and binds the conversion and revenue-model sources into its release digest. It performs no activation and remains blocked until live restricted-key configuration and attributable merchant/refund/privacy/tax/fulfillment approval exist.
- [x] Add a deterministic daily loop with immutable receipts, a canonical live
- [x] Add `/sitemap.xml` for approved public utility and revenue surfaces while
- [ ] Configure an operator-owned production checkout whose hosted page states
- [ ] Accumulate genuine eligible sessions and provider-confirmed settled
- [x] Add the commercial-system ingestion boundary for normalized production
The repository also includes a self-contained, leadership-facing architecture
and evaluation snapshot at
`metamuseum-project-architecture-overview.html`(../metamuseum-project-architecture-overview.html).
It summarizes the implemented stack and workflows while preserving the strict
boundary between local engineering capability and unearned production,
audience, or revenue claims.
The homepage trust-page regression test now accepts the repository's supported
one-provider fixture as well as multi-provider data by matching the rendered
`museum source` / `museum sources` grammar; CI no longer fails when managed
storage contains records from exactly one provider.
A+ public-content program
- [x] Replace the graph's 483-row relationship wall with a searchable, progressively disclosed index. The default page now exposes 24 relationship rows and 55 main-region links instead of 973, retains all source/relationship/target links on demand, and reflows its Cytoscape canvas without horizontal overflow at a 390 px viewport.
- [x] Reframe `/entities` as public cultural discovery and replace its 477-row DOM wall with eight-item server-addressable windows per entity type. The measured route now renders 66 rows and 90 main-region links instead of 477/491, adds people/place/material question paths, retains search/facets/tenant scope and previous/next access to every entry, states the source-identity boundary, and passes a 390 px filtered-page overflow check.
- [x] Reframe `/insights` from an operational analysis dashboard into an evidence-bounded artwork journey. Three reader questions lead into recorded dates, places, and relationships; each visualization identifies museum facts or interface summaries, rejects unsupported influence and movement claims, and exposes recorded places as keyboard-operable filters when place data exists. Desktop and 390 px checks show no horizontal overflow.
- [x] Reframe `/issues` from a 100-row operational triage queue into a public standards decision trail. The route now explains how Linked Art proposals evolve, separates GitHub source facts from Meta Museum interface readings, keeps all 321 issues searchable in twelve-item pages, focuses the selected interpretation for keyboard users, and reduces the default main-region load from 102 links/103 buttons to 14 links/15 buttons without removing source access.
- [x] Replace `/roadmap`'s internal slice ledger and filesystem-derived date with a current public plan organized as Now, Next, and Not yet proven. The route derives its completed public-content count from the authoritative roadmap, links five inspectable discovery surfaces, retains maintained-source and structured-JSON access, and explicitly refuses to turn technical capability into claims of adoption, endorsement, revenue, or impact. Desktop and 390 px layouts remain overflow-free.
- [x] Make the Rosenberg research deployment self-contained: Vercel retains the five exact tracked JSON inputs imported by production code while continuing to exclude all other generated evidence, with a clean packaging regression test.
- [x] Complete Stage 1 of the external provenance-research network: six versioned, validated research-only source records now cover scope, geography, dates, searchable fields, access, language, reuse rights, primary-document availability, discovery provenance, and fail-closed automation. See the registry contract(product/provenance-research-source-registry.md).
- [x] Complete Stage 2 as a safe manual federated-research lane: dated source-specific access assessments keep all six sources manual-only; bounded queries, raw-envelope SHA-256 receipts, Rosenberg collision controls, agreement/conflict/not-comparable matrices, fixtures, provenance-governance ownership, and release-ineligible outputs are enforced. See the Stage 2 workflow(product/provenance-federated-research.md).
- [x] Complete Stage 3 with a shared fail-closed network contract and one officially supported connector: Getty Provenance Index SPARQL is bounded, CC0-attributed, allowlisted, timeout/retry/size/result limited, kill-switched, receipt-backed, operator-gated, and permanently research-only. A live exact-label smoke passed; five unverified sources remain explicitly disabled. See the connector runbook(product/provenance-network-connectors.md).
- [x] Expand Getty Rosenberg discovery with fixed, bounded templates for Linked Art names/identifiers, actor types, objects, title-transfer sales activities, date ranges, and known entity IDs. The first retained live run finds one Getty Group and 25 sales-activity leads; review and primary-document linkage remain open.
- [ ] Complete the decisive Rosenberg research benchmark across at least three sources, including manual receipts for the five non-network sources, primary documents, a reviewed dossier, an uncertainty-first interface, and genuine researcher/non-specialist reproducibility testing.
- Publishing surface: the homepage now exposes a reader-first research spotlight, and the reusable report leads with the finding, significance, changed understanding, and unresolved question before the audit layer. A six-beat trail distinguishes documented February/May correspondence and the RBS identity anchor from the later-reported September handover, shows the bounded V.A.8 correction, and adds the official POP record's two hashed object images without calling them photograph 3492 or handover proof. Permanent HTML, report JSON, JSON Feed 1.1, and RSS 2.0 keep AI-assisted `research-leads` separate from researcher-controlled `reviewed-findings`; the latter remains empty until attributable review. Feed items expose researcher-decision state, limitations, missing evidence, an explicit validation boundary, and exact methodology, correction, and contribution routes. Four shared participation contracts route agents and researchers into reproduction, correction, method challenge, or independent review with exact evidence and privacy requirements, while denying publication, contact, novelty, legal, payment, impact, and income authority. The report page renders the same contracts as responsive evidence cards with direct, instrumented actions; browser proof covers accessibility, no horizontal overflow, exact GitHub targets, privacy, authority, and the unchanged noindex state. That proof exposed and removed a nested duplicate `main` landmark inherited from the report page. Canonical metadata, article social metadata, and citation-aware `Report` JSON-LD now present the proposed answer only as an under-review suggested answer; desktop/mobile proof parses the live JSON-LD and confirms no accepted answer, while indexing remains unavailable until publication approval. The pending value-release candidate binds the exact changed service and page digests. Canonical research routes bypass locale rewriting so dynamic report slugs and citations do not 404. Research pages, API routes, and the novelty command are registered in operational ownership controls. Email and webhook delivery remain future work.
- Approval-bound organic handoff: SEO candidates now require the publication envelope's validation release digest to equal the current research release, closing stale-but-valid approval reuse. A replay-verified ready candidate can produce one non-overwriting, privacy-safe organic-publication manifest with exact canonical, robots, sitemap, report-channel, structured-data hash, citation count, internal-link, feed, implementation-target, and verification expectations. A separate no-write verification mode rebuilds the complete manifest from the retained SEO candidate and rejects semantically forged state even when an attacker recomputes the outer hash. The manifest retains no private validation rows and authorizes no file write, publication, deployment, indexing, sitemap submission, outreach, or payment.
- Reports-index gate: `/research/reports` now remains canonical but `noindex,follow` and absent from the sitemap while no attributable publication-approved reviewed finding exists. Its state-honest `CollectionPage` JSON-LD exposes lead/reviewed counts and boundaries. Only reports satisfying publication-approved status, reviewed-findings channel, explicit approval, a dated accountable reviewer, and a retained publication decision with approver code, public HTTPS evidence, source-report/validation SHA-256 bindings, approval time, and four affirmative content decisions can activate hub/report sitemap entries; toggled flags without the decision fail. Both hub and detail page are direct Research Commons release-manifest members; desktop/mobile proof covers canonical, robots, JSON-LD, WCAG, and overflow.
- Current evidence: dossier v1 answers an actor-level question across Getty, ERR, Legacy Explorer, Rose Valland, Lost Art, Proveana, the official RBS, Archives diplomatiques, and the MoMA Paul Rosenberg Archives. Human-operated browser searches retain visible-text hashes, expanded-query behavior, bounded negatives, aliases, and collision warnings. RBS entry *5034 / OBIP 37.954 matches Braque, the mandoline/fruits/bottle description, 97 × 130 cm, and Paul Rosenberg as named owner. The diplomatic correspondence supplies object-specific primary corroboration for Brussels recovery and intended restitution. The updated official JDP-0056 record embeds two image files now retained locally with SHA-256 receipts, providing relevant image-based registration evidence; image-specific republication rights remain uncleared. A full review of MoMA V.A.8 found no secure photograph-3492 match and positively disconfirmed its strongest apparent candidate as _L'intérieur au vase noir_. The dossier remains a draft because final physical handover and independent review/testing remain unproved.
- Interface evidence: `/research/rosenberg` translates supported ordinary-language questions into visible fixed searches, shows uncertainty and all source outcomes before synthesis, rejects raw query syntax, and remains no-index. `/research/rosenberg/dossier` presents the frozen answer, primary evidence, claims matrix, source ledger, receipt hashes, disagreements, missing evidence, and limits for open checking; `/research/rosenberg/validate` presents the revision-bound independent-review and participant protocol. Genuine usability and reproducibility outcomes are still unmeasured.
- External-validation protocol: the current dossier revision plus primary, manual, 1947 archive-bridge, and official image-registration receipt hashes; independent-review task; three-researcher counterbalanced timing study; five-non-specialist open-response study; human coding rule; and fail-closed thresholds are frozen in `artifacts/provenance-research/rosenberg-validation-protocol-v1.json`. The evaluator rejects synthetic, unconsented, duplicated, revision-mismatched, undersized, self-reviewed, slow, irreproducible, or poorly understood evidence. The validation workspace supplies a client-local independent-review return form that binds the review to the frozen digest, requires three distinct HTTPS sources, captures corrections and a decision, and requires the facilitator to attest that a private identity-to-code mapping exists. A separate qualified-researcher study randomizes baseline/interface order, times both conditions, requires three citations and a substantive answer per condition, retains the facilitator-confirmed washout, and exports a richer revision-bound session. `pnpm research:validation:assemble` now validates and projects returned files into an evaluator bundle while retaining SHA-256 receipts for every review, researcher response, raw reader response, and human-coding decision; it rejects duplicates and refuses overwrite. None of these workflows uploads responses or establishes real-world eligibility by itself. No real rows have been entered yet.
- Participant collection: `/research/rosenberg/study` is a no-index, client-local five-minute comprehension flow. It collects explicit consent and eligibility, generates a pseudonymous code, opens the frozen report, and downloads two open responses bound to the dossier digest without uploading, retaining, or scoring them. The evaluator now retains both human coders' decisions and requires a distinct adjudicator for disagreement. Desktop and 375 px browser checks pass with no horizontal overflow; five genuine returned sessions remain required.
- Current-cycle recommendation packet: impact is limited to the public research-study journey and Rosenberg validation contract. Highest risks are self-selected participants, response-file loss before facilitator handoff, and private attribution mappings not being retained with review custody. A versioned unsent recruitment packet names three current official professional routing channels, exact professional/non-specialist/coder copy, uncompensated and privacy disclosures, non-endorsement language, and send-evidence rules. The approval-ready copy is now personalized to the canonical public operator Joseph Chirum and `art@sunandrainworks.com`; it discloses that email return metadata identifies the sender to the facilitator. Its exact dossier, protocol, candidate, and copy digests plus requested scope are frozen with `approval: null`. `pnpm research:recruitment:approval` now reports only `Accountable human approval is required before recruitment outreach`; nothing has been sent. After genuine approval, recruit one qualified independent reviewer, three eligible research practitioners, five eligible adults, two coders, and an adjudicator while retaining provider receipts and recruitment channels. Acceptance remains one attributable independent review, three counterbalanced researcher records, five unique revision-matched reader files, complete two-coder decisions, and a passing fail-closed evaluator.
- Novelty research: the reusable `research:novelty` workbench independently scores object-match signals and chronology conflicts while refusing verified identity, novelty, or publication. A newly discovered collision—two distinct 1938 Braques with identical 81 × 100 cm dimensions on adjacent Mangin references—corrected RBS *5036 versus the Lasker work from 85/99 to 60/99 (`possible`) and removed the qualified chronology conflict. The MoMA finding aid narrows decisive inspection to III.A.2.2.42, III.D.1.c, III.D.1.h, V.A.2, V.A.12, II.QQ.1, and II.QQ.8; a complete reproduction request is prepared, but MoMA reference intake is closed until September 8, 2026. The official French restitution workbook also exposes a 1945 Braque/Rosenberg row superficially overlapping the 1947 POP record JDP-0056. The correct primary target is now `209SUP/1`, dossier 45.15, views 199–1173—not the rejected E–I volume `209SUP/392`. Its inventory distinguishes the 1945 judicial-method and 1947 Belgium restitution groups, so the dates alone are not a contradiction. Internal pages 33–38 and 45–48, OBIP 37.954, Mangin pages 33–34, and transaction records remain required.
- Primary-document advance: `209SUP/1` internal page 33 (viewer media 1154) identifies a Braque _Nature Morte_, 97 × 130 cm, found with a Brussels dealer and claimed by Rosenberg; page 48 (media 1173) identifies _Mandoline, fruits et bouteille_, reports Brussels recovery, and describes restitution as intended after formalities. The retained hashes and continuous dossier sequence create an object-specific primary bridge to RBS *5034 and JDP-0056. Two official JDP-0056 images now close the relevant-image registration gap, but a signed September receipt is still required to prove final physical handover and photograph 3492 remains necessary to establish the cited archive-photo lineage.
- [x] Enforce five explicit reader-value promises and three-headline/two-opening package preparation.
- [x] Reject synthetic, specialist, unconsented, and undersized package-test evidence.
- [x] Put the story path before assurance detail without removing citations or boundaries.
- [x] Instrument headline impression, open, 25/50/90% depth, continuation, share, save, and learned-something signals.
- [x] Add a privacy-safe production evidence importer with denominators, monotonic-funnel checks, explicit thresholds, and independent-review/operator-release gates.
- [ ] Complete five real non-specialist Rosenberg package sessions; three headlines and two openings are prepared and correctly remain ineligible for drafting.
- [x] Reconcile Rosenberg identities and dated roles without collapsing Galerie Paul Rosenberg (Group) into Paul Rosenberg (Person); retain a sourced `operated-by` bridge and fresh live receipts for all three objects.
- [ ] Complete contradiction, historical-context, visual-rights, editorial, subject-matter, and operator reviews.
- [ ] Observe a valid production window and earn independent 95/100 with every rubric dimension at least 9/10.
See connection-reader-value-and-audience.md(product/connection-reader-value-and-audience.md).
Instrumentation and fixtures never count as A+ outcome evidence.
---
Status (evaluation baseline retained; operational checkpoint August 20, 2026)
<!-- BEGIN:PROJECT_STATS -->
<!-- Generated by `pnpm docs:stats`; do not edit by hand. -->
| Generated project stats | Current value |
|---|---|
| Next.js | `16.2.12` |
| React | `19.2.4` |
| App page files | root homepage + `76` non-root page files (`77` total) |
| API route handlers | `175` `app/api` route handlers |
<!-- END:PROJECT_STATS -->
Executive Assessment
Meta Museum is a strong, unusually complete Linked Art product and a credible
controlled-beta system. It is not yet defensible as institution-grade or
repeatable SaaS. The gap is mostly evidence quality, operational history, and
commercial proof rather than missing core product breadth.
| Category | Score | Evidence-based assessment | Immediate implication |
| ------------------------------- | ----------------------------------------------------------: | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------ |
| Product experience | 6.5/10 for the public; 8.2/10 for museum and research users | The source-backed professional journeys are coherent, but public discovery still exposes workspace language, operational controls, and specialist navigation before establishing a reason to explore or return. | Separate the public museum from the professional workspace and build one curiosity-led discovery loop before adding more platform breadth. |
| Linked Art and data semantics | 9.0/10 | Canonical JSON-LD, event-centric modeling, rights, provenance, equivalent identities, HAL relations, IIIF, and provider boundaries are deeply implemented and test-backed. | Sustain conformance; do not add another provider until readiness work clears. |
| Evaluator and trust story | 8.4/10 | `/projects`, `/evidence`, `/datasets`, `/docs`, and machine-readable APIs make the work inspectable. The public evidence ledger currently mixes contradictory artifact baselines. | Repair proof consistency before adding more evaluator copy. |
| Accessibility | 9.3/10 | `pnpm a11y:check` passed 18 routes with 0 violations on July 12. Auth redirects for protected routes were handled as expected. | Keep zero severe violations and add accessibility checks to any hero or navigation polish. |
| Engineering architecture | 8.0/10 | Boundaries, contracts, strict TypeScript, storage abstractions, tests, and evidence automation are mature. The route and script surface is large and operationally expensive. | Prefer consolidation and owner clarity over new surface area. |
| Reproducible local quality gate | 6.5/10 | Direct ESLint passed and the local review-goals gate passed, but the audit began with a stale closeout guard, direct TypeScript failed in generated `.next/dev/types/validator.ts`, and a direct full test run did not finish within 10 minutes. | Re-establish one clean, repeatable canonical gate before feature work. |
| Documentation and governance | 8.5/10 | The docs surface is rich, agent-readable, and guarded for drift. The previous roadmap had become a completion ledger rather than a decision document. | Keep this file current and concise; send completed detail to history. |
| Operational readiness | 6.0/10 | Production preflight is 20/20 and launch review is 7/8, but long-window SLO, uptime, KPI, and artifact-coherence proof remain red. | Clear deploy-environment proof, then let time-based evidence accumulate without manufacturing results. |
| SaaS and business readiness | 5.5/10 | The managed pilot offer and tenant-aware technical foundation are credible. There is no invoice-backed pilot, repeatable onboarding, retention, or margin proof. | Sell and execute one bounded concierge pilot before building self-serve billing. |
Overall decision: **8.1/10 as a local product and portfolio case study; 5.8/10
for strict public-production readiness.** The project is safe to keep demoing and
iterating. Broad institution-grade or profitable-SaaS claims remain blocked.
The July 21 live review confirmed that the primary public journeys render without
browser console errors. The same review found that `pnpm typecheck:diagnostic`
failed in test contracts and fixtures. Those type errors are now repaired:
`pnpm typecheck:diagnostic`, 59 focused tests, `pnpm lint`, and `pnpm build` pass.
The full serial `pnpm test` run completed `1,677/1,677` tests in `157.5s` on
July 28 after stale deployment-preflight and documentation-contract expectations
were aligned with canonical evidence. `pnpm review:goals:check`
now reports `22` production-proof blockers (`0` local-refresh, `0`
deployment-environment, and `22` real-world-evidence). Completing the canonical
local test run is no longer a P0 blocker.
The July 28 timeout investigation confirmed `1,704/1,704` tests pass in 154
seconds with a compact reporter. The canonical test script now uses that
reporter to prevent captured verbose output from stalling the command channel.
August 5 quality-gate refresh: P0.2 is green again after the 24-hour closeout
guard expired. Lint passed in `19.3s`, diagnostic TypeScript in `5.4s`, the full
serial suite in `150.4s`, and the Next.js production build in `50s`. The repair
keeps intentionally untracked `data/source-repos` mirrors out of tracked-text
drift scans and anchors synthetic SLO samples to each evaluation timestamp so
the gate remains reproducible after the original June fixture window. The three
new standalone evaluator briefings are preserved in the same clean worktree.
The architecture evaluation briefing now marks the canonical-gate risk closed
and carries these measured results instead of its earlier stale warning.
Readiness Boundary
reports `status: external-evidence-required`, `local gate status: passed`, and
`strict 10/10 gate status: failed`.
launch review passes `7/8`, with the launch-review Era C dependency still red.
paid-pilot proof, retention, and gross-margin evidence.
- [ ] Strict 10/10 readiness is not green yet: `pnpm review:goals:check`
- [x] The July 10 production deployment preflight passes `20/20`; the latest
- [x] The latest strict handoff reports `0` local-refresh, `0` deployment-environment, and `22` real-world-evidence blockers, or `22` production-proof blockers in total.
- [x] The evidence ledger, readiness dashboard, Era C, long-term, launch-review, and review-goals artifacts now share the same canonical blocker, SLO, uptime, and ActivityStreams values.
- [ ] Remaining strict proof includes 30-day SLO/uptime evidence, production KPI exports, durable ActivityStreams syndication including a genuine `Delete`,
- [x] Current launch status is governed by `pnpm review:goals:check`.
- [x] `pnpm review:goals:check` must be green before broad public SaaS claims.
The allowed claim is: **local gates pass and strong external proof exists; strict
10/10 production readiness remains blocked by external evidence.**
Evaluation Findings That Change Priority
- Public evidence consistency is repaired and regression-guarded. The
canonical handoff, README, roadmap, and live ledger now agree on `21 = 0 + 3
generated ActivityStreams evidence, so its pending Delete row cannot repeat
confirmed`Create`or`Update` types as missing.
- The local gate is not yet reproducible from one clean command sequence.
The closeout guard, a malformed generated Next dev validator, and a long-running
direct test invocation obscure whether a fresh checkout is truly green.
- The product breadth is sufficient. Fourteen provider lanes, 34 pages, 146
API routes, public docs, validation, reconciliation, ActivityStreams, IIIF,
agent review, and tenant-aware pilot controls are enough for the next learning
cycle. More breadth would dilute the evidence and customer work.
- The next product-polish gain is quality, not more sections. The current
hero can enlarge a low-resolution source image, while Core Web Vitals and
primary-journey performance do not yet have a concise public baseline.
- The public product needs a distinct reason to return. Cross-provider search
is useful, but it does not yet turn the underlying data advantage into stories,
surprising discoveries, personal collections, or connections that no single
museum site can show.
- 18`strict blockers. The ledger also reconciles partner confirmations with
---
Priority Plan
Priority is determined by trust and dependency, not by implementation novelty.
Do not start a lower tier while an actionable higher-tier exit condition is red.
Time-bound external evidence may continue accumulating in parallel.
P0 - Restore A Trustworthy Baseline (Now, 0-7 Days)
| ID | Outcome | Owner | Exit criteria | Verification |
| ---- | ----------------------------------------------------- | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| P0.1 | Keep public readiness evidence internally consistent. | Platform + Evidence | `/api/evidence/ledger`, `/evidence`, `/readiness`, review-goals, launch review, and the long-term artifacts agree on SLO sample/day counts, uptime, ActivityStreams observed/missing types, and current blocker scope. Fallbacks disclose missing artifacts instead of substituting incompatible values. | `tests/services/readiness-evidence-consistency.test.ts`; focused evidence-ledger, readiness, review-goals, and documentation-drift tests; `pnpm evidence:ledger:probe:check`. |
| P0.2 | Restore one clean local quality gate. | Platform | From a fresh generated state, `pnpm session:closeout:check`, `pnpm lint`, `pnpm test`, `pnpm build`, and `pnpm typecheck:diagnostic` all complete successfully. The generated `app/api/vanda/search/route.js` validator fragment is valid after regeneration, and test duration is recorded. | Run the five canonical commands and attach elapsed time plus the first failing test if any. |
| P0.3 | Clear deployment-environment proof drift. | Operations | Rerun production preflight, launch evidence, launch review, and public Era C evidence with production environment present; reduce the deployment-environment lane from `4` blockers to `0` without changing the real-world claim boundary. | `pnpm launch:preflight:production`; `pnpm launch:evidence:production`; `pnpm launch:review:production`; `pnpm era-c:exit-gate:public`; `pnpm review:goals:check`. |
| P0.4 | Preserve nightly k6 evidence artifacts. | Platform + Evidence | The Docker fallback writes `artifacts/performance/k6-slo-summary.json` as the host runner user, deletes stale summaries before each run, and fails when no fresh summary is produced. | `pnpm exec tsx --test tests/scripts/k6-slo-runner.test.ts`; confirm the next Era C workflow ingests one fresh SLO sample. |
Production database SSL drift was repaired on July 27: Vercel now uses the
existing `sslmode=verify-full` connection value, the production artifact was
redeployed, and `/api/ai/query` returned `200` through the public alias.
P0 exit gate: all local commands are reproducibly green, public evidence has no
cross-surface contradictions, and deployment-environment blockers are zero.
P1 - Convert Reliability And Demand Into Proof (Next, 1-6 Weeks)
| ID | Outcome | Owner | Exit criteria | Verification |
| ---- | --------------------------------------- | ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| P1.1 | Complete the 30-day reliability window. | Operations | One canonical source reports 30 distinct UTC days of complete passing deployed SLO samples, including the cold-record scenario, and at least 99.9% public-read uptime with failed rows aged out of the retained window. | Scheduled probes plus `pnpm longterm:evidence:public` and `pnpm era-c:exit-gate:public`. |
| P1.2 | Produce real SOTA KPI evidence. | Data + Curation | Production/Postgres or warehouse exports meet the reconciliation auto-approve and reviewed-precision thresholds; every capture row identifies its production source. | `pnpm monitoring:kpi-evidence:production`; `pnpm era-c:exit-gate:public`. |
| P1.3 | Close one invoice-backed managed pilot. | Founder + Product | One real buyer has a signed scope and invoice reference, one collection is activated within seven days, and support, required KPI, retention, and gross-margin rows are captured without placeholders. | `pnpm pilot:buyer-pack`; `pnpm pilot:activation`; `pnpm pilot:support`; `pnpm pilot:kpi`; `pnpm pilot:evidence --check`. |
The architecture evaluator now participates in the commercial pre-revenue claim control: it must disclose the absent invoice-backed pilot, remain `External evidence required`, and point to `pnpm pilot:buyer-pack` plus the tenant-scoped `pnpm pilot:evidence --check` acceptance gate. This improves the handoff but does not count outreach or local tooling as revenue proof.
August 5 ownership and operator-experience pass: the compact public mobile header and product-specific agent sign-in context are implemented and regression-tested. The operations-risk report now covers `35/35` page routes across `7` ownership/consolidation lanes alongside `61/61` API families and fully classified package-script namespaces. `pnpm ops:profile` makes the Next.js + Postgres portable baseline explicit and keeps Python services, AG2, Solr, GraphDB, and publication workers optional until their readiness gates justify enablement. The evaluator marks mobile and authentication closed while keeping broad surface area and specialist topology honestly managed rather than eliminated.
| P1.4 | Preserve honest ActivityStreams adoption. | Platform + Partnerships | Keep `3/3` real external consumers, `3/3` verified durable callbacks, and zero rejected subscriptions fresh. Add `Delete` only after a genuine upstream `404`/`410` tombstone and partner read; never synthesize it to satisfy the gate. | `pnpm providers:coverage:seed`; `pnpm activity:tombstone:scan`; `pnpm activity:subscriptions:guard`; `pnpm activity:syndication:evidence`. |
| P1.5 | Raise visible product quality and establish a performance baseline. | Product + Frontend | Home, Explore, one artwork detail, Projects, and Pilot pass mobile/desktop visual review; no hero image is rendered above a defensible intrinsic size; no text overlaps; LCP <= 2.5 s, CLS <= 0.1, and INP <= 200 ms on the agreed production profile. | Refresh public-trust screenshots, run a production performance audit, rerun `pnpm a11y:check`, and retain the metrics artifact. |
P1.5 performance checkpoint (July 28): all ten production cold-load traces pass
the agreed lab budgets. Explore mobile is the limiting LCP at `2,286 ms`, Pilot
desktop is the largest CLS at `0.0665`, and the representative interaction trace
is `28 ms`. The route matrix and profile are retained in
`docs/ops/frontend-performance-baseline.md`(ops/frontend-performance-baseline.md).
P1.5 remains open pending the mobile/desktop visual review and fresh axe run.
Step 6 remediation (July 28): the fresh axe run passes `18/18`. The two Home
blockers are fixed locally: the source band is now a normal full-width sibling
of the padded content container, and V&A IIIF services are promoted to a
1,200 px image derivative before thumbnail fallbacks. The repeated local
production-build matrix passes all ten mobile/desktop traces with 0.00 CLS, a
maximum 1,628 ms LCP, and 29 ms INP; the 18-route axe gate also remains green.
Repeat this matrix against the deployed revision to close P1.5.
Linked Art 1.1 watch checkpoint (July 28): the standalone
`linked-art-1-1-agenda-impact-tracker.html`(../public/linked-art-1-1-agenda-impact-tracker.html)
maps all 26 August 5 agenda issues to current support, expected impact, required
fixtures or schema changes, and pending community decisions. Update its decision
column and the canonical Linked Art reference ledger after the meeting before
changing validators or production mappings. Issue #637 now has an internal,
provenance-bearing
confirmed-negative reconciliation contract. It keeps curator-confirmed
non-matches separate from unresolved candidates and withholds Linked Art
projection until the community settles the property name and assertion pattern.
Issue #362 now has a provisional internal response-profile contract. Its
server-defined brief projection preserves canonical identity, marks itself
incomplete, and links deterministically to the full record; no new public
profile parameter is enabled before the community decision.
The pre-meeting implementation evidence packet now combines issues #362, #637,
and #780 with executable references and decision questions. Use it during the
August 5 discussion, then replace its pending questions with resolution links
before promoting any candidate behavior.
Linked Art 1.1 meeting checkpoint (August 19): the saved reference checkout now
tracks the upstream `v1.1` branch at `fded7e7`, 31 commits beyond the prior
`3ed503b` master snapshot. The standalone
`linked-art-1-1-august-19-2026-meeting-briefing.html`(../public/linked-art-1-1-august-19-2026-meeting-briefing.html)
records the supplied logistics and 21 agenda issues, current milestone counts,
post-August-5 label changes, issue #637's same-day reference-placement question,
and tested decision prompts. Agenda proposals and Meta Museum implementation
evidence remain explicitly non-normative until the community records outcomes.
An August 20 post-meeting refresh verifies that the reference branch is still at
`fded7e7` while the milestone has grown from 43 to 45 open issues. It records the
explicit #366 agreement to place the relationship on the affected thing; keeps
#524 and #637 gated with their clarified file-level and concept-versus-individual
boundaries; adds #804 and #806 to the watch list; and cites the Getty auction
example now recorded in #493.
The machine-readable meeting decision ledger covers all 26 agenda issues.
Issue #366 is now the sole resolved row. Agenda proposals are recorded separately
from outcomes; resolved rows require a
matching Linked Art issue URL, target release, and explicit local action before
they can drive post-meeting changes.
The Step 3 conformance pass now covers nine focused patterns under
`tests/fixtures/linked-art-1.1/`: qualified `AttributeAssignment` ambiguity,
inscribed `Name` evidence, `Name.created_by`, prototype-level provenance for an
unenumerated `Set`, member-side Addition and Removal, Person Joining and
Leaving, and auction selling/purchase separation.
Endpoint-family inspection now preserves the active terms-ontology inverse links
`added_member_by` and `removed_member_by`. The lifecycle fixture declares its
extension context explicitly and stays non-normative until the 1.1 meeting
decision is recorded.
The compatibility pass is now executable through
`LINKED_ART_1_1_COMPATIBILITY_BOUNDARY` and
`tests/quality/linked-art-1-1-compatibility-boundary.test.ts`. The audit keeps
six representative pending property placements rejected, leaves proposed and
deferred classes outside the endpoint map, and documents the post-meeting
promotion procedure in
`docs/linked-art/1.1-compatibility-audit.md`(linked-art/1.1-compatibility-audit.md).
P1 exit gate: 30-day reliability and production KPI rows pass, one real paid
pilot reaches first value, and strict ActivityStreams evidence is either complete
or explicitly waiting on a genuine upstream tombstone with all other rows fresh.
July 27 checkpoint: the three accepted durable callback rows are restored and
`pnpm activity:subscriptions:guard` passes `3/3`. Syndication remains honestly
blocked only on a real `Delete` activity read.
The nightly Actions environment now supplies all three production consumer IDs
to `activity:adoption:matrix`. Its production verification passed `12/12` feed
probes and resolved `3/3` declared consumers; `Delete` remains the sole missing
observed activity type.
A write-enabled July 27 tombstone scan checked 68 canonical upstream targets
across 14 provider lanes with zero errors and found no genuine `404`/`410`.
Accordingly, no `Delete` was minted; the scheduled scan must continue until a
real upstream removal can be observed and read by the three consumers.
The nightly workflow now runs the durable callback guard exactly once through
`activity:syndication:evidence`; the redundant standalone guard step was removed
without weakening its failure behavior.
Collection and readiness now have separate Actions semantics. The nightly
evidence workflow succeeds when probes and artifact generation work even if the
recorded status is red. The following `Era C Readiness Gate` workflow reports
those known external-evidence blockers without producing a failed scheduled job;
a manual dispatch remains fail-fast and owns strict Era C thresholds, durable
callback enforcement, and production launch review.
The Actions matrix now uses `actions/setup-node@v7`; execution-policy tests own
that major consistently, and runtime file metadata no longer depends on an
overload-derived Node type that can become optional in newer type packages.
Public Discovery Product Track (Next, staged behind the P0 gate)
Goal: turn Meta Museum's cross-provider data advantage into a welcoming public
museum built around curiosity, storytelling, and repeat visits. The professional
workspace remains available, but it must no longer dominate the anonymous public
journey. The defining promise is **connections no single museum website can
show**.
This track does not authorize a new provider, runtime dependency, or fully
autonomous publishing path. Each phase ships behind the existing rights,
provenance, accessibility, performance, citation, and accountable project-
operator release controls. Outside specialists improve assurance but are not a
prerequisite for ordinary source-backed research and bounded public experiments.
Connections use three explicit assurance tiers:
- Independent research — agents may discover, reproduce, rank, and draft
candidates without outside reviewers. Results remain internal and make no
novelty, historical-causation, rights-clearance, or scholarly-validation claim.
- Operator-reviewed public experiment — the project operator may release a
narrowly factual Connection when every material claim resolves to museum
sources, fact/inference labels and uncertainty are visible, media rights are
safe, sensitive claims are absent, and the page says it has not received
external expert review. This is the default independent operating lane.
- Externally validated — distinct qualified editorial, rights, and subject-
matter reviewers approve retained evidence. This tier is required before
claiming scholarly novelty, expert validation, sensitive provenance or
identity conclusions, externally cleared rights, or commercial licensing.
The project may advance through tiers 1 and 2 independently. Missing external
review blocks only tier 3 claims; it must remain visible as missing evidence and
must never be filled by an agent or fixture identity.
Independent-lane implementation checkpoint (August 9): the typed Connection
contract, public page, and machine record now expose assurance tier, external-
review status, and a revision-bound operator-release record. The current Flowers
journey is visibly labeled independent research and not externally reviewed.
Promotion to an operator-reviewed experiment requires a non-synthetic `OP-...`
human operator code, timestamp, release notes, and completed citation, rights-
boundary, sensitivity, and disclosure checks; agents cannot satisfy the record.
Live-source checkpoint (August 9): the reviewed independent-lane manifest made
four representative channel calls plus four candidate-source calls and retained
timestamped, hashed receipts for a Met API
record, bounded Getty SPARQL result, Getty IIIF manifest, and Getty Linked Art/LOD
record. All four returned 200 with expected media types; the capture report has
zero failures and zero ingestion rejections. Native payloads are receipt-only
until channel-specific mappings are reviewed, so capture evidence does not imply
local Linked Art conformance.
Independent discovery checkpoint (August 9): a four-record Met/AIC normalized
input passes deterministic pattern analysis and produces a medium-confidence Van
Gogh entity-reconciliation candidate plus bounded provenance-coverage notices.
Agent-assisted research ranks two public-story candidates: matching museum-
supplied year and maker metadata for _Wheat Field with Cypresses_ and _The
Bedroom_, and matching supplied wave labels for _Rough Waves_ and _Under the Wave
off Kanagawa_. The shortlist retains live hashes, facts, inferences, uncertainty,
rights boundaries, and refused novelty/causation/sensitive/licensing claims. Both
remain independent research with zero publication actions pending operator review.
Draft-contract checkpoint (August 9): both shortlisted candidates now satisfy the
full `CuratedConnectionJourney` contract, including synchronized story/claim
references, research method and rejected hypotheses, assurance disclosure,
pending revision-bound operator release, source-controlled rights, corrections,
licensing, and measurement. They remain in a separate draft collection; tests
prove neither slug is available from the public resolver or feed. The operator
packet at `docs/product/independent-operator-release.md` is the next gate.
Source-monitor checkpoint (August 9): `pnpm connections:source-monitor --
--check` compares semantic claim fields for all four Met/AIC sources while
excluding volatile response timestamps. The first live run is clean for 4/4
sources with no publication hold. Any field change, fetch failure, unsafe
redirect, invalid JSON, or oversized response opens a hold and requires a new
journey revision plus operator decision before release.
| Phase | Outcome | Initial scope | Exit criteria | Verification |
| ----- | ----------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| PD0 | Establish the public baseline and content boundary. | Record first-time task completion, artwork-to-artwork continuation, return visits, public-domain downloads, and share events. Define the minimum quality bar for public records: usable image, intelligible title/date, source, attribution, and resolved reuse message. | Baseline artifact exists; public and professional audiences, routes, vocabulary, and analytics events are explicitly separated; records below the quality bar cannot enter featured feeds. | Public-route inventory, analytics event contract, quality-filter fixtures, privacy review, and five non-specialist usability sessions. |
| PD1 | Separate the public museum from the professional workspace. | Public navigation becomes Home, Explore, Stories, Connections, My Collection, and About. Evidence, APIs, agents, imports, annotations, org status, and operational controls move behind one clearly labeled “For museums and researchers” entry point. Rewrite the homepage in plain language around art and discovery. | A first-time visitor can explain the product and reach an artwork without encountering workspace status or specialist implementation language; professional routes remain directly reachable and unchanged in capability. | Mobile and desktop visual review, keyboard pass, `pnpm a11y:check`, public-navigation tests, and moderated five-second comprehension checks. |
| PD2 | Ship the minimum delightful discovery loop. | Add `Surprise me`, a rights-safe Artwork of the Day with a stable dated URL, and public artwork pages led by image, essential facts, “Why this is interesting,” related works, and previous/next discovery. Collapse technical metadata and researcher feedback below the public story. | The complete loop works: entry point → artwork → short story → related discovery → another artwork. Featured records never have broken media or ambiguous reuse messaging. | Deterministic selection tests, record-quality tests, mobile/desktop E2E, share-preview checks, and measured artwork-to-artwork continuation. |
| PD3 | Launch Connections as the signature feature. | Discover non-obvious event, provenance, exhibition, authority, material, temporal, and graph relationships; reject bare shared-word, maker, year, place, material, or classification matches. Present the strongest findings as visually led, layered articles with accessible stories and expandable evidence. | At least three operator-reviewed public experiments span three or more institutions collectively. Each passes frozen rubric v1 at 95/100 or higher with no dimension below 9/10, demonstrates knowledge unavailable from one source page, contains no unsupported causal or sensitive claims, and preserves human-controlled release. | Triviality-filter and rubric tests; live hashed source receipts; candidate contradiction checks; independent final score; citation, rights, accessibility, adversarial, and journey E2E checks; five-session comprehension and production engagement evidence. |
| PD4 | Add source-backed Stories and guided exploration. | Publish short image-led stories using reusable formats such as “One artwork, three details,” “Same year, different worlds,” “A disputed identity,” and “How an object changed hands.” Add exploration by subject, place, century, and color; add mood only as clearly labeled interpretation. Hide empty maps and timelines. | A minimum viable editorial cadence is sustainable; every factual claim resolves to a source; guided filters return useful results; empty analytical surfaces do not appear publicly. | Story-schema and citation tests, filter-quality samples, editorial review log, structured-data/share-card validation, and completion-rate measurement. |
| PD5 | Make trustworthy reuse and participation useful. | Add public-domain image download with source, rights, attribution, and metadata. Add “What changed?” with plain-language record version comparisons. Allow local-first saved collections, then optional account sync and read-only sharing after demand is observed. | Downloads package correct rights context; record changes identify source, time, and change origin; a visitor can save and share a coherent collection without being forced to sign in first. | Rights/download fixtures, version-diff tests, local-storage and account-migration tests, privacy review, and collection share E2E. |
| PD6 | Prove retention before expanding. | Evaluate artwork continuation, Surprise Me use, story completion, connection opens, downloads, collections created/shared, and 7-day return visits. Improve the strongest loop; retire or revise weak entry points. | Two consecutive measurement windows show a credible repeat-use signal and no regression in accessibility, performance, rights, or citation quality. Any further personalization or recommendation work has a measured hypothesis. | Analytics review, usability replay, public performance/a11y matrix, editorial quality audit, and a written continue/change/stop decision. |
Recommended first release: PD0 + PD1 + PD2 + three PD3 journeys. It must
demonstrate one complete public loop:
Interesting entry point → beautiful artwork → understandable source-backed
story → unexpected cross-museum connection → another discovery → save or share.
Public discovery stop conditions: pause expansion if featured-record quality
cannot be guaranteed, if connection evidence cannot support the displayed claim,
if rights context is separated from a download, if public pages regress the
agreed accessibility or performance budgets, or if measured use shows no
improvement after two iterations. Missing outside reviewers is not an independent-
research stop condition; it prevents promotion to the externally validated tier.
PD1 implementation checkpoint (August 6): the primary public navigation is now
Home, Explore, Stories, Connections, My Collection, About, and one quiet “For
museums and researchers” entry. Anonymous Explore and artwork journeys no longer
render organization status or professional workspace chrome, and public Explore
suppresses import prompts, roadmap language, and provider implementation notes.
The homepage now leads with cross-museum discovery in plain language; dedicated
Stories, Connections, My Collection, and professional-workspace landing pages
make every navigation destination intentional. Automated navigation, homepage,
workspace-boundary, and Explore acceptance tests are green. The local production
build passes, the 18-route accessibility matrix reports zero severe violations,
and a 375 px browser review finds no horizontal overflow. A fresh deployed visual
review and non-specialist comprehension sessions remain PD1 evidence tasks rather
than reasons to reopen its implementation scope.
PD2 implementation checkpoint (August 6): `/surprise` selects from the same
image-backed, publication-eligible local artwork pool as the homepage and sends
visitors directly into a public artwork journey. `/today` resolves to a stable
UTC-dated `/today/YYYY-MM-DD` page whose selection is deterministic for that date.
The homepage exposes both entry points. Anonymous artwork pages now lead with
“Why this is interesting,” keep facts and source detail in an expandable section,
offer previous, next, Surprise Me, and related-artwork paths, and withhold
researcher annotations and operational relationship tools. Signed-in researchers
retain the complete professional view. Selection and surface acceptance tests are
green. The production build passes, the 18-route accessibility matrix reports
zero severe violations, Surprise Me resolves into an eligible local artwork, and
homepage, daily, and artwork routes show no horizontal overflow at 375 px. A
deployed review remains necessary before promoting this local checkpoint to
production proof.
August 7 deployed checkpoint: PD1/PD2 is live on the production alias. The full
serial test suite, lint, diagnostic typecheck, local and Vercel builds,
production 18-route axe audit, 66-check crawler preview, 20/20 deployment
preflight, public Explore smoke, repeated public-trust screenshot baseline, and
zero-advisory dependency audit pass. Ten retained Lighthouse captures report
100 accessibility and 0 CLS; the desktop routes pass the LCP budget, while the
stricter Lighthouse mobile profile reports 3.7-5.9 second LCP and keeps P1.5
open for remediation and an agreed-profile recapture. The refreshed adoption
matrix passes 12/12 operator-run endpoint probes for all three named consumer
IDs and remains blocked on genuine `Delete`. Those probes do not substitute for
fresh reads made by the external consumers themselves: the retained declared
consumer reads are outside the 30-day adoption window, so Era C correctly
reports `0/3`. Production preflight is 20/20 with zero deployment-environment
failures; the remaining launch-review blockers are time-bound SLO samples and
real-world adoption/KPI evidence.
The concrete production preflight is zero-failure on deployment
`dpl_2WH2w1hrM4zftM84kn3ouU3xu4N4`. The broader `review:goals:local` roll-up now
also reports zero deployment-environment blockers: aggregate launch-review and
Era C wrappers inherit real-world-evidence scope, while concrete preflight,
auth, smoke, IIIF, and k6 failures remain deployment-scoped when present. The
remaining 21 strict blockers are explicitly time-bound or human/external
evidence rather than deployment configuration failures.
August 7 PD0/PD3 checkpoint: the five public outcome events now have a typed,
consent-gated contract that excludes direct identifiers and professional
routes. First artwork completion, artwork continuation, 24-hour return,
rights-qualified download selection, and successful share are instrumented.
A dated baseline artifact and five-session non-specialist
protocol are present, while production observation and the five human sessions
remain open. PD3 has exactly one curated journey—Flowers across two centuries—
spanning Getty and Met records with citations, a high-confidence metadata
label, an explicit no-influence boundary, contract tests, and an editorial
decision packet awaiting external sign-off. That packet now represents the
optional externally validated tier: it separates editorial, rights, and subject-
matter decisions, and `pnpm connections:first-review` rejects direct identifiers,
synthetic evidence, duplicate reviewers, and incomplete approvals. Its absence
does not block independent research or operator-reviewed public experiments, but
the journey must disclose that external expert review and novelty validation are
absent. Before adding journeys two and three, implement the operator-release
record and public assurance-tier label, then keep claims inside the tier-2 boundary.
August 9 A+ quality reset: the prior Flowers, Van Gogh, and Waves concepts score
38, 40, and 43 under frozen rubric v1 and fail the triviality gate. They are no
longer eligible merely because their citations are reliable. PD3 now prioritizes
20 live-source candidates across at least three institutions, requires two
independent substantive signals plus knowledge unavailable from one source page,
and promotes only candidates reaching 95/100 with no dimension below 9/10.
The first deeper-discovery pass added a fail-closed 20-candidate portfolio gate
and expanded live capture from 8/8 to 14/14 jobs across four institutions. Exact
provenance-role detection and API/SPARQL/IIIF/Linked Art research surfaced Paul
Rosenberg and Wildenstein three-museum leads. Rosenberg now has typed actors, a
sourced person-gallery bridge, dated events, fresh receipts, and zero authority
contradictions; it remains unqualified pending real package observation,
human editorial review, and audience evidence. The automated visual-rights
preflight is fail-closed: one receipt-backed Getty image may display, while the
AIC surrogate awaits image-level terms and the copyrighted Braque is withheld;
accountable rights sign-off remains human.
The real-session handoff now includes a blinded five-session field guide and a
fail-closed importer; it records no direct identifiers and cannot create or
substitute participant evidence. Ties and continuation below 60% now block the
draft, preventing the team from overriding observed package preferences.
A no-index, client-only test surface now randomizes option order and downloads
candidate-bound responses without server collection; the CLI directly merges
those envelopes and rejects files from another candidate or schema version.
The consent step now validates both attestations at form submission rather than
depending on controlled-checkbox state, closing an interaction defect found in
live local-browser review. Static/type checks pass; that browser run remained
inconclusive because its React tree did not hydrate anywhere on the page.
Five synthetic specialist lanes then scored the unfinished Rosenberg work from
2/10 (comprehension evidence design) to 8/10 (non-obviousness/reproducibility).
The consolidated report is `docs/product/rosenberg-five-agent-specialist-review.md`.
Immediate changes removed unsupported promotional/causal wording, diversified
the frozen package frames, withheld the AIC surrogate pending image-level terms,
and added no-JS/download failure boundaries. Agent scores remain non-qualifying.
August 8 cultural-intelligence checkpoint: the first journey now produces three
synchronized representations from one typed contract: a public visual story, a
research dossier exposing fact/inference labels, method, uncertainties, rejected
hypotheses, novelty status, and rights boundary, plus an evaluation-only JSON
record with the same claim/citation graph and human-review state. Consent-gated
story-completion and reuse-interest signals are instrumented outside the five PD0
outcomes. Institutional usefulness, agent/editorial minutes, independent novelty
verification, five usability sessions, and attributable external approval remain
evidence needed for externally validated or commercially validated claims; they
do not block the independent operating lane.
The expansion gate is now executable through `pnpm connections:evidence`: its
privacy-safe artifact requires observed completion plus sharing/reuse, qualified
editorial sign-off, institutional usefulness, agent/editorial labor and cost,
independent novelty review, a real price response, and the valid synchronized
machine record. It reports tested gross value before labor separately from labor
minutes and cannot call that result profit. No real evidence has been imported,
so promotion to externally validated or commercially validated status remains
unauthorized. Independent research and operator-reviewed public experiments
remain authorized within their stated tier.
Internal hidden-pattern work can now proceed without violating that public gate:
`pnpm connections:patterns` emits a review-only collection-intelligence report
from normalized records plus explicit source rows. Deterministic candidates cover
equivalent-record conflicts, possible entity reconciliation, shared materials,
owner/custodian or set references, structured-provenance coverage gaps, and
geographic contrasts. Every lead carries citations, confidence, and a refusal
boundary; uncited records are rejected and unsupported demographic or market
conclusions are listed as refused analyses. No second public journey was added.
The review-only report now also consumes explicit Linked Art event evidence:
matching exhibition identifiers, `used_specific_object` groupings, structured
acquisition/transfer parts, dated event places, and shared activity actors. Its
event graph preserves source-record IDs on every edge. Geographic sequences are
movement candidates rather than transport claims; ownership histories do not
assert completeness, authenticity, custody, or legal title.
Additional explicit-evidence candidates now cover alternative maker assignments,
reversed event timespans, repeated technique identifiers, separate `represents`
and `about` iconographic concepts, and unidentified depicted people. Boundaries
prevent authorship resolution, invented corrected dates, workshop/influence
claims, collapsed depiction semantics, or demographic/underrepresentation
inference from these candidates.
The machine layer now also exposes `/api/cultural-intelligence` as a lifecycle-
aware collection feed. Each item preserves revision, production provenance,
review history, corrections, and separate editorial/licensing decisions. The
feed currently reports one evaluation item and zero licensable items; approval
metadata must be complete and all corrections resolved before eligibility can
change. Underlying source-record and media rights remain explicitly separate.
Five derivative formats now compile from the same versioned claim graph:
newsletter, daily feed, narrated visual essay, classroom package, and licensed
article. Evidence sections preserve claim/citation IDs and all formats repeat the
rights boundary. The derivative endpoint returns only an HTTP 409 release
manifest—not internal copy—while the first record lacks editorial and licensing
approval. Format availability is therefore implemented but audience demand,
quality, labor, accessibility, and price remain unproven external evidence.
Value-based pricing now has an executable evidence ladder through
`pnpm connections:pricing`: hypothesis, tested-no-signal, market signal, one
invoice-validated delivery, and repeatable price evidence. Repeatability requires
three scoped offers and two paid, accepted, value-confirmed, positive-contribution
deliveries across buyer segments. Labor is fully costed at an attributable rate;
interest is never revenue. No real offer artifact exists yet, so pricing remains
unvalidated.
The `pd0:evidence` intake command now validates a real GA4 export and moderated
session records, rejects direct identifiers or invented counts, and derives
completion only from five sessions plus attributable product-owner approval.
Its package namespace, script, artifact directory, and product-governance owner
are registered in the executable evidence-ownership and operations-risk controls.
Responsive browser review found the initial action block below the full image on
mobile; it now precedes the image and remains overflow-free at 390 px and 1440 px.
The Vercel ignore contract excludes local provider source mirrors, the `.tools`
binary cache, and the upstream `linked.art` checkout except its required schema
subtree. However, deployment `dpl_GiNFmNpo2UZs4uGEm4Y3B54Bya3b` still archived
79,077 files (669.1 MB), so CLI archive filtering remains an open packaging
optimization; the deploy itself completed and passed its runtime build.
P2 - Productize Only After The Pilot Loop Works (Later, 6-12 Weeks)
| ID | Outcome | Trigger | Exit criteria |
| ---- | --------------------------------------------------------- | ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| P2.1 | Guided organization onboarding and first-value dashboard. | One invoice-backed pilot completes activation and its friction is documented. | A new managed org can be provisioned with a sample or customer dataset in under 15 minutes; progress and first value are visible without engineering inspection. |
| P2.2 | Repeatable subscriptions and usage visibility. | Pricing, support load, and gross margin are validated on at least one pilot. | Checkout or invoice-backed subscription sync, webhook/audit evidence, quotas, usage, billing state, cancellation reason, and customer portal are supportable. |
| P2.3 | Institution procurement package. | A buyer starts security/legal review. | Deployment-specific subprocessors, DPA/legal artifacts, access review, incident drill, retention controls, backup/restore proof, status reporting, and SLA/SLO packet are buyer-reviewable. |
| P2.4 | Production agent bridge decision. | A named operator accepts the review workload and risk boundary. | AG2 bridge has explicit sign-off, eval evidence, rollback, auditability, and human-publication approval. A2A/AG-UI remain deferred. |
---
Readiness Scorecards
Launch Readiness
| Lane | Score | Decision | Next evidence |
| --------------------------------------- | ---------: | ------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| Internal development and portfolio demo | 9.0/10 | Safe to use and present with the strict-readiness caveat. | Reproduce the canonical local gate and resolve the public evidence inconsistencies. |
| Controlled public beta | 8.4/10 | Technically credible on Vercel + Neon, but the formal beta gate remains evidence-red. | Clear P0, then maintain narrow acceptance criteria while the 30-day window accumulates. |
| General public production | 6.5/10 | Do not claim complete readiness. | Passing long-window SLO/uptime, production KPI, and coherent launch evidence. |
| Institution-grade / strict 10/10 | 5.5-6.0/10 | Blocked by external and time-based proof. | `pnpm review:goals:check` passes with no production-proof blockers. |
SaaS Readiness
| Lane | Score | Decision | Next evidence |
| ------------------------- | -----: | ------------------------------------------------------------ | ---------------------------------------------------------------------------------------- |
| Technical SaaS foundation | 7.0/10 | Strong enough for concierge pilots. | Prove onboarding, support load, usage, and tenant operations with one buyer. |
| Paid pilot readiness | 8.2/10 | The offer and operator path are ready; revenue proof is not. | Reply or qualified follow-up, signed scope, invoice-backed entitlement, real activation. |
| Self-serve SaaS readiness | 3.0/10 | Deferred. | Start only after the pilot validates pricing and activation friction. |
| Profitable SaaS business | 4/10 | Credible wedge, but repeatable revenue is not proven yet. | Retention, support minutes, infrastructure cost, conversion, and gross-margin evidence. |
The primary wedge remains the Managed Linked Art Launch Pilot for small and
mid-size museums, archives, galleries, digital-humanities labs, and artist estates
that need standards-compliant collection publication without a semantic-web team.
Manual invoicing is correct for the first 1-3 pilots. Creator-side provenance and
self-serve billing stay deferred until the B2B pilot loop produces evidence.
---
Evidence Workstreams
| Workstream | Current state | Completion condition |
| -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Local quality | Review-goals local status passes and direct ESLint passes; full canonical reproducibility was not demonstrated in the July 12 audit. | P0.2 is green from a clean generated state. |
| Deployment | Preflight `20/20`, both Render probes, all public smokes, and deployed k6 pass; launch review is `7/8`. The strict handoff has zero deployment-environment blockers because its remaining launch and Era C wrappers depend only on real-world evidence. | Preserve the green concrete deployment matrix while the real-world evidence windows mature. |
| Long-window SLO and uptime | The latest strict handoff reports `11` retained deployed SLO samples across `10/30` distinct UTC days, with `11` passing samples and no failed or incomplete rows in the active report window; the Era C artifact still reports only `15/30` samples toward its exit gate. | Continue distinct-day collection until the complete 30-day threshold is genuinely met. |
| ActivityStreams | Operator endpoint probes pass for three named IDs, but the retained real external consumer evidence is stale and therefore counts as `0/3`; genuine `Delete` evidence is pending. | Collect three fresh real external consumers covering `Create`, `Update`, and `Delete`, with verified callbacks. |
| Production KPI | Local enrichment is promising; production reconciliation distribution and reviewed precision are incomplete. | P1.2 passes the SOTA KPI acceptance rows from named production sources. |
| Managed pilot | The offer, runbook, entitlement, activation, support, and evidence tooling exist. The latest no-pricing buyer pack is specific to the recorded Te Papa outreach and has a real account, organization, and owner, but no paid-pilot tenant or invoice-backed entitlement exists. | Obtain a real buyer reply plus signed scope or invoice reference, provision the tenant, then use P1.3 tooling to record activation, retention, support load, and margin evidence. |
The buyer-review surface now includes a standalone ten-record demonstration at
`/museum-linked-art-pilot-demonstration.html`. It uses traceable public API records
to show source preservation, event-centric Linked Art JSON-LD, rights review
boundaries, validation findings, and museum questions. It proves a review pattern,
not a completed customer engagement or permission to reuse source images.
A one-page buyer brief at `/managed-linked-art-pilot-brief.html` now packages the
problem, five-day process, required inputs, deliverables, privacy and security
boundaries, and post-pilot decision into a printable pre-call handout. It links to
the ten-record demonstration and preserves the same evaluation-only claim boundary.
The first-call workflow now has a timed guide at
`/museum-pilot-discovery-call-guide.html`: a one-minute permission-based opening,
five fit questions, a boundary recap, three explicit decision paths, and a
follow-up record. The call qualifies a bounded pilot before any product tour and
links directly to the buyer brief and demonstration when supporting proof is useful.
The post-call handoff now has a public-data request at
`/museum-pilot-data-request-template.html`. It supplies a copy-ready museum message,
accepts CSV, JSON, XML, LIDO, or a public API, distinguishes minimum from optional
fields, excludes credentials and restricted material, and records the reviewer,
publication boundary, transfer method, receipt evidence, and agreed deletion date.
The conversion and delivery packet now adds a counsel-review sample agreement,
four-level introductory pricing, and a reusable results report. The public pilot
offer uses the same `$0` evaluation, `$3,500` fixed paid pilot, implementation from
`$12,000`, and ongoing service from `$1,250` monthly hypothesis, so buyer surfaces
no longer conflict. These remain unvalidated prices until invoice-backed delivery,
acceptance, retention, and gross-margin evidence exists.
The refined demonstration presents one primary path on desktop and mobile:
collection record, Linked Art mapping, validation, then reviewable result. Source
evidence is collapsed beneath the interaction, and the final state names open
museum decisions and the human publication gate.
The current outreach ledger records the eight user-confirmed August 5 submissions,
their real recipient or form channel, zero assumed replies, and August 12 follow-up
dates. `docs/sales/museum-outreach-pipeline.md` contains eight unsent follow-up
drafts plus a second official-source-researched group of eight prospects that must
remain `research_only` until first-round feedback is reviewed and sending is
authorized. `docs/ops/paid-pilot-commercial-evidence-process.md` closes the
invoice-to-margin capture design without treating placeholders as proof.
Standing Evidence Controls
`generatedAt` separate from source `sourceUpdatedAt` and `sourceUpdatedDoc`
checksum metadata.
`wikidataexplorer-metamuseum-prod` and the other real consumers distinct,
retain durable callback evidence, and preserve the zero rejected subscriptions
state without allowing placeholders to satisfy strict proof.
`pnpm longterm:evidence:public` output remain strict gates; a frontend Core Web
Vitals baseline is added in P1.5 rather than inferred from API SLO evidence.
remain separate scopes. No aggregate badge may silently promote one scope into
another.
- Public docs metadata freshness: `/api/docs/manifest` must keep response
- ActivityStreams onboarding ledger: partner rows must keep
- Performance evidence: the cold-record budget, 30-day SLO depth, and
- Claim boundary: local success, deployment success, and real-world success
---
Product And Engineering Guardrails
- Linked Art JSON-LD remains canonical; UI DTOs are projections at boundaries.
- Preserve rights, source attribution, provenance, multi-value arrays, event
semantics, carrier/content/surrogate separation, and opaque URI handling.
- Adapters do not import each other; provider parsing stays in adapters;
cross-provider mapping stays in `src/utils/artwork-builder.ts`; contracts remain
leaf modules.
- AIDD + TDD remains mandatory for behavior changes. Standards-critical work
cites reference rounds and fixture anchors before implementation.
- Public publication and agent-generated claims require citations, refusal paths,
audit evidence, and human approval.
- Cultural-intelligence candidates use an executable deterministic routing policy:
weak candidates archive without an agent, strong ambiguity permits at most one
bounded evidence-packet pass, detector reports derive their own scoring inputs,
and every traceable review-queue publication route remains human-gated.
- Probabilistic cultural-intelligence signals are versioned, thresholded,
fixture-calibrated, source-backed review candidates with zero LLM calls; they
never establish identity, influence, movement, meaning, or historical truth.
- Cultural-intelligence ingestion validates source envelopes and Linked Art,
retains immutable hashed snapshots, normalizes only a comparison projection,
and recognizes exact authority IDs without network or LLM calls.
- Review-ledger runs are append-only and attributable; only fully approved
entries generate synchronized JSON-LD, API, dossier, timeline, and accessible
page artifacts, all still requiring an operator to publish.
- Scheduled cultural capture evaluates explicit cadences and enforces HTTPS
allowlists, redirect/media/size bounds, hashed receipts, and separate
replay/live evidence before ingestion.
- The integrated production-like replay proves four of five candidates route
without an LLM (80%), one bounded pass is allowed but not executed, all nine
routine stages are deterministic, and external outcome evidence remains null.
- Cultural graphs preserve cited entity/activity/concept/equivalence edges and
their assertion certainty; contradictory rights become review candidates,
while unknown rights remain a separate reuse blocker.
- Completion readiness is machine-audited: technical controls and fixture proof
remain distinct from attributable live captures, human decisions, production
outcomes, observed costs, institutional usefulness, and commercial signals.
- Provenance rules detect broken dated transfer chains and ownership before
production; multiple-agent review is executable only for consequential
conflicts with measured value above cost and two independent typed results.
- Keep Next.js, React, TypeScript, custom CSS, Postgres/JSONB, Solr, GraphDB, and
canonical ID decisions locked as documented in CLAUDE.md(../CLAUDE.md).
- No new runtime dependency, provider, service, database, or architecture era is
started while P0 is red without explicit approval.
Deliberately Deferred
economics and activation are real.
by measured scale or customer evidence.
- New provider integrations beyond the current 14 production lanes.
- Self-serve signup, checkout, billing portal, and growth automation before pilot
- Synthetic ActivityStreams `Delete` evidence.
- Broad production agent autonomy or public publishing without operator sign-off.
- A microservice, triple-store, vector-store, or framework expansion not justified
Active 30-day revenue experiment
The supporter, direct-sponsor, and institutional-pilot funnel is implemented
locally with transparent public offers, a one-page sponsor packet, consent-aware
revision-bound/deduplicated browser events, five first-party-sourced sponsor
candidates, an exact unsent message, human approval/send-evidence gates, and a
fail-closed revenue-report command. It does not yet
claim a live checkout, approved outreach, inquiries, contracts, invoices,
payments, accepted delivery, or revenue. The next gate is accountable human
approval of the production revision, introductory pricing, target list, and
outbound copy, followed by a predeclared exact 30-day window and honest outcome
report. See ops/30-day-revenue-experiment.md(ops/30-day-revenue-experiment.md).
The complete local WCAG browser audit passes, while the public-trust visual
smoke retains one expected privacy-baseline failure for human review. Direct
production probes currently return 404 for support, sponsor, and sponsor packet,
and show the previous pilot copy; release and the 30-day clock therefore remain
unstarted. A declaration command now prevents the clock from starting without a
released git SHA, production analytics identity, exact dates, and human approval.
Rosenberg primary-document checkpoint
The MoMA V.A.8 returned-paintings folder has now received a complete bounded
review, not merely a finding-aid citation. Pages 2-65 were inspected; the
apparently promising Braque sequence on pages 4-8 was disconfirmed by its own
notarized declaration as _L'intérieur au vase noir_. High-resolution review of
the grouped catalogue sheets likewise produced no secure photograph-3492 or
target-title match. This useful negative is retained with a file hash and rights
boundary in `rosenberg-primary-document-receipts-v1.json`. It narrows the next
research action to photograph 3492, RA1, or a handover receipt and prevents the
project from presenting V.A.8 as object-specific evidence. The case still awaits
independent human review and does not claim scholarly novelty.
Research-outcome instrumentation checkpoint
The Rosenberg report now has consent-gated, revision-bound measurement for
report view, actual historical-trail 25/50/90% depth, successful share,
browser-local save, “learned something new,” and study start. Actions are
session-deduplicated, carry no direct identifiers, and saving remains available
without analytics consent. Desktop interaction and a 375 px runtime check pass
with no horizontal overflow. Production probes remain decisive: the report,
reader study, support, and sponsor routes return 404 after locale routing; only
the older pilot page returns 200. Therefore current audience, comprehension,
support, and sponsor outcomes remain zero/unobserved rather than inferred from
local instrumentation. An accountable approved release is the next dependency.
Release-decision checkpoint: `value-release-candidate-v1.json` consolidates the
exact dossier, receipt, protocol, report, study, engagement, support, sponsor,
pilot, and production-probe hashes. Tests recompute every digest and the
candidate fails closed with eight incomplete human attestations, no operator
code, no approved commit, and `releaseEligible: false`. The companion
`product/value-release-decision.md` gives the operator one bounded decision and
makes clear that approval does not authorize novelty, legal, image-rights,
outreach, audience, or revenue claims. No further local feature is required to
begin the experiment; genuine approval and deployment are now the limiting
inputs.
---
Cadence And Ownership
Research Commons organic acquisition — local release candidate complete
The private AI question-collection approval workspace now converts the first evidence-operations blocker into an exact, attributable human decision. Four explicit approvals, a public attributable HTTPS evidence record, a pseudonymous accountable code, and a substantive note are canonically SHA-256 signed in-browser and downloaded locally. Browser validation now reuses the authoritative public-HTTPS predicate and rejects local, private, and example/test/sandbox/demo/fixture hosts before download; the question-publication and open-release deployment workspaces share the same rule. The workspace has no persistence or activation authority; approval verification, deployment, and runtime enablement remain separate. This improves operator usability but earns no external A+ point by itself.
The acquisition workstream now also exposes a private exact-contract editorial workspace covering all six question-led guides, their claims, sources, limitations, FAQ data, indexing effects, and excluded claims. Its downloaded approval is canonically signed and mutation-sensitive, while publication, indexing, sitemap inclusion, campaign activation, and deployment remain separate human actions. It removes review friction without counting editorial approval as traffic, validation, or income.
Both approval workspaces now avoid approval-only choice architecture: an accountable reviewer can sign a rejection or changes-required decision with all approval dimensions false. Rejections remain integrity-verifiable while every activation and publication checker stays fail-closed.
The external-validation workstream now has a device-local JSON preflight for returned validation and evidence-envelope files. It checks the exact pseudonymous three-researcher/one-expert mapping, facilitator verification, consent, independence, synthetic exclusions, direct-identifier keys, receipt binding, corrected-v2 packet, novelty decision, and published limitations. The privacy-safe receipt excludes file content, and the authoritative CLI repeats the shared inspection before frozen-study projection.
The conventional-baseline manifest now closes the post-outcome-freeze loophole: it requires an accountable `REG-…` preregistration code, attributable HTTPS record, SHA-256 receipt, and a preregistration timestamp that predates all observed sessions. Digest-valid paired results cannot pass if the study was only frozen after outcomes were visible.
The production AI evaluator now rejects any bundle whose dataset freeze or independent label approval does not strictly predate every production trace. This prevents expected claims and refusal labels from being retrofitted after observing model behavior, while retaining the existing 50-case, 10-answer/10-refusal, two-reviewer, citation, latency, and cost gates.
The production AI evidence chain now shares the strict public-HTTPS validator used
by other A+ lanes. Question-export attestations, manifest preparation, frozen dataset
sources, label approvals, run receipts, citations, and independent reviews reject
localhost, private/link-local addresses, `.local`/`.invalid`, and
example/test/sandbox/demo/fixture hosts. Focused service and subprocess fixtures use
non-fixture public-shaped hosts, while explicit regression cases prove that rehashed
non-public evidence cannot advance assembly or scoring.
External impact now derives the reviewed-release timestamp from the signed validation artifact and requires every citation, substantive reuse, or accepted correction to occur strictly afterward and no later than evaluation time. Events can no longer be retroactively attached to a release that did not yet exist or projected from future timestamps.
Six question-led research methods, canonical Article/FAQ metadata, internal
qualification paths, aggregate cockpit events, and a digest-bound publication
gate are implemented locally. The routes fail closed as `noindex,nofollow` and
remain absent from the sitemap. The next milestone is independent editorial and
research review of the exact release; only then may a human approve indexing.
Production impressions, qualified researcher participation, external reuse,
settled income, and the remaining A+ requirements remain unobserved.
Consented organic campaign — local controls complete, production proof pending
The exact double-opt-in audience, confirmation purpose, five-message copy,
privacy boundary, and aggregate measurement contract now produce a digest-bound
readiness artifact. Durable failed-delivery receipts stop automatic retry, and
receipt-backed permanent bounces or complaints can suppress by hash without raw
recipient data; temporary failures stay held for operator review. The gate still lacks six production delivery receipts and
attributable human approval, so lead capture and sending remain unauthorized.
Independent validation projection — local controls complete, external study pending
The Rosenberg comparison can now project into the A+ investigation, reviewer,
and frozen-baseline fields only after the underlying study passes and exact
external evidence proves three facilitator-verified researchers, one independent
subject expert, a distinct materially corrected v2 packet, a novelty decision,
published limitations, and unique receipt hashes. Direct identifiers,
self-reported-only qualifications, unchanged dossier digests, and synthetic
substitutes are rejected. The current readiness artifact has four blockers;
there are still no genuine returned sessions or expert decision to credit.
The artifact now has a canonical integrity digest and automatically supplies
the corrected investigation plus reviewer set to the A+ collector only when it
is ready and untampered. Measured improvement still comes exclusively from the
separate row-level conventional-baseline importer, preventing a weaker
projection from overriding that evidence lane.
Production AI evaluation — evaluator complete, production evidence pending
A separate A+ evaluator now refuses to treat the 120-prompt local golden run as
production proof. It requires at least 50 unique fresh production traces bound
to a pre-output, independently approved label manifest, with at least ten
answerable and ten refusal cases, two independent reviewers, 25 reviewed
citations, one model/prompt revision, provider or instrumented cost, and unique
trace/review receipts. Precision, recall, citation and refusal accuracy, p95
latency, and mean cost are calculated from rows. The current artifact is blocked
on the absent frozen manifest and production trace bundle.
| Cadence | Required action |
| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| Every behavior change | Red-green-refactor tests, focused smoke/evidence, README + roadmap update, and `pnpm session:closeout`. |
| Every 72 hours during active shipping | Canonical local gate plus fresh P0 evidence checks. Stop expansion when a required gate is missing. |
| Weekly | Refresh long-window evidence, inspect failed-sample age-out dates, review pilot pipeline, and update only changed roadmap decisions. |
| Monthly or before a buyer review | Refresh production preflight, launch evidence/review, procurement packet, access/DR evidence, and the strict handoff. |
The roadmap records current decisions and measurable outcomes, not every merged
change. Completed implementation detail belongs in
progress/era-history.md(progress/era-history.md), specialized docs, generated
artifacts, and git history.
Research release quality checkpoint (2026-08-14)
The A+ program now has a single-digest technical release gate covering tests,
lint, build, dependency security, docs validation, and diff hygiene. The point
remains unearned until a fresh successful artifact exists against the current
canonical Research Commons manifest; this does not relax any external-evidence
or production-income requirement.
The release receipt is now semantically replayable rather than outer-hash-only. Its
verifier recomputes the digest from the exact canonical file set, requires the fixed
six commands, validates command chronology and exact durations, and reconstructs
blockers/status. Rehashed blocked-to-passed, command-substitution, and manifest-
omission edits fail before scoring.
The initial gate exposed high-severity advisory 1139346 in Lighthouse's
`extract-zip` dependency. The redundant three-route Lighthouse check was
replaced by the retained 23-route axe/Playwright WCAG A/AA gate, and patched
transitive versions of PostCSS, Sharp, OpenTelemetry Core, and UUID are pinned.
The dependency audit now reports zero advisories without a baseline waiver. The
canonical release manifest includes the package manifest, lockfile, CI,
accessibility command, security command, and security baseline.
The repository and CI now share pinned pnpm 10.34.5 resolution semantics, so
workspace security overrides cannot be silently ignored by an older client.
Organic researcher contribution checkpoint (2026-08-14)
An indexable `/research/contribute` hub now gives researchers and experts a
single canonical path into packet reproduction, material correction,
counterbalanced workflow comparison, and method challenge. It requires no
account, uploads no response, exposes `ResearchProject` structured data, and is
linked from the sitemap, footer, and Research Commons. Controlled study tools
remain `noindex`, while the contribution guide makes privacy, custody,
qualification, correction, and human-decision boundaries explicit. This
improves discoverability but earns no external-validation point until genuine
receipt-bound evidence is returned and independently qualified.
Machine-readable Research Commons checkpoint (2026-08-14)
The indexable Commons now advertises a public, digest-protected JSON-LD catalog
and directly fetchable packet JSON Schema whose declared `$id` resolves. The catalog enumerates only public
methodology, contribution, correction, review, and schema resources plus
no-account actions; it deliberately excludes draft questions, unapproved
findings, private packets, and participant data. Both resources have explicit
media types, public caching, CORS, sitemap entries, and visible links.
Technical recommendation packet: the change adds two static research routes
and one catalog service without changing evidence or publication decisions.
Ranked risks are (1) leaking draft research through discovery, controlled by an
allowlist catalog and exclusion tests; (2) catalog drift, controlled by its
canonical digest and release manifest; and (3) schema divergence, controlled by
serving the exact release-bound packet schema. Next actions are to validate
external consumption (research operations; acceptance: attributable fetch or
reuse receipt), monitor crawler discovery (acquisition operator; acceptance:
fresh provider export), and keep catalog entries tied to publication state
(research lead; acceptance: no unapproved resource in the catalog). Validate
with `node --import tsx --test tests/services/research-commons-catalog.test.ts
tests/pages/research-commons-page.test.ts`, `pnpm research:commons:browser-proof`,
and `pnpm research:release-quality`. The next-cycle hypothesis is that an
external researcher can discover the packet contract without traversing a
private API; falsify with a clean-session fetch of both advertised alternates.
The catalog now also exposes a deterministic, schema-complete synthetic sample
packet. It uses a fictional object and non-executing provider, carries explicit
permanent exclusions from every qualifying evidence lane, and detects any
post-generation mutation through the same canonical digest contract used by
device-local packets. This closes the reproducible-sample requirement without
manufacturing a research result.
Technical recommendation packet: ownership remains within the packet and public
catalog services. Ranked risks are (1) a sample being mistaken for evidence,
controlled by permanent machine-readable exclusions and fictional content; (2)
silent mutation, controlled by canonical SHA-256 verification; and (3) sample
drift from the public schema, controlled by required-field and release-manifest
tests. Next actions are external clean-session reproduction (research
operations; acceptance: independently recomputed matching digest), schema-client
import (developer advocate; acceptance: generated client accepts the packet),
and catalog discovery measurement (acquisition operator; acceptance: fresh
attributable provider receipt). Validate with `node --import tsx --test
tests/services/research-commons-sample.test.ts
tests/services/research-commons-catalog.test.ts`,
`pnpm research:commons:browser-proof`, and `pnpm research:release-quality`. The
next-cycle hypothesis is that a third party can reproduce the digest without
project code; falsify by following only the public sample instructions.
Research correction propagation checkpoint (2026-08-14)
Corrections now use a public version-1 envelope schema, two independent digests,
raw-file receipts, allowlisted field targets, pseudonymous human submission,
HTTPS evidence, and explicit consent. `research:correction-propagate` preserves
the original packet and creates a distinct candidate revision with the proposed
change and an append-only pending decision. It rejects synthetic samples,
direct identifiers, digest drift, unsupported targets, and conclusive claims;
it cannot accept, publish, qualify, contact, or convert a proposal into impact.
Technical recommendation packet: ownership sits in the correction-propagation
service and file-receipt command. Ranked risks are (1) silent historical
overwrite, controlled by retaining the exact original and a previous-digest
link; (2) automated acceptance, controlled by a fixed pending human decision;
and (3) correction laundering into impact evidence, controlled by synthetic,
qualification, release, and acceptance separation. Next actions are independent
correction review (subject expert; acceptance: reproduced evidence and explicit
decision), released-revision binding (research lead; acceptance: distinct
approved digest), and impact import only after external acceptance (evidence
operator; acceptance: attributable receipt). Validate with `node --import tsx
--test tests/services/research-correction-propagation.test.ts`, `pnpm
research:correction-propagate`, and `pnpm research:release-quality`. The
next-cycle hypothesis is that any one-field correction can be propagated without
mutating its source; falsify by comparing the retained source bytes and digest.
The indexable contribution hub now contains a device-local correction-envelope
builder that shares the operator validator's browser-safe core. It requests no
name, email, account, upload, or artwork file; validates one allowlisted field,
HTTPS evidence, pseudonymous contributor code, human submission, consent, and
AI disclosure; computes the canonical digest with Web Crypto; and downloads the
file under researcher custody. Desktop and mobile Chromium prove the complete
download path, accessibility, and no horizontal overflow. That proof also
exposed and fixed a pre-existing nested-main landmark defect on the contribution
page.
Technical recommendation packet: ownership is split cleanly between the pure
envelope core, client builder, and Node-only propagation finalizer. Ranked risks
are (1) browser/server validation drift, controlled by the shared pure core; (2)
unnecessary personal-data capture, controlled by the field allowlist and direct
identifier rejection; and (3) proposal/acceptance confusion, controlled by
pending-only wording and disabled publication. Next actions are genuine expert
use (research operations; acceptance: independently returned envelope),
accountable review (subject expert; acceptance: explicit accept/reject receipt),
and candidate propagation (operator; acceptance: distinct retained digest).
Validate with `node --import tsx --test
tests/services/research-correction-envelope.test.ts
tests/services/research-correction-propagation.test.ts`, `pnpm
research:commons:browser-proof`, and `pnpm research:release-quality`. The
next-cycle hypothesis is that an expert can return a valid correction without
sharing contact or artwork data; falsify through a consented usability session.
Release diagnostic observability checkpoint (2026-08-14)
The canonical release gate no longer reduces an intermittent test, lint, or
build failure to an opaque exit code. Failed receipts retain at most five
prefix-filtered, redacted, individually hashed diagnostic lines plus an exact
gate-specific next action. Successful output remains digest-only. The gate does
not retry automatically, and its test instruction states that a passing retry
does not prove a flaky failure resolved.
Technical recommendation packet: ownership remains in the release-quality
service and runner. Ranked risks are (1) leaking command output, controlled by
prefix selection, redaction, length/count limits, and pass-output exclusion;
(2) hiding intermittent failures through retries, controlled by no automatic
retry and explicit operator wording; and (3) tampered diagnostics, controlled by
per-line and artifact digests. Next actions are to use the serial spec reporter
on recurrence (quality owner; acceptance: named failing test), repair or
formally quarantine its root cause (module owner; acceptance: repeated clean
canonical runs), and retain failure/pass history (release operator; acceptance:
both receipts remain attributable). Validate with `node --import tsx --test
tests/services/research-release-quality.test.ts`, `pnpm lint`, and `pnpm
research:release-quality`. The next-cycle hypothesis is that the next command
failure yields a safe named diagnostic; falsify with the unit-test error fixture.
Conventional-baseline integrity checkpoint (2026-08-14)
Measured improvement can no longer enter A+ scoring as unlinked aggregate
numbers. A dedicated importer now freezes cases, workflows, protocol and
thresholds before observation; derives results from unique paired session rows;
requires both counterbalanced orders and receipt hashes; and cross-links each
participant to the qualified-researcher evidence set. No genuine bundle is
present, so the external requirement remains unearned.
Production AI evidence assembly checkpoint (2026-08-14)
The production evaluator now has a privacy-safe staging path. Frozen labels,
production observations, and independent citation reviews remain separate until
an exact receipt-bound join succeeds. Retained rows exclude prompts and output
prose, include observed latency and cost receipts, and prevent the run operator
from self-review. No production observations exist yet, so the metric remains
unearned and no model execution is authorized.
The operator path now includes `research:ai-production-run-kit`, which converts
only a valid independently approved 50-case manifest into case-complete pending
observation and balanced two-lane review worksheets. It retains no prompts or
outputs and invents no human reviewer data. The run kit, assembler, and final
evaluator now share a canonical parsed-manifest digest, eliminating an
interoperability defect where harmless JSON formatting could prevent a genuine
bundle from aligning. The default configuration remains blocked because no real
approved manifest exists.
External-impact integrity checkpoint (2026-08-14)
External impact can no longer pass from aggregate counts. A dedicated importer
requires unique attributable public events, independent substantive
verification receipts, and an exact match to the reviewed investigation ID,
version, and packet digest. Meta Museum-owned and synthetic hosts are excluded.
No genuine external citation, reuse, or accepted correction is retained, so the
requirement remains unearned.
Research acquisition integrity checkpoint (2026-08-14)
A+ acquisition now requires one receipt-bound join across provider SEO, GA4
funnel, exact approved campaign controls, and pseudonymous human-verified
qualified actions tied to the current Research Commons release. Clicks, opens,
downloads, signups, and AI activity remain funnel signals only. No genuine
joined provider and qualified-action evidence is retained, so the requirement
remains unearned.
The contribution hub now closes the aggregate measurement gap with four
consent-dependent events for reproduction, correction, workflow comparison,
and method challenge. Normalized GA4 import and the operator cockpit preserve
each count and the aggregate contribution-path total. A+ acquisition now
requires at least one such fresh path signal alongside the separate
human-verified qualified-action ledger, so click activity remains explicitly
ineligible as scholarly validation.
Unified research evidence operations checkpoint (2026-08-14)
`research:evidence-operations` now reads the A+ scorecard and all six retained
external-evidence lanes, verifies available artifact integrity, calculates
requirement-specific freshness, and emits deterministic 30/60/90-day windows.
Seven missing criteria are grouped into six exact command workstreams because
external investigation review and reviewer qualification share one validation
bundle. Missing, blocked, invalid, stale, and ready-but-unaccepted states remain
distinct, and every action explicitly denies external side-effect authority.
The operations artifact now retains the readiness requirements, lane snapshots, and
stage-specific command overrides and semantically reconstructs freshness, reporting
windows, grouped workstreams, priorities, invocation contracts, alerts, and authority.
Rehashed command, priority, or authority substitutions fail before cockpit display.
The four older fallback writers now use the shared canonical integrity
finalizer, eliminating false `invalid` alerts while keeping every absent-input
lane blocked and ineligible.
The production-AI workstream now resolves its command from verified stage
evidence. It now begins by converting a receipt-bound 50-question export into
an integrity-protected, text-free labeling worksheet with blank labels and no
approval. A valid candidate advances to the separately human-approved run-kit
stage; a valid ready run kit advances to evidence assembly; and a valid ready
assembly advances to production evaluation. Missing, blocked, or digest-invalid
stages return the operator to the earliest safe command.
Technical recommendation packet: the behavior change is confined to evidence
operations and preserves every external-action boundary. Ranked risks are (1)
retaining sensitive question text, controlled by hash-only output; (2)
advancing from a modified intermediate artifact, controlled by canonical
verification; and (3) mistaking preparation for approval or qualifying
evidence, controlled by null labels and the unchanged A+ evaluator. Next actions
are to prepare a genuine 50-question candidate export (research operator;
acceptance: verified text-free worksheet), obtain independent label approval
(research lead; acceptance: receipt-bound frozen manifest), assemble
receipt-bound observations and reviews (evaluation operator; acceptance: ready
verified assembly), and run the unchanged production scorer
(research lead; acceptance: row-derived passing metrics). Validate with
`node --import tsx --test tests/services/research-evidence-operations.test.ts`,
`pnpm research:evidence-operations`, and `pnpm research:release-quality`. The
next-cycle hypothesis is that a completed verified run kit changes the first
command to evidence assembly; falsify by inserting a digest mismatch and
confirming the queue returns to run-kit preparation.
Evidence command-contract drift checkpoint (2026-08-14)
Evidence operations now have an executable maintenance contract across all nine
stage and external-evidence commands. The focused test resolves each documented
command through `package.json`, verifies its exact implementation target,
confirms that source parses every machine-readable input-contract flag, and
requires the implementation in the canonical research-release manifest. This
turns documentation/CLI drift into a release failure instead of an operator
surprise. The focused suite passes 6/6; this adds no external action authority
and does not change the evidence-backed A+ score.
Private evidence-queue cockpit checkpoint (2026-08-14)
The editor-only organic-income cockpit now consumes the signed research-evidence
operations artifact instead of leaving the prioritized queue accessible only as
generated JSON/Markdown. Authorized operators see each safe invocation plus
its structured private/non-private and human-attested/machine-derived inputs.
Missing, digest-modified, and malformed artifacts expose zero commands and a
clear fail-closed explanation. The page, loader, and their tests are now bound
into the canonical research release. Focused page, loader, and release-manifest
tests pass 5/5; no model, recruitment, publication, email, or payment authority
is introduced.
Production-question export contract checkpoint (2026-08-14)
The first production-evaluation handoff no longer trusts a TypeScript cast over
operator JSON. A release-bound Draft 2020-12 schema and matching runtime
validator enforce the exact root and case fields, reject additional or direct-
identifier fields before creating an output directory, and use row-number-only
diagnostics so malformed identifiers and private values are never reflected.
The CLI still hashes valid questions into blank, unapproved labeling rows and
performs no model call. Service, subprocess, schema, non-reflection, and release
mutation coverage passes 9/9 focused tests.
Canonical production-question schema discovery checkpoint (2026-08-14)
The production-question schema no longer claims an unserved, noncanonical URL.
Its `$id` now resolves through `/schemas/production-question-export/v1` on the
canonical `www` origin with `application/schema+json`, public CORS, and bounded
cache controls. The Research Commons page, signed JSON-LD catalog, and sitemap
all discover the same route, and the route/schema are release-bound. Ten focused
route, catalog, page, validator, and manifest tests pass. Only the data-free
contract is public; production questions remain private.
Canonical-host equality is asserted from the served schema response rather than
inferred from filenames.
Attested production-query transformer checkpoint (2026-08-14)
The first genuine-data handoff can now be produced from an explicitly exported
Meta Museum `ai-query-log.json` without hand-editing or implicit live-storage
access. `research:ai-question-export` requires a separate human production
attestation bound to the exact raw digest, accepts only fresh successful rows,
deduplicates normalized questions, rejects contact data, pseudonymizes raw query
IDs, and writes the private strict-schema export only at 50 eligible cases.
Aggregate-only stdout and no-output-on-blocked behavior keep questions out of
operator logs. Eleven focused service, subprocess, schema, and release tests
pass; the transformer cannot establish production genuineness or approve labels.
Production-question aggregate readiness checkpoint (2026-08-14)
`research:ai-question-attestation` now prepares a digest-bound, text-free
candidate before any human production claim. It shares the transformer's exact
freshness, success, length, contact-data, and deduplication analysis; retains
only aggregate exclusion counts plus one privacy-safe outcome code per row; and leaves environment, source, operator, and
collection approval null. The consent-hardened current local log yields zero
eligible rows from the legacy input because none satisfies the exact current export
shape and consent boundary. The candidate remains
blocked and makes no production claim. The current replay found 2,497 legacy rows,
all malformed for the exact current evaluation-export contract and therefore zero
eligible. Evidence operations verify its digest but correctly keep consented
collection approval ahead of this downstream prerequisite. Candidate, subprocess,
resolver, command-contract, and manifest-chain tests pass.
Verification reconstructs counts, blockers, readiness, and the blank attestation
without retaining questions, raw IDs, timestamps, or contact values. Rehashed count,
blocker, or blocked-to-ready substitutions fail.
Consent-bound AI evaluation logging checkpoint (2026-08-14)
The AI-query request contract now has an optional strict evaluation envelope:
`consent: true`, consent version `research-ai-evaluation-v1`, explicit human
submission, and an allowlisted collection lane. Partial, false, or extended
claims fail request validation. Ordinary future queries retain only a question
SHA-256 and null text; exact text is retained only for the consented lane. The
aggregate readiness analyzer excludes every ordinary or legacy row as
`unconsented`, while the separate exact-digest human production attestation
remains mandatory. API, logger, analyzer, transformer, and operations coverage
passes 25/25 focused tests, and these trusted inputs are release-bound.
Approval-gated AI question surface checkpoint (2026-08-14)
`/research/evaluate-ai` now provides the technically complete path for future
diverse consented questions without requesting identity, but remains no-index
and inactive. A stable contract digest covers the exact consent copy, purpose,
data fields, exclusions, and 35-day retention. Activation requires both an
attributable checked-in approval of every dimension and a separate runtime
switch; either alone fails closed. The inactive page renders no form, and the
API independently returns 403 for crafted collection requests. Contact-bearing
questions fail before execution, ordinary logs are hash-only, and the writer
nulls expired consented text while preserving receipts. Twenty-six focused
page, gate, API, logger, analyzer, transformer, and release tests pass. No
approval, deployment, publication, recruitment, or collection occurred.
Research browser-proof scoring checkpoint (2026-08-14)
A+ readiness now verifies the three required Research Commons journeys by
exact title across both desktop and mobile Chromium, while also requiring the
report totals to reconcile and unexpected/flaky counts to remain zero. This
replaces the obsolete assumption that a healthy report always contains exactly
four executions. The score therefore stays strict when required coverage is
lost and remains stable when additive browser coverage is introduced. Focused
regression tests prove both cases; the current artifact-backed readiness score is 3/11,
with all eight remaining requirements dependent on attributable external
production, validation, acquisition, impact, or income evidence.
A+ scorer trust-boundary checkpoint (2026-08-14)
The canonical research release now binds the A+ evaluator implementation,
operator CLI, blank evidence configuration, service and CLI tests, plus a
manifest regression test. The regression test mutates each trusted scoring
input independently and proves the release digest changes. This prevents an
easier threshold, altered requirement, or changed evidence loader from being
applied under an older release identity. As designed, expanding this boundary
temporarily invalidates the prior release-quality receipt until the canonical
gate regenerates evidence for the new digest.
Production-question intake usability checkpoint (2026-08-14)
The first external-evidence handoff now accepts a genuine production-question
export directly through `--input`, computes its exact raw-file receipt and a
truthful preparation timestamp, and emits only the integrity artifact plus a
question-hash-only blank labeling worksheet. Controlled runs may override the
timestamp and output directory, while configuration mode remains available for
scheduled operation. An end-to-end subprocess test proves that 50 cases become
a ready worksheet and that neither stdout nor either generated artifact
contains the retained question text. No model call, label inference, approval,
contact, or publication authority is added.
The private status now retains a hash/length-only replay projection with valid
pseudonymous case IDs and redacts invalid IDs. Its verifier reconstructs source
eligibility, uniqueness, normalized-length gates, blockers, candidate count, status,
and every blank labeling row; rehashed readiness or worksheet edits fail.
Approved-manifest run-kit usability checkpoint (2026-08-14)
The second production-evaluation handoff now accepts an independently approved
manifest and pseudonymous run-operator code directly through CLI arguments. It
derives the canonical dataset identity and truthful preparation timestamp,
rejects the label approver as run operator, and writes only blank case-complete
observation and balanced independent-review worksheets. A subprocess test
proves the separately exported worksheets omit question hashes, approval references,
and the approver code. The private status artifact retains the exact approved
manifest and reconstructs the dataset digest, approval chronology, operator
separation, blockers, status, case-complete observations, and balanced review lanes;
rehashed blocked-to-ready or row-removal edits fail. The command still performs no
model call and running it without genuine approved inputs remains blocked.
Production evidence assembly usability checkpoint (2026-08-14)
The third production-evaluation handoff now accepts the approved manifest,
completed production observations, and separate independent reviews directly.
It derives the canonical manifest identity and both raw file receipts, removing
six copy-prone configuration values. The operator must still provide the
attributable HTTPS production evidence URL, its retained receipt, and a
pseudonymous run code; the assembler cannot infer those facts. End-to-end proof
assembles 50 aligned traces and reviews, while absent real inputs still produce
only an integrity-protected blocked status and no evaluation bundle.
The assembly is now self-contained and semantically replayable: its private artifact
retains only the sanitized manifest, observations, reviews, release/operator fields,
and external receipt inputs, then reconstructs all row joins, chronology, reviewer
separation, citations, costs, blockers, counts, status, and bundle. Rehashed blocked-
to-ready, row-removal, or bundle-substitution edits fail before scoring.
Production AI scoring usability checkpoint (2026-08-14)
The final production scorer now accepts the approved manifest and ready bundle
directly, derives both receipts, revalidates frozen-label alignment, and emits
the same integrity-protected evaluation artifact without configuration edits.
An end-to-end subprocess test proves a 50-case, two-reviewer bundle yields only
row-derived precision, recall, citation, refusal, latency, and cost metrics.
The scorer still performs no model call, ignores aggregate assertions, and
fails closed when either genuine input is absent or any threshold is missed.
Together, all four production-evaluation handoffs now have direct CLI paths.
External validation bundle-binding checkpoint (2026-08-14)
The highest-leverage external lane now accepts a returned validation bundle
and private evidence envelope directly, keeping qualifications, attribution,
and the versioned correction decision out of tracked configuration. The signed
readiness artifact retains both the frozen protocol digest and the exact raw
validation-bundle digest before it can project the corrected investigation and
four qualified reviewers. End-to-end proof covers three researchers, one
subject expert, five independently coded readers, a distinct corrected v2,
explicit novelty/limitations decisions, and measured comparison outcomes.
Absent genuine evidence still produces four blockers and imports nothing.
Conventional baseline direct-reconciliation checkpoint (2026-08-14)
The separate measured-improvement lane now accepts a frozen manifest and
paired-session bundle directly and derives their canonical/raw receipts. It
still requires the attributable study source, measurement time, both
counterbalanced orders, unique consented independent researchers, three cited
sources per condition, facilitator receipts, 20% median improvement, and 80%
reproduction. End-to-end proof yields 50% median improvement and 100%
reproduction from rows and preserves the three participant codes that the A+
scorer must reconcile against the validation lane's qualified reviewers. No
aggregate metric or validation projection can bypass this independent check.
External impact cross-lane binding checkpoint (2026-08-14)
External citation/reuse/correction evidence now imports against the signed
validation readiness artifact rather than three retyped investigation fields.
The importer derives the exact corrected investigation ID, version, and packet
digest plus the raw event-file receipt; A+ scoring independently repeats the
release match. Public-URL validation now rejects localhost, private and
link-local IPv4, and private/loopback IPv6 in addition to internal and
synthetic hosts. End-to-end proof binds a verified citation to the corrected v2
packet, while the default lane remains blocked with no event to credit.
Acquisition release-binding checkpoint (2026-08-14)
Attributable acquisition now requires the integrity-verified organic provider
report itself—not only the qualified action ledger—to match the current
research release. Direct organic-report, campaign, and qualified-action inputs
derive raw receipts that remain inside the signed artifact. Shared public-URL
validation excludes localhost, private/link-local IPv4, and private/loopback
IPv6 verification references. End-to-end proof joins provider SEO/funnel data,
the exact approval-gated campaign controls, and one consented human-qualified
organic action; clicks, page opens, bots, and activity alone remain ineligible.
Net-income receipt and freshness checkpoint (2026-08-14)
The economics lane now accepts settlement and observed-cost exports directly,
derives both raw-file receipts, and retains them inside the signed artifact.
Live rows cannot cite Stripe test-mode dashboard paths; cost evidence must be
fresh, public-network, observed, order-attributable, and receipt-bound. The
row-level proof computes $29.00 gross less $1.14 fees and $2.00 observed cost
as $25.86 net while excluding a $99.99 sandbox settlement. The checked-in lane
still has no live settlement and therefore remains blocked; no checkout or
payment capability was activated.
Evidence-operations invocation checkpoint (2026-08-14)
The six-workstream operations artifact now exposes both a stable command ID and
the complete non-executing direct-input invocation for each current stage. The
production-AI template changes as verified artifacts advance; validation,
baseline, impact, acquisition, and economics templates enumerate their exact
bundle, evidence, provider, action, settlement, and cost inputs. The Markdown
handoff includes the same templates, while authority flags remain false and no
secret, external file, model run, contact, publication, or payment action is
performed.
Agent-safe evidence input contracts checkpoint (2026-08-14)
Every stage-aware invocation now carries structured contracts for all required
flags: evidence kind, description, private/non-private handling, and whether a
human attestation is mandatory or the value may be machine-derived. All nine
possible stage commands have non-empty, duplicate-free contracts. This lets a
bounded agent assemble local tasks and route sensitive files correctly without
inventing reviewers, receipts, production observations, approvals, or external
authority; the signed operations artifact retains the complete contracts.
Research net-income integrity checkpoint (2026-08-14)
A+ economics now derives positive net income from current-release live Stripe
settlements and separate observed order-cost receipts. Currency, offer revision,
refunds, fees, disputes, attribution, and sandbox exclusion remain explicit;
gross receipts and forecasts cannot substitute. No genuine live settlement is
retained, so the income requirement remains unearned.
AI question-collection prerequisite checkpoint (2026-08-14)
Production-evaluation operations now detect that the consented question lane
is inactive before recommending log attestation. The new
`research:ai-question-collection-readiness` command emits the exact stable
collection contract, its digest, a blank approval record, explicit operator
sequence, runtime-switch state, blockers, and a tamper-evident packet receipt.
It distinguishes awaiting human approval, ready for separate runtime
activation, and active states without approving research, deploying, enabling
collection, or contacting anyone. This removes a dead-end loop while retaining
the human approval and deployment boundaries.
Question-led SEO approval-contract checkpoint (2026-08-14)
The six substantive research-question guides no longer depend on an impossible
self-referential full-release approval digest. A stable content contract now
binds every title, question, section, source, FAQ, canonical publication effect,
and excluded claim. The publication-readiness command emits a blank human
review record and tamper-evident packet; indexing requires attributable HTTPS
review evidence and explicit claims/source, usefulness, rights/privacy, and
desktop/mobile accessibility decisions. Evidence operations selects this gate
before provider acquisition imports while the pages remain noindex and absent
from the sitemap. Approved pages also stop displaying the contradictory
"publication is not authorized" notice.
The packet now semantically replays the exact contract, explicit signed decision,
preparation chronology, derived blockers and status, operator sequence, and false
authority boundary. Its default blank is generated from the current contract;
checked-in configuration supplies no runtime approval. Non-public evidence URLs,
future approvals, and rehashed awaiting-to-approved substitutions fail closed.
Open-source research discovery checkpoint (2026-08-14)
Researchers can now discover one synchronized software identity through a
canonical, indexable `/research/software` landing page, repository-native
`CITATION.cff`, CodeMeta, and public JSON-LD at `/api/research/software`. The
human page gives direct reproduction, correction, expert-review, repository,
methodology, schema, dataset, and citation paths while sharing the machine
record's explicit claim boundary. The Research Commons and footer link it.
GitHub issue routing disables unstructured blank reports and directs people to
bounded research or private-support paths. Release tests reject repository or
version drift and verify public cache, CORS, OPTIONS, and artifact integrity.
The metadata intentionally asserts no DOI, deposit, named authorship, external
reuse, endorsement, validation, demand, impact, revenue, or income.
Approved-research SEO candidate checkpoint (2026-08-14)
The publication layer now has a deterministic bridge from an exact approved
report to a reviewable SEO landing-page candidate. It requires the reviewed-
findings channel, publication-approved status and flag, an attributable dated
researcher decision, two HTTPS-cited claims, limitations, and the current
research release. Output binds the full report digest and carries canonical
metadata, Article JSON-LD, citations, missing evidence, and internal links while
publication, deployment, indexing, and outreach authority remain false. The
current Rosenberg report correctly produces only a blocked artifact; a verified
approved-report envelope can be supplied without changing application code.
The candidate now retains the exact report and full publication envelope and
semantically rebuilds all copy, citations, JSON-LD, links, and robots state during
verification; rehashed landing-page edits cannot pass.
Review-to-publication envelope checkpoint (2026-08-14)
The report layer now binds the missing transition from genuine external review
to an SEO-eligible approved report. A ready validation artifact must prove the
materially corrected, limitations-bound non-synthetic revision. Separate human
publication approval must bind both exact report and validation digests and
affirm citations, limitations, rights, and novelty framing. The resulting
approved-report envelope changes report status/channel only inside a signed
handoff while retaining false publication and deployment authority. SEO input
now rejects approved-looking bare JSON and accepts only this verified envelope.
The current run remains blocked with no approved report emitted.
The envelope is now self-contained: it retains the unapproved source report and
complete validation artifact and replays correction/version requirements, approval
chronology, reviewer separation, editorial decisions, and the approved projection.
The preparation packet is self-contained as well and canonically reconstructs blocked
or ready state from the retained report, validation, explicit approval, and generation
time. Approval evidence must be public attributable HTTPS and cannot postdate packet
generation; rehashed status, blocker, or projection substitutions fail closed.
Dependency-aware approval-register checkpoint (2026-08-15)
Six previously fragmented approval and evidence lanes now feed one canonical,
tamper-evident register. It orders the core research path before optional email
and commercial activation, models prerequisites explicitly, and supplies the
private income cockpit with the first safe preparation command only after
digest verification. Every authority flag remains false. The current first
action is consented AI question-collection approval; external validation and
the post-review publication decision remain genuine external requirements.
Campaign and commercial readiness artifacts now sign their full persisted
records. The register therefore distinguishes integrity-valid blocked work from
tampered or unsigned artifacts without weakening any activation gate.
The register itself now retains those six source artifacts and reruns every
lane-specific verifier before reconstructing states, dependencies, and first action;
caller-supplied integrity booleans are ignored. Ready reviewed publication is sourced
from the full replayable approval envelope, while the preparation packet is used only
for blocked state. The private artifact can therefore reject rehashed routing edits
and forged ready summaries rather than trusting a one-time script check.
AI question-collection readiness now reconstructs its retained contract, explicit
signed approval, preparation chronology, runtime switch, blockers, and status during
verification. Its inactive template is generated from the current contract rather
than trusting stale null configuration fields. Public attributable approval URLs and
approval-at-or-before-preparation are mandatory; rehashed activation-state edits fail.
Organic expert-return checkpoint (2026-08-15)
The indexable contribution journey now closes the previously disconnected
handoff from device-local review to the repository's structured independent-
review form. The form captures a pseudonymous reviewer code, exact packet
digest, at least three checked public HTTPS sources, corrections, alternative
hypotheses, novelty, limitations, and AI-assistance disclosure while excluding
identity evidence and private files. The frozen Rosenberg workspace links the
same route but remains noindex. Public contribution is explicitly not accepted
as facilitator-verified qualification, A+ validation, publication approval, or
compensation.
It now also removes GitHub and prior facilitator contact as prerequisites for
creating a review. The no-account browser form collects only a packet digest,
pseudonymous code, reproduction, three public-source findings, corrections,
alternative hypotheses, provisional novelty, limitations, AI disclosure, and
four explicit declarations; it rejects direct identifiers and placeholder
hosts and downloads an integrity-signed envelope with every consequential
authority false. The page, sitemap, JSON-LD catalog, and public JSON Schema make
the pathway discoverable to humans and agents. Any A+ use still requires
separate facilitator verification of identity mapping, qualification,
conflicts, independence, consent, and custody.
The handoff now has its own consent-gated aggregate event and cockpit metric,
kept separate from contribution-path totals to prevent double counting. No
reviewer code, packet digest, account identifier, or external-form outcome is
collected, and an open never counts as a returned or qualified review.
The private cockpit now derives a fail-closed expert-review funnel state from
the verified aggregate report and signed approval register. It prefers genuine
returned validation evidence over click activity and limits operator actions to
evidence refresh, handoff inspection, approved-channel inspection, or distinct
human review—never visitor identification or contact.
The reproducible browser proof builds the current worktree before starting its
production server, then covers seven Research Commons journeys across desktop
Chromium and Pixel 5. The expert-return coverage exercises the no-account,
device-local review envelope and verifies its exact packet binding, authority
denials, optional public return route, privacy boundaries, accessibility, and
horizontal overflow. All 14/14 executions pass with no unexpected or flaky runs.
The release-handoff regression asserts all sixteen open targets, including the
expert-review schema, and fails if implementation and capture coverage diverge.
The no-account expert-review journey now continues through a non-overwriting,
offline envelope importer. It replay-verifies the returned device-local artifact
and routes it to the existing facilitator decision while requiring an exact
envelope-receipt match and retaining zero qualification or validation authority.
Facilitator decisions now also carry the exact candidate artifact SHA-256; the
verifier rejects missing bindings and reuse against a different returned review.
Operators can now generate a replay-verified, non-overwriting decision worksheet
from any supported return candidate. It copies the exact digest but leaves all
human evidence and checks blank, reducing transcription risk without automating
qualification or approval.
An editor-protected, no-index facilitator workspace now verifies the blank
worksheet entirely on-device, displays its exact candidate binding, requires
all applicable human checks, and downloads the completed decision without an
upload. The server verifier also rejects contact details in the retained
qualification basis, while the CLI remains the authoritative replay boundary.
The browser now also reconstructs and semantically checks GitHub and device-local
candidate projections, source receipts, nested signatures, and authority denials;
a rehashed candidate that grants itself qualification fails before form unlock.
Post-publication acquisition checkpoint (2026-08-15)
The attributable-acquisition join now requires a fourth, digest-bound evidence
lane: accountable human approval followed by evidenced public publication of
the exact Research Commons release. Qualified researcher or expert actions must
occur after that publication and inside the provider window, while every SEO
and analytics export must be captured only after the window closes. This
prevents internal tests, launch preparation, pre-publication visits, premature
exports, clicks, and opens from being credited as researcher acquisition. The
command remains blocked by default and grants no publication, campaign,
contact, qualification, or income authority.
The action-level attribution contract now closes the remaining self-assertion
gap. Every qualified row needs a unique consented first-touch receipt from a
Google or Bing organic visit to an approved research path, explicit internal-
traffic exclusion, independent-external participant status, and a human
attribution verifier distinct from the eligibility verifier. Missing evidence
fails closed; aggregate traffic and an `organic-search` string cannot establish
acquisition on their own.
The same ledger contract is now an open Draft 2020-12 JSON Schema served with
public caching and CORS, linked from the Research Commons, advertised in its
signed JSON-LD catalog, and discoverable through the sitemap. It encodes the
pseudonymous action, first-touch, independence, consent, and verification shape
while leaving cross-row uniqueness, distinct-verifier, chronology, freshness,
public-network, and digest checks to the authoritative importer. This reduces
real evidence-return friction without publishing any participant evidence.
Raw action JSON now crosses a strict runtime structural boundary before the
evaluator reads nested fields. Missing and additional fields, invalid enums or
constants, malformed pseudonymous codes, and incorrect hashes return a
canonical blocked artifact rather than throwing. Contract tests compare the
public schema's complete action, attribution, and qualification required-key
sets against a passing runtime fixture so documentation drift fails CI.
A+ projection-injection checkpoint (2026-08-15)
The authoritative A+ CLI now accepts only version and boundary metadata from
configuration. Attempted scoring projections are listed as rejected and never
evaluated; every scoreable external or technical projection must instead come
through its dedicated signed-artifact verifier. An adversarial CLI test injects
a perfect-looking AI evaluation and proves the requirement remains missing.
Hand-authored configuration can no longer substitute for production evaluation,
review, baseline, impact, acquisition, income, browser, or release-quality
artifacts.
The persisted A+ scorecard now retains its verified evidence projection and replays
all eleven requirements at the recorded evaluation time. Grade, summary, evidence
copy, next actions, rejected configuration fields, and authority are derived rather
than trusted. Evidence operations verifies this artifact and falls back to zero
passed requirements when it is missing or modified, so a rehashed grade cannot
misroute the private operator queue.
Release-bound production evaluation checkpoint (2026-08-15)
Production AI evidence now forms one release-consistent chain. The privacy-safe
assembler derives the canonical Research Commons release digest into its bundle;
the metric evaluator independently recomputes and checks that digest; and the
A+ scorer requires the projected AI digest to equal the open-research digest.
An older valid 50-case evaluation can no longer be combined with a newer code,
prompt, workflow, schema, or documentation release. Focused assembly, evaluator,
CLI, scorer, chronology, and mismatch tests pass without executing a model.
Release-bound external validation checkpoint (2026-08-15)
Independent validation now carries the same release identity. The readiness CLI
derives the canonical Research Commons digest into the signed corrected-v2
investigation projection, and the A+ scorer requires it to match the current
open-research release. Packet-level correction, reviewer qualification, and
artifact integrity remain mandatory, but they can no longer validate a changed
workflow merely because an older packet still verifies. Direct validation,
signed import, CLI, A+ mismatch, and publication-handoff tests retain coverage.
Preregistered release-bound baseline checkpoint (2026-08-15)
Measured workflow improvement now names the tested implementation before any
session occurs. The conventional-comparison manifest freezes the exact Research
Commons release digest with its protocol, cases, workflows, thresholds, and
preregistration receipt. The importer independently checks that digest against
the current release, retains it in the signed projection, and A+ requires it to
match the open-research release. An older speed/reproduction result cannot be
credited after the workflow changes, even when its study artifact still verifies.
---
History And References
Evidence correction: open does not mean locally implemented
The Research Commons open-release point now requires exact-release production
deployment approval and successful post-deployment response receipts for every
public resource. Local file existence cannot score. The importer is non-executing,
so publication remains a human operation and the point stays unearned until genuine
production evidence is supplied.
The deterministic evidence-operations report now covers all eight currently
missing evidence lanes: open-release production proof plus the seven external
evaluation, review, baseline, impact, acquisition, and income requirements. Its
first safe action is an importer invocation, never an automatic deployment.
Bounded agent work orders
The device-local Commons journey now converts its narrative agent plan into eight
standard AgentTask envelopes bound to the packet digest. Dependencies, candidate
ceilings, model-call ceilings, rights warnings, and human-trigger requirements are
machine-readable. Retrieval starts blocked and no task authorizes execution,
conclusions, publication, contact, or payment. The work-order schema is cataloged,
sitemapped, CORS-readable, and part of exact-release public probe evidence.
The client now verifies the signed artifact before download by recomputing the
canonical packet and work-order digests and reconstructing the expected DAG from the
packet. Rehashed dependency removal, inflated candidate/model budgets, source-plan
substitution, or cross-packet reuse fails even when generic AgentTask fields remain valid.
The public schema now closes task/input/budget/authority objects and fixes the hard
ceilings and false authorization flags rather than documenting them only in prose.
Open-release artifacts now retain their complete normalized deployment and probe
bundle. The verifier recomputes the evaluation and compares canonical outputs, so
an outer integrity hash no longer substitutes for semantic revalidation. The A+
next action names the exact evidence importer and required post-deployment receipts.
External citation/reuse/correction artifacts now retain their full normalized
event ledger and exact reviewed-investigation binding. Integrity verification
replays the evaluator at the retained evaluation time, preventing rehashed changes
to projected citation, reuse, correction, source, or event-ID results.
Economics artifacts now retain their row-level Stripe settlement and observed
cost ledgers. Verification replays live eligibility, freshness, sandbox exclusion,
refunds, fees, attributable costs, and net income, preventing rehashed profit edits.
Acquisition artifacts now retain the complete provider, campaign, publication,
first-touch, and qualification join. Semantic replay prevents rehashed traffic,
contribution-open, or qualified-review inflation from entering the A+ scorecard.
Measured-improvement artifacts now retain the frozen manifest and complete paired
session bundle. Semantic replay recomputes medians and reproduction while enforcing
session-before-measurement/evaluation chronology, preventing rehashed performance claims.
Production AI evaluation artifacts now retain the approved frozen label manifest
and complete trace bundle. Verification rechecks alignment and recomputes all
quality, latency, cost, coverage, and reviewer-independence metrics row by row.
External validation now enforces attestation chronology: qualifications and the
material-correction decision cannot postdate measurement, and correction cannot
predate independent review or returned-evidence assembly.
External-validation artifacts now retain the complete identifier-free returned
evidence and semantically replay the A+ projection at every downstream ingestion
boundary. Rehashed edits to reviewers, correction, investigation version, or study
results therefore fail before publication, impact, or A+ scoring can consume them.
The open-release handoff now starts with `pnpm research:open-release-handoff`, a
non-network preparation command that binds the current release, fixed sixteen-route
allowlist, blank human decision, exact sequence, and false authority flags in a
non-overwriting, semantically replayable packet. Evidence operations selects this
safe first action while no integrity-valid capture exists. Only after attributable
approval and production deployment may the bounded receipt-capture command probe the
fixed sixteen-route
allowlist covering the human Commons, schemas, JSON Feed, RSS, and JSON-LD software
metadata; refuses redirects, failures, unexpected content, time inversions, and
oversized bodies, and writes non-overwriting exact-body hashes for the independent
evidence importer. The operations queue exposes the two-stage capture/import path
without granting deployment or publication authority. Its stage resolver now replays
the configured capture through the authoritative evaluator: only an integrity-valid,
exact-current-release bundle advances from capture to import, preventing both repeated
network work and premature progression from malformed receipts.
The open-release queue now also closes the duplicate-handoff loop. It searches retained
non-overwriting handoffs, accepts only a canonically verified packet for the exact
current release prepared no more than 35 days ago and not in the future, and then points
to the approval-gated capture invocation. A stale, forged, malformed, or previous-release
packet returns to preparation. This improves operator usability without treating the
blank handoff as approval or granting deployment, network, publication, or indexing
authority.
The approval boundary no longer depends on manually setting `humanApproved: true`.
An editor-only, noindex, device-local workspace builds the exact current handoff,
requires an attributable APR-coded approve/reject decision with five explicit safety
checks and a public receipt digest, and downloads a signed artifact without persistence
or execution. Receipt capture now replays that decision against the current release and
requires its decision time to match the deployment configuration. A forged decision,
rejection, stale release, timestamp substitution, or standalone boolean cannot reach
the network capture stage.
The post-deployment handoff is now non-manual as well. A no-network, non-overwriting
assembler consumes the signed decision and production deployment receipt, replays exact
release and chronology, validates attributable public receipts and distinct APR/OPS
roles, and writes an ignored private capture config with false execution authority.
Capture re-verifies it before networking and automatically emits the completed importer
config beside the bounded probe receipts. Evidence operations can consume an explicitly
named private config to advance deterministically from config assembly to capture and
then import without editing timestamps or relative paths.
The private deployment workspace now offers the same join as a device-local browser
flow. It accepts the two JSON files without upload, checks both canonical signatures,
the current release and fixed target set, public evidence hosts, APR/OPS roles,
production/synthetic flags, and 35-day chronology, then downloads a configuration that
the capture CLI must replay before networking. Browser proof feeds its output to the
authoritative server verifier on desktop and mobile. This removes shell-only friction
without creating deployment, capture, publication, indexing, or A+ authority.
Open-release receipts are now self-contained rather than hash-only: every bounded
public response body is retained as canonical base64, exact hashes are recomputed at
ingestion, and media-specific JSON/JSON Feed/RSS/HTML/Markdown structure is reparsed.
This closes the forged-hash and status-200 error-page gap while retaining only public,
allowlisted content and enforcing the existing per-resource size bound.
Correction propagation now emits a signed, self-contained artifact rather than an
unsigned derived packet. It retains both exact input files within a 2 MB-per-file
bound, recomputes raw receipts, validates the original/envelope digests, reruns the
allowlisted field edit, and compares the entire revision canonically. Rehashed
publication authorization, revision, correction, or source-file substitutions fail;
the separately written candidate still carries no acceptance or publication authority.
The expert-validation flywheel now has a deterministic return preflight between the
public structured issue form and private facilitator review. The offline command binds
one return to its packet and exact repository issue, enforces public-source, privacy,
consent, pseudonym-role, and AI-disclosure rules, writes only to a new path, and signs
the complete retained input. Its output is explicitly non-qualifying until a separate
accountable facilitator verifies identity, expertise, independence, conflicts, and
consent; no review status can authorize contact, publication, payment, or an A+ claim.
Facilitator qualification now forms a second signed link rather than an editable
configuration row. The verification command retains the complete candidate and a
distinct post-capture facilitator decision, requires an attributable receipt and
eight explicit checks, and deterministically projects one pseudonymous attestation.
Validation readiness rejects inline qualifications and requires the three researcher
plus one subject-expert artifacts to replay and bind the exact dossier packet before
the external-review projection can advance.
Material correction and novelty are no longer editable validation projections. A
new non-overwriting decision artifact replays the signed correction propagation,
binds the same verified subject expert and original packet, requires a later human
acceptance with six explicit checks, and derives the distinct v2 packet identity.
Validation readiness now rejects inline correction decisions and consumes only this
replayable artifact, while publication, title, authenticity, legal status, contact,
and payment authority remain false.
The measured-improvement lane now closes the cross-study pseudonym-reuse gap. A
same-study join replays the frozen comparison and external validation, requiring
their release, protocol, packet, measurement, participant rows, timing, durations,
reproduction, and citations to match. A+ retains the join and validation sources,
so recomputing a scorecard hash cannot substitute copied codes or aggregate metrics.
The scorecard trust root now covers every external lane, not only validation. Its
private provenance retains the complete verified public release, production AI
evaluation, validation, same-study baseline join, impact, acquisition, and economics
artifacts. Verification replays every source and requires exact projection equality;
missing, unused, or semantically modified provenance invalidates the scorecard even
when an attacker recomputes its outer digest.
The same trust root now covers the three technical projections. Workflow replays the
full responsive browser report; automation retains exact package bytes tied to the
release manifest; quality retains and replays the six-command release artifact. This
removes the final summary-only path by which forged technical booleans could complete
an otherwise source-valid A+ scorecard.
The public Commons now routes each visitor by intent across seven low-friction paths:
read, reproduce, build, contribute, private assessment, self-serve kit, and bounded
human work. Its machine-readable catalog advertises the same discovery and commercial
options. Typed consent-gated aggregate events measure route choice without identity or
artwork data, while the evidence boundary prevents navigation from becoming qualified
interest, scholarly validation, demand, a sale, settlement, revenue, or income.
Commons intent measurement now continues through the provider-export boundary rather
than ending at browser markers. Strict ingestion retains the seven route counts,
rejects unknown or duplicate rows, derives their aggregate, and exposes both levels in
the private cockpit. This closes the instrumentation-to-operator gap while preserving
the separate qualified-action, settlement, and net-income evidence lanes.
A repository-public JSON Schema now specifies the complete normalized organic evidence
bundle, and the private cockpit links directly to it. A drift test compares the schema
event enum with the runtime allowlist, eliminating TypeScript/test inference as the
operator contract while retaining runtime rejection as the authoritative gate.
The authoritative importer now closes the remaining structural and chronology
gaps. Exact keys are required at every bundle level, exports must be captured
after their reporting window, lead statuses and settlement IDs are unique, and
payment state/mode/source/receipt values are validated before any projection.
Payment rows also require explicit USD currency, and a non-overwriting offline
settlement-ledger normalizer now rejects identifying/unknown fields, duplicate
IDs, invalid receipts, impossible amounts, and chronology inversions before
bundle assembly. Normalization does not authenticate provider evidence.
A non-writing organic evidence preflight now checks that full runtime contract,
exact release binding, and 30/60/90-day window before an operator can replace a
report. Its output is deliberately limited to digests and row counts, and the
writing command reruns the same gate to prevent preflight/import drift.
The legacy standalone report writer can no longer bypass that gate. Both
writers require the period, exact current release, and non-future provider
timestamps before their first write, with parity locked into the release
manifest and regression suite.
Raw-provider automation now begins with a bounded GA4 adapter. A filtered
two-column aggregate export becomes a non-overwriting normalized fragment
offline, with the runtime event allowlist shared directly with ingestion and
strict rejection of dimensions, duplicates, unknown events, invalid counts,
future capture times, and non-Analytics evidence hosts.
A deterministic offline assembler now joins all five normalized provider
fragments, injects the current release and declared reporting window, and runs
the complete preflight before a non-overwriting write. This removes manual
bundle editing while retaining provider authentication and evidence acceptance
outside the assembler. Every fragment now independently binds the exact window,
so valid evidence from different periods fails assembly rather than silently
mixing. A paired Search Console/Bing CSV normalizer also validates exact daily
click/impression rows, provider hosts, indexed-page counts, dates, and ratios
offline without provider or evidence authority.
The signed acquisition-evidence regression fixtures now replay those same
window-bound provider sources, closing the legacy-report bypass at the next
evidence layer.
All five fragments and signed source receipts now retain the raw-input SHA-256,
enabling replay against privately retained exports without placing sensitive
lead or settlement rows in normalized or public artifacts.
A lane-aware, no-write fragment replay command now regenerates every search,
GA4, lead, or settlement fragment from retained raw input, fails on any drift,
and emits only privacy-safe digests, row count, and false authority.
Temporal attribution is now enforced below replay: each lead state requires its
status-specific event in the half-open window, and every payment carries a
window-validated recognition instant through normalization, import, and report
calculation. Period metrics can no longer silently use lifetime status rows or
out-of-period settlements.
Period accounting now separates settled-order gross from refund/dispute
movements. Refund-only periods preserve negative net income, fees, and costs
without increasing gross, order count, or purchase conversion.
Five separate replay steps are now composed by an all-or-nothing custody
preparer. It validates every raw/fragment pair and exact-release bundle before
creating a new directory containing only the bundle and privacy-safe custody
manifest; drift creates no partial package and grants no authority.
as separate external steps.
The first external-evidence handoff no longer requires Release Operations to
hand-author deployment-receipt JSON. The private open-release approval workspace
downloads an exact-release-bound worksheet whose evidentiary fields remain null
and explicitly require provider proof. Both the browser and authoritative CLI
assemblers reject the incomplete worksheet, so this usability bridge cannot be
mistaken for deployment or an earned open-Research-Commons readiness point.
Desktop Chromium and Pixel 5 browser proof now download that worksheet, bind it to
the visibly rendered release digest, verify every provider fact is still null, and
pass axe plus horizontal-overflow checks. This is reproducible usability evidence,
not production-deployment evidence.
The federated provenance query UI and API were deployed to the canonical
production domain on August 18. Vercel build `metamuseum-mwcrf3ypb` completed
all 216 routes and the canonical supported/refusal probes passed. Compile-time
public research inputs now live under `src/data/research`, separating deployable
product data from ignored generated evidence. The deployment clears the runtime
availability blocker; a new exact-release quality digest, signed decision, and
provider receipt remain required before the open-release A+ point can be earned.
Offline normalization now covers the sensitive lead lane as well. Known
encrypted store rows become only three consent-state counts; plaintext or
unknown fields, duplicate IDs, unsupported states, time inversions, and output
overwrite fail closed, and no identifier-like value enters the fragment.
The public-page quality gate now includes an explicit regression contract for
research contribution controls at narrow mobile widths. The 390px Linux
Chromium failure is reproduced and corrected with a zero-overflow measurement;
GitHub Actions run `32589651885` and canonical smoke run `32589705508` passed.
The next public-surface audit increment replaces the former partial route list
with a 45-route manifest and dynamic coverage for every published Connections
and Stories page. The resulting 150-case desktop, zoom-equivalent, and mobile
matrix found two previously invisible defects: `/docs` escaped its viewport and
the Connections chapter numerals missed contrast requirements. Both now pass.
Homepage rotation also uses verified open-image records to repair sparse stored
provider groups, increasing observed representation from two sources to five in
the running app. Full fourteen-provider image representation remains open and
must be earned with item-level rights evidence and reliable display URLs.
The next provider-qualified increment adds Smithsonian American Art Museum to
the fallback rotation through a verified CC0 object and media record, bringing
the clean-deployment fallback to six museums. The quality gate is also now
reproducible on developer machines: regression checks use an isolated
seed-backed record store, while `--managed-storage` remains available for an
explicit operational drift audit of imported records. This prevents mutable
local collection data from masquerading as a CI code regression.
The first Connections editorial redesign is complete locally. Its public index
now previews three decisive turns, while the flagship detail page follows six
question-led steps from shared catalog subject through documented museum facts
to a bounded interpretive reveal. Every turn references declared claims, anchor
navigation is keyboard-addressable, and the route passes Axe and overflow checks
at desktop, zoom-equivalent, and mobile sizes. Expanding beyond one published
journey remains gated on equally strong source evidence and review discipline.
The public-route audit now includes Surprise Me. The former redirect could eject
visitors to an arbitrary provider page, including slow or incomplete third-party
states, while dropping Meta Museum's explanatory and rights context. The route
is now a force-dynamic internal artwork page selected from one verified lead per
provider, with explicit attribution, rights and source links, a non-ranking
boundary, and onward discovery choices. Deterministic tests prove that all
fourteen curated fallback providers are reachable and that later records from a
repeated provider cannot displace its verified lead. Canonical desktop and
mobile browser proof remains required for the exact deployed revision.
The static public-page audit manifest now names `/surprise` because the route is
intentionally absent from the canonical sitemap; this closes the browser-gate
coverage gap without presenting random responses as indexable editorial pages.
The next public-route increment replaces the sparse Research Reports list with
an editorial research desk. The sole Rosenberg report remains an under-review,
non-indexable lead, but its public index now exposes the motivating question,
source-backed significance, four evidence counters, unresolved evidence, exact
HTML and JSON routes, and a three-stage lead/review/approval ladder. The empty
reviewed-findings lane is deliberately prominent, and participation links route
to evidence contribution, the validation protocol, and the published method.
No reviewed finding, scholarly novelty, completed restitution, adoption, or
revenue is claimed. Full browser and canonical deployment proof remains open
for the exact revision.
The following public-route increment addresses the Linked Art Inspector. The
deployed tool was functional but supplied almost no orientation beyond its
editor: it did not state what was checked, distinguish inspection from
certification, or offer a route from the result to the model, datasets, or
records. The revised surface frames three explicit validation areas, a
read-only default, a three-step inspect/import workflow, and four inspectable
onward paths. Import authorization is unchanged and stays fail-closed.
Technical recommendation packet: behavior ownership remains split between the
server-rendered route framing and the existing client workbench. The three
highest observed risks are (1) a specialist editor appearing before a novice
can understand its purpose, (2) inspection being mistaken for certification or
publication, and (3) a successful result becoming a dead end. Curator
Experience owns the explanatory hierarchy and responsive layout; API Platform
owns inspect/import semantics and must preserve the permission gate; Quality
owns source-contract, API, browser, and accessibility regression checks.
Acceptance requires the exact public copy and four onward links, unchanged
import authorization, zero severe Axe findings, no horizontal overflow at all
three audited widths, and passing `pnpm test`, `pnpm lint`, and optimized
`pnpm build`. The next-cycle hypothesis is that another public tool route with
low context density will be the weakest remaining surface; falsify it by
re-running the canonical route inventory and direct browser comparison after
deployment.
The Getty provider workspace is the next repaired canonical route. Its former
layout exposed useful capabilities through repeated generic cards and raw
endpoint output, but did not connect those capabilities to an artwork or tell a
reader what the outputs could and could not establish. The revised experience
begins with a verified public-domain Getty image and source record, then follows
a numbered path from Linked Art and IIIF through read-only SPARQL to
ActivityStreams. Query results remain distinct from rights clearance,
attribution, provenance conclusions, and entity reconciliation; stream events
remain distinct from changes to, destruction of, or loss of a physical object.
Technical recommendation packet: the route remains a server component with URL
state, and the Getty adapter retains all provider parsing and network behavior.
The three highest risks were (1) technical capability without cultural context,
(2) raw query success being overread as a research or rights conclusion, and
(3) ActivityStreams deletion being mistaken for object destruction. Curator
Experience owns artwork-led hierarchy and responsive styling; Provider Platform
owns the adapter and read-only enforcement; Quality owns page, API, adapter,
browser, and accessibility evidence. Acceptance requires the verified featured
image/source/rights tuple, preserved shareable query parameters, explicit
semantic boundaries, no inline page styling, zero severe Axe findings, no
horizontal overflow, and successful full lint/test/build/smoke gates. The next
cycle should test whether the remaining lowest-context provider/tool route is
now `/research/query`, `/collections`, or another canonical surface rather than
assuming sparsity alone proves weakness.
The homepage visual-coverage gap is now closed in the release candidate. A
rights-qualified fallback round covers all fourteen production providers once
before any provider repeats, uses fourteen distinct image URLs, and keeps
collection-record provenance separate from image-source and license provenance.
Unknown or review-required rights are excluded. The carousel also drops failed
media during a session and disables automatic movement when reduced motion is
requested. Production completion still requires the full quality gate,
deployment, and canonical browser verification for this exact revision.
The first canonical exercise caught the AIC fallback IIIF endpoint returning
403; the release correction preserves the AIC object record while sourcing a
verified public-domain image of that exact painting from Wikimedia Commons.
The next low-context public-tool repair replaces `/graph`'s brief technical
header and pointer-dependent canvas with an artwork-led relationship atlas. A
verified public-domain Met sculpture demonstrates four documented node types,
then the live graph exposes every encoded edge through a text relationship
index linking both endpoints. Museum fact, catalog relationship, editorial
sequence, uncertainty, image rights, and inference boundaries remain visibly
distinct. The Cytoscape node-pivot behavior and tenant preview scope are
preserved.
Technical recommendation packet: the server route continues to own record
loading and graph projection, while the existing client component owns only
visual layout and selected-node state. The three highest risks were (1) a
specialist graph appearing without a cultural question, (2) pointer-only
navigation excluding keyboard and screen-reader users, and (3) shared concepts
being overread as influence, attribution, provenance, identity, or historical
contact. Curator Experience owns the artwork-led example and claim labels;
Graph Platform owns the derived nodes and edges; Quality owns page, component,
responsive, image, accessibility, and complete-link evidence. Acceptance
requires the verified image/source/rights tuple, preserved live pivot behavior,
all graph edges available as text routes, no page-authored inline styles, zero
severe Axe findings, no horizontal overflow, and successful full
lint/test/build/smoke gates.
The reconciliation review route is the next public-method repair. It now asks
the reader-facing question of when two records describe the same event, explains
the evidence reviewers compare, and states that confidence is a routing signal
rather than proof of identity, attribution, ownership, provenance, or catalog
authority. The four service-configured threshold bands, live decision counts,
human-review links, and optional AI-tiebreaker status remain intact. A zero-item
queue is explicitly a dataset status rather than evidence of completeness and
offers direct paths to entities, source datasets, and the canonical B6.1 method.
The route is read-only and continues to preserve rather than merge or rewrite
source records.
Technical recommendation packet: Curation continues to own threshold decisions
and reversible review, Data Platform owns candidate generation and provider
evidence, and Public Experience owns explanation and navigation. The three
highest risks were (1) a score being mistaken for proof, (2) an empty queue being
mistaken for a clean corpus, and (3) an optional model tiebreaker being mistaken
for publication or merge authority. Acceptance requires unchanged tested
threshold behavior, an explicit read-only and non-proof boundary, useful empty
state routes, one main landmark, no page-authored inline styles, responsive and
accessible browser evidence, and complete quality/deployment gates.
The route audit also closed an authorization mismatch on
`/curator/annotations`: anonymous visitors could render reviewer identity and
decision controls even though every review submission correctly failed closed at
the editor-only API. The workbench now requires an editor for `GET`, while the
read-only reconciliation methods route deliberately remains public. Acceptance
requires role-unit coverage, an anonymous browser redirect to the custom sign-in
surface, authenticated editor reachability, and unchanged API write protection.
The next public-route increment closes an irreversible-action defect on My
Collection. “Clear collection” no longer deletes every browser-local save on the
first click: it opens a labelled inline alert, focuses the non-destructive choice,
states that clearing cannot be undone, and requires a separate confirmation.
Collection data remains local to the browser and no account, sync, upload, or
sharing behavior is introduced. Acceptance requires reducer coverage for every
state transition, keyboard-operable controls, distinct destructive styling, zero
severe Axe findings or mobile overflow, and the complete build/deployment gates.
The visual-comparison route is the next repaired public research surface. It
now asks what changes when two museum images meet at the same scale, explains a
three-step close-looking method, and keeps the current pair's museum records,
rights language, and attribution adjacent to the viewer. Public selection
round-robins only the image-ready providers measured in the current record
store, retains one preferred rendition per artwork, and defaults to two
different works. The current seed/store truth is one image-ready provider;
fourteen governed connectors are documented separately and are not represented
as fourteen-provider image coverage. Optional rectangles are off by default and
labelled as Meta Museum demonstration guides, never museum annotations.
Technical recommendation packet: Public Experience owns the question-led
hierarchy and plain-language controls; Data Platform owns provider-aware
selection and rendition deduplication; Quality owns localized navigation,
desktop/mobile interaction, image loading, and accessibility evidence. The
three highest risks were (1) two renditions of one artwork masquerading as a
comparison, (2) connector count being mistaken for image-ready coverage, and
(3) visual similarity being overread as attribution, influence, identity, date,
or provenance. The optimized-server audit also found and repaired a Next.js 16
locale rewrite loop by distinguishing the propagated internal locale pass from
a new public navigation. Acceptance requires distinct default artworks, one
choice per artwork, truthful provider counts, source and rights links for the
active pair, demo guides off by default, a localized 200 without redirect loops,
zero severe Axe findings or horizontal overflow, and the complete
test/lint/build/deployment gates. The next-cycle hypothesis is that a canonical
Connections or Stories surface now carries the weakest combination of cultural
question, source visibility, and useful onward action; falsify it through the
route inventory and direct comparative browser audit after this exact revision
deploys.
The Connections audit found that the flagship episode itself is visually and
structurally credible, but its surrounding index stretched one approved journey
without an editorially useful onward route. Both the index and flower journey
now present the three existing source-backed Stories as companion trails. Each
card begins with a different question, retains its lead artwork, museum sources,
and public-domain statement, and links directly to the corresponding Story.
Visible copy prevents readers from treating those Stories as additional
approved Connections. Public Experience owns the shared companion component;
Editorial owns the question framing and format distinction; Quality owns exact
link, source, rights, responsive, accessibility, and image checks. The three
highest risks are (1) breadth being implied by relabelling Stories as episodes,
(2) an onward-action grid becoming generic card filler, and (3) source names or
rights disappearing at the format boundary. Acceptance requires all three
canonical Story routes on both Connections surfaces, explicit format language,
distinct questions and images, retained source and rights text, zero severe Axe
findings or horizontal overflow, and complete test/lint/build/public-smoke
gates. The next-cycle hypothesis is that the three Story detail pages now have
the weakest onward path into the wider collection; falsify it by comparing
their terminal actions and completion behavior after this exact revision is
deployed.
The companion-Story audit then found an exact terminal-path defect across all
three details: each finished with the same generic index, Explore, checklist,
and paid-kit links, leaving no narrative handoff from the clue the reader had
just followed. The shared Story template now computes a deterministic next
Story and previews it with a distinct entry question, artwork, provider set,
rights statement, and direct link. A separate longer-trail band routes readers
into the six-turn flower Connection and retains its museum-fact, inference, and
no-influence framing. Both discovery routes precede the optional paid method,
which remains clearly identified and secondary. Public Experience owns the
cyclic continuation and responsive hierarchy; Editorial owns entry questions
and handoff language; Data/Trust owns museum attribution and image-rights
retention; Quality owns route, ordering, image, focus, mobile, and accessibility
evidence. The three highest risks were (1) a paid offer interrupting cultural
discovery, (2) a generic “read another” action hiding the next intellectual
question, and (3) a related-story image losing its source or reuse boundary.
Acceptance requires one unique next-story route per detail, a complete three-
story cycle, the longer Connection route, discovery before product, retained
source and rights text, working high-resolution images, no severe Axe findings
or overflow, and full test/lint/build/public-smoke gates. The next-cycle
hypothesis is that the Stories index now understates the three-story sequence
because it presents a lead plus two supplements rather than a connected reading
route; falsify it through a production hierarchy and navigation audit after
this exact revision deploys.
The next flagship-Connection audit found a different narrative break: the
opening comparison was visually strong, but every one of the six explanatory
chapters became text-only precisely when the reader was asked to inspect a
flower, crop, horizon, or format. Each public turn now declares one or more
artwork evidence views in the journey model, including an intentional crop,
plain-language label, source-visible caption, and exact museum artwork target.
The final reveal restores both paintings side by side so “time reconstructed”
and “space reconstructed” remain a comparison rather than an unsupported
influence claim. Public Experience owns the responsive evidence-view system;
Editorial owns the crop labels and bounded captions; Data/Trust owns artwork-ID,
source, and rights integrity; Quality owns model validation, link, image,
keyboard, mobile, and accessibility evidence. Key risks are decorative image
repetition, crops implying facts beyond the record, and draft material becoming
publishable without visual evidence. Acceptance requires evidence imagery in
all six public turns, seven valid declared-artwork views, a two-work reveal,
visible rights attribution, no overflow at desktop or 390 pixels, and the full
test/lint/build/accessibility/interaction/deployment gates. The next-cycle
hypothesis is that the Connections index still overstates breadth around one
approved episode; test whether a clearer “one released investigation” frame and
an editorial preview queue improve credibility without promoting drafts.
That index hypothesis was confirmed: “Episode 01” and “Tonight’s mystery” made
one released investigation feel like a thin entertainment series rather than a
deliberate research boundary. The index now names exactly one released
independent investigation and introduces two future topics only as questions
under source review. Each notebook entry explains intellectual significance,
lists the verification still required, and links to both official museum
records; neither unpublished slug is routable or linked, and no unchecked image
or conclusion appears. Editorial owns question quality and the distinction
between inquiry and finding; Data/Trust owns official-record targets and status
derivation from the two isolated drafts; Public Experience owns the ruled
notebook hierarchy and responsive behavior; Quality owns exact-count, no-draft-
link, source-link, focus, overflow, and route checks. Principal risks are a
research queue being mistaken for publication, static questions going stale as
drafts change, and non-clickable cards suggesting missing interactions.
Acceptance requires explicit “not published” labels, four functioning official
record links, zero public draft routes, one and only one released feature,
desktop/mobile accessibility, and all deployment gates. The next-cycle
hypothesis is that the Stories index, which still frames one lead and two
supplements, now understates the deterministic three-story reading sequence.
That hypothesis is now implemented locally: `/stories` presents all three
released stories as one numbered route with six artwork images, equal editorial
weight, explicit source and rights context, and two visible next-clue handoffs.
The route ends in the longer flower Connection rather than a generic card grid.
Acceptance requires focused and full tests, lint, optimized build, desktop and
mobile browser inspection, accessibility and interaction audits, exact-head CI,
canonical smoke, and direct production verification before this slice closes.
The August 23 completion audit verified all four notebook identities through
the museums' collection APIs and received image/jpeg 200 responses from all
three companion image proxies. Direct 1440px and 390px browser inspection found
the released feature and notebook contained with no horizontal overflow. Local
launch evidence in `artifacts/launch/connections-index-a11y-2026-08-23.json`
passes 207/207 route-viewport cases with zero violations, while
`artifacts/launch/connections-index-interactions-2026-08-23.json` passes 67
routes, 4,795 visible links, 906 unique internal targets, and 45 fragment targets
with zero failures. The focused 27-test editorial contract, full serial suite,
ESLint, and the 218-page production build pass; the build used an ephemeral
process-only `AUTH_SECRET` and wrote no secret to the repository.
The current public-quality slice also replaces the raw documentation catalog
with a curated six-entry start path, searchable complete inventory, and readable
source view. The route inventory now enumerates 77 pages, 190 handlers, and 351
HTTP methods; the interaction audit covers 90 routes and 685 internal targets.
Duplicate nested main landmarks found by that expanded audit were repaired on
the docs hub and six Rosenberg research routes before release.
Production reader inspection additionally caught a duplicated source-title H1
that automated WCAG checks did not flag. The reader now removes only the leading
Markdown title from its rendered body and preserves the substantive section
anchors beneath one page-level heading.
The next public-discovery repair addresses the original cross-provider flower
search complaint. Ranking now removes presentation-identical records, excludes
zero-relevance filler when genuine query matches exist, places image-backed
works first, and interleaves equally relevant providers before repeating one
source. In all-provider mode, relevant no-image records move into a separate
linked source list with an explicit combined-provider count; they no longer
weaken the visual gallery or masquerade as missing artwork thumbnails.
Local browser evidence for `flowers&source=all&limit=18` now shows 18
image-backed gallery cards, eight separately labelled no-image source records,
six represented providers overall, and zero desktop/mobile overflow. Full tests,
lint, the optimized build, accessibility, and interaction gates pass locally;
canonical verification remains required after deployment.
The next route-quality slice replaces `/patterns`' raw twelve-card heuristic
dump. The public page now excludes administrative collection-page groupings and
generic or singular/plural duplicate labels, exposes six reviewable questions,
states the record count and non-inference boundary for every lead, and provides
twenty-four named artwork links. The record-quality queue is capped at eight
clearly labelled tasks, while the reviewed flower Connection demonstrates the
separate publication standard for a lead that survives source review.
Acceptance evidence now passes locally: focused selection/rendering tests, the
full serial suite, ESLint, the 218-page optimized build, 90-route interaction
integrity, and all desktop, 200%-equivalent, and mobile accessibility checks.
Exact-head CI, canonical smoke, and direct production inspection remain the
release boundary for this slice.
The Contact route now replaces an undifferentiated message panel with three
explicit, privacy-bounded paths: collection pilot, provenance review, and
support for the public work. Query input is resolved through a fixed allowlist,
only the approved subject reaches the visitor's own email app, and no contact
form or upload surface is introduced. The route explains what context belongs
in a first note, keeps confidential material out of scope, and makes clear that
an inquiry grants no engagement, publication, data-access, or payment
authority. Unit coverage locks the allowlist and rejects arbitrary subject
text; the complete organic funnel passes six desktop/mobile Chromium cases with
zero Axe or overflow failures. Canonical production inspection remains the
post-deployment boundary.
The first governed Calliope comparison run now covers both registered held-out
cases across five repetitions using Anthropic Haiku 4.5. Exact response usage
totals 8,395 input and 2,631 output tokens ($0.021550); every per-run cost stays
below policy. The maximum 13,130 ms latency breaches the 12,000 ms ceiling, and
independent blinded scoring remains incomplete, so the candidate retains no
promotion, deployment, or publication authority. Runtime responses now expose
only provider/model/token metadata needed for cost evidence and fail closed when
required usage is absent or malformed.
Calliope's first defect-repair cycle now passes the previously failed local
operational gates. Explicit malformed Linked Art rights statements refuse before
model spend, already-refused tasks cannot enter the provider branch, and the
Anthropic timeout is clamped below the 12-second value ceiling. The repeated
held-out rerun recorded five representative provider calls and five adversarial
pre-model refusals: maximum latency 3,920 ms, total exact cost $0.010505, and no
adversarial provider calls. Independent blinded quality and reviewer-time scores
remain the next gate; the result grants no promotion or runtime authority.
Calliope's v3 governed rerun now closes the missing reviewer-instrument gap. The
same representative and adversarial fixtures produced ten system observations:
five paid Haiku 4.5 object-label drafts and five repeated pre-provider rights
refusals. Exact usage was 5,360 input and 1,043 output tokens ($0.010575 total;
$0.002237 maximum per paid run); maximum system latency was 3,409 ms. Provider
prompts now carry a bounded source note, source URL, and rights summary with
explicit heritage no-inference boundaries, and only `end_turn` plus non-empty
content counts as model output. A private randomized packet, separate key, and
two-reviewer worksheet bind SHA-256 `b389e698…d50b94c`; public v3 trial and
readiness receipts contain aggregates only. No reviewer evidence was invented,
so Calliope remains constrained pending two genuine independent returns.
The OpenAI reconciliation tiebreaker safety cycle is locally complete. Exact
evidence-equivalent ties now abstain before provider spend; provider responses
may abstain explicitly, cannot select outside the supplied candidate set, and
must return valid model/token metadata. Requests are capped at five seconds and
latency plus usage are retained on the review decision. The focused 14-test
reconciliation suite passes, and the canonical ten-observation worksheet is
prepared. A real adjudicated comparison remains blocked on an `OPENAI_API_KEY`
and independent judgments, so the surface remains unproven rather than promoted.
The reconciliation trial is now a v2 orchestration test rather than ten direct
provider calls. Five evidence-equivalent observations abstain before spend; only
five distinguishable observations may call OpenAI. Its comparator is a strong
identifier-aware deterministic evidence rule that explains its selection, so a
model cannot claim value merely by matching an obvious authority identifier.
Janus runtime reconciliation and diagnostic tools also emit instrumented
receipts. With no OpenAI key or independent review, the decision remains to keep
the model disabled unless blinded reviewer effort improves materially.
Janus no longer trusts model-written reconciliation explanations. OpenAI may
return only a confined candidate or abstention plus supplied evidence-field
names; local postflight proves those fields distinguish the candidate and renders
the explanation deterministically. Paid postflight failures preserve returned
usage, cost, latency, and model identity instead of disappearing as free errors.
Voyage embeddings now have a credible conventional comparator: normalized
lexical unigram/bigram feature hashing replaces the former SHA-byte vector, so
token overlap produces meaningful cosine behavior. Provider calls declare
document-retrieval intent, disable silent truncation, enforce a 2.5-second cap,
validate response indices, finite consistent dimensions, model identity, and
usage, then expose aggregate latency/token telemetry. Deterministic fallback no
longer consumes AI-call quota. The balanced worksheet is prepared, but no
Voyage credential or independent relevance judgments exist; the grade remains
unproven.
The Voyage lane now has an enforceable pre-provider egress boundary. Model use
requires exact public-catalog, public-safe, and cultural-care-review literals;
missing declarations and obvious email, labeled telephone, private/non-public,
or culturally restricted content return a structured refusal before network
access or AI quota evaluation. API-rate gating remains first. Clearly reviewed
public catalog text still reaches the bounded, metered provider path, while local
deterministic embeddings remain
available without egress. No credential is configured, so this is safety/readiness
evidence only and does not change the 0/5 AI-value grade.
The embedding lane now also has a strict executable provider trial: two held-out
ranking cases (representative and adversarial false friend), five repetitions per
case, paired query/document telemetry, dated list-price accounting, and private
randomized A/B review materials with an aggregate-only public receipt. The trial
launcher was exercised, created no evidence artifacts, and stopped before
network access because `VOYAGE_API_KEY` is absent. Next is a governed provider run followed by two
independent cultural-heritage relevance/effort reviews; until then the grade
remains 0/5.
Voyage trial v2 repairs the evaluation itself. Opaque document IDs remove answer
leakage, private packets include the query and corpus needed for genuine judgment,
BM25 replaces hashed cosine as the deterministic comparator, and parallel
provider calls are measured by wall-clock pair latency. The receipt adds mean
reciprocal rank and ranking-stability gates; runtime rejects zero, oversized, or
query/document-incompatible vectors and retains paid postflight usage/cost. The
canonical launcher still stops before artifact creation because no Voyage key is
configured, so no provider advantage is claimed.
Voyage trial v3 expands the preregistered held-out corpus from two cases to six:
three semantic representative queries and three adversarial museum-assertion
false friends, each repeated five times. The runner derives all dataset, call,
safety, and tool totals from the manifest, enforces the registered $0.01 pair-cost
ceiling, and emits a strict replayable v3 receipt. Trial readiness no longer
credits the Voyage filename alone; it verifies the receipt against the exact v3
manifest before reporting provider evidence. Provider and independent-review
evidence remain absent, so the grade is unchanged.
Visual similarity now fails closed at its actual SigLIP boundary. Only public
HTTPS image URLs may leave the app, and provider use now requires exact public-
image classification, rights review, provider-fetch permission, and cultural-
care review before AI quota evaluation or network access. Route and service
callers share that contract. Model output must return an identified model and
one unique, finite, range-valid score for every supplied candidate, with no
invented URLs. Calls are capped at 3.5 seconds and return latency/model/candidate
telemetry plus an advisory-only contract that visual resemblance cannot establish
identity, attribution, influence, or provenance. The no-service fallback is now
truthfully named IIIF derivative/topology ranking rather than a visual heuristic.
The balanced worksheet is prepared, but no configured service or independent
expert judgments exist, so the model remains constrained and unproven.
The visual lane now has an executable rights-documented trial over public-domain
Art Institute of Chicago records and IIIF images: representative and adversarial
false-friend cases run five times each against normalized metadata overlap, with
candidate confinement, top-one stability, provider-reported-cost completeness,
latency, and private randomized expert review. Missing provider cost remains null
and blocks the gate. The launcher stopped before network access and created no
artifacts because `SIGLIP_SERVICE_URL` is absent; next is a governed service run
and two independent cultural-heritage relevance/effort reviews. Grade remains
0/5.
Visual trial v2 replaces the two-case design that gave the deterministic baseline
10/10 and therefore made the registered quality lift impossible. Six official AIC
series comparisons now produce thirty observations; normalized title, creator,
medium, and subject metadata scores 25/30, so SigLIP must be perfect to clear the
10-point lift threshold. Before model spend, the runner replays all 24 allowlisted
AIC public-domain records with bounded redirect-free requests and retains only an
aggregate digest. The cost gate now matches the audit's $0.02 ceiling, all counts
are manifest-derived, and readiness strictly replays v2 receipts rather than
trusting a filename. Live rights replay passed 24/24; no SigLIP evidence exists.
Clio's social-editor cycle now separates deterministic safety from generative
copy judgment. A reproducible house-style preflight flags hype, unsupported
certainty, endorsement language, and excessive punctuation; provider egress
requires a bounded public-safe evidence attestation. Anthropic responses must
end normally, identify the returned model, include exact token usage, obey the
ten-second cap, preserve retained copy exactly, materially change revisions,
and introduce neither unapproved URLs nor deterministic safety failures. Model
self-scores were removed. The real balanced Haiku 4.5 trial completed ten calls
for $0.016060 total with 3.899-second maximum latency, five retains, five
adversarial revisions, and zero local boundary failures. Its private blinded A/B
packet still needs independent quality and reviewer-effort scoring, so Clio
remains unproven.
Agent execution now produces a persistent evidence loop rather than a persona-
only response. Every Clio, Mercator, Janus, Themis, and Calliope run writes the
same `runTelemetry` receipt into its organization-scoped AgentTask: execution
kind, requested/executed provider state, returned model/tokens, known Haiku cost
or honest unknown pricing, latency, actual tools, source count, HTTP(S) Linked
Open Data record IDs, and citation URLs. Deterministic and provider-fallback runs
retain zero model cost. This upgrades latency/cost evidence from missing to
partial without changing any AI-value grade; held-out aggregates and genuine
reviewer time remain required.
The agent workbench now closes the previously missing human-measurement handoff:
each run exposes task and telemetry SHA-256 targets, starts a reviewer timer, and
offers a strict pseudonymous scorecard for decisions, quality, tool usefulness,
Linked Open Data grounding, and bounded correction reasons. The separate
append-only receipt rejects digest drift, duplicates, unmeasured time, free text,
and incomplete declarations; it is organization-scoped and explicitly cannot
change the collection, publish output, or promote an agent. This is measurement
infrastructure, not a completed observation, so the 0/5 grade remains unchanged.
OpenAI reconciliation now has an executable promotion-candidate lane rather than
an unpriced optional call. Exact returned token totals must reconcile, known
GPT-4.1 mini snapshots receive a dated official USD basis, and unknown models
remain unpriced. A strict held-out manifest drives five repetitions over one
adjudicated selection and one ambiguity/abstention case against the stable
highest-score-or-abstain baseline, producing a private randomized A/B packet and
aggregate-only false-authority receipt. The attempted run stopped before network
access because no `OPENAI_API_KEY` exists; provider and review evidence is pending.
The independent-review handoff is now executable rather than a labeled
baseline/model spreadsheet. `audit:ai-agent-value:blind-review:prepare` binds a
private A/B packet into a strict blank worksheet while withholding the
randomization key. `audit:ai-agent-value:blind-review:aggregate` requires two
distinct pseudonymous reviewers, exact packet and candidate-text replay, eight
complete output-quality dimensions, pair preference, and directly observed
review seconds before unblinding. Its public receipt contains only content
hashes and aggregates. The real Clio packet has a ten-row private worksheet and
privacy-safe readiness receipt; no human scores were fabricated, so the grade
remains 0/5 pending genuine returns.
Independent review now has a usable offline interface rather than requiring raw
JSON editing. The packet-bound form preserves blinding and strict response replay
while supplying accessible labeled controls, anchored scores, live progress,
automatic per-pair timing, declarations, and private JSON export under a no-
network content-security policy. Browser verification completed the workflow,
found and repaired long-digest overflow, and proved a 360px layout with no
horizontal overflow or console errors. Future Calliope, reconciliation,
embeddings, visual-similarity, and Clio trial launchers generate the form beside
their private worksheet; genuine two-reviewer returns remain the next gate.
Calliope's evidence loop now distinguishes executed tools from declared
capabilities. Instrumented receipts expose provider, grounding, rights, quality,
SEO, analysis, and fallback outcomes with privacy-safe hashes and bounded
metrics; the workbench shows their status and attribution. Other agents retain
derived tool attribution as an explicit evidence gap rather than inheriting a
false claim of equivalent instrumentation.
history, including the milestones consumed by `/api/roadmap`.
artifact handoff.
contract.
command ownership and preferred entry points.
standards rounds and fixture anchors.
architecture and SOTA acceptance criteria.
- progress/era-history.md(progress/era-history.md): full Era A, B, and C slice
- roadmap-to-10.md(roadmap-to-10.md): executable strict-readiness checklist and
- risk-register.md(risk-register.md): open engineering and operating risks.
- ops/review-goals.md(ops/review-goals.md): review-goals policy and command
- ops/evidence-script-ownership.md(ops/evidence-script-ownership.md): evidence
- linked-art/LinkedArtModel1.0-Reference.md(linked-art/LinkedArtModel1.0-Reference.md):
- linked-art/LinkedArtSOTAWebApp.md(linked-art/LinkedArtSOTAWebApp.md): target
Visual-similarity trial v2 removes answer-bearing case names, item-label leaks,
and model/baseline score signatures from the private A/B review packet. The
metadata baseline normalizes simple plural variants, the aggregate manifest
digest binds the rights declaration, and paid SigLIP responses rejected after
receipt retain model/cost/latency telemetry through the API. No
`SIGLIP_SERVICE_URL` is configured, so no provider result or AI advantage is
claimed; next is a governed live run and two independent heritage reviews.
Mercator and Themis now report executed tools rather than persona-derived lists.
Mercator separately commits mapping-plan and MappingTemplate-validation results
with bounded column/error metrics; Themis commits rights summaries, provider
named-graph resolution, and final provenance review with known-rights and blocked
record counts. Janus and Calliope were already instrumented. This improves
reviewability without reclassifying deterministic agents as AI or claiming
benefit. Clio now closes the last runtime gap with separate instrumented receipts
for pattern analysis, conditional sparse-scope diagnosis, diagnostic synthesis,
and local drafting. All five agents distinguish observed execution from static
capability declarations; independent human value evidence is still required.
Clio social-editor trial v3 replaces the original straw-baseline packet with a
usable deterministic safe edit and gives both A/B sides identical original,
approved-source, and evidence context. Opaque case IDs and alternating placement
produce an exact 5/5 label balance. Ten real Haiku calls passed machine gates for
$0.015965 total with 4.603-second maximum latency; paid postflight failures now
retain model/tokens/latency. The old packet is superseded, and no benefit is
claimed until two independent reviewers score the current packet.
The shared independent-review aggregator now publishes privacy-safe per-dimension
model/baseline means and lifts, plus strict non-regression gates. A model cannot
pass by averaging a grounding or cultural-care loss against better prose scores.
Protected gates cover grounding, citation quality, calibration, robustness, and
cultural care; overall quality, pairwise preference, agreement, and review burden
remain visible, with all consequential authority still false.
Clio evidence-cited reader hook — provider delta ready, value unproven
The retired free-form social editor has been replaced experimentally by a
narrower `clio-evidence-hook` hypothesis. Haiku may reorder exact factual words
from bounded public-safe evidence into one or two concise sentences, but every
sentence must cite one or two known evidence IDs. Deterministic postflight rejects
unknown IDs, unsupported factual tokens, schema drift, unsafe evidence, and any
implied publication, attribution, or provenance authority. The comparator is a
strong deterministic concatenation of the same first two evidence statements.
V1 retained a $0.00052/1.199-second paid JSON-envelope failure; v2 retained a
$0.000515/1.128-second discourse-token mismatch. Although v3 completed, an
aggregate private-packet audit found that all five adversarial model hooks omitted
the required negated boundary while the deterministic baseline preserved it.
V3 is superseded and must not be assigned. Contract v2 marks fact versus boundary
evidence, requires every boundary ID, and requires explicit negation in every
boundary-citing sentence. V4 confirmed the repair; v5 adds public derived
grounding counts so readiness can replay 10 required boundary citations, 10
observed citations, and 10 preserved negations. V5 completed ten balanced calls;
all ten pairs differ from the deterministic baseline. Exact usage was 2,215 input
and 717 output tokens, $0.0058 total, $0.000617 maximum per run, and 1.388 seconds
maximum latency. The public receipt is semantically replayed
against both fixture hashes and recomputed pricing before readiness routes it to
review. No AI advantage is claimed: the private packet still requires two
independent blinded reviewers and must clear quality, grounding, citation,
calibration, robustness, cultural-care, cost, latency, variance, and reviewer-
effort gates.
The v5 assignment attempt exposed an orchestration mismatch: trial rows carried
extra private context fields rejected by the shared exact-schema parser. V6 moves
that context exclusively into the candidate envelope and emits the canonical
four-field packet row. Ten fresh calls retain 10/10/10 boundary counts, cost
$0.00581 total, and peak at 1.198 seconds. The v6 packet now has a privacy-safe arm-aware materiality receipt. It binds the
exact private packet, withheld randomization key, and public trial receipt by
SHA-256, then retains no candidate text or labels. All 10 pairs differ after
punctuation/case normalization, each case has one stable repeated model output,
mean token-set Jaccard is 0.942, and the model is 5.1% longer (143.5 versus 136.5
characters). This prevents cosmetic churn from being called reviewable while
making the unresolved tradeoff explicit; only independent reviewers can decide
whether the structural rewrite improves clarity enough to offset added length.
The shared assignment command accepts v6 and produced two distinct offline forms.
Their public receipt binds packet `1a89247b…b77cbcb` and both form digests while
retaining no reviewer codes, paths, text, labels, or identities. Readiness parses
that receipt against the v6 review-readiness receipt and routes to aggregation;
completed independent responses remain absent.
The canonical AI-value audit now consumes the already verified provider evidence
rather than displaying null machine metrics. Calliope v10 replays to 10 system
observations, $0.000979 mean cost, 747 ms mean latency, zero safety failures, and
passing machine gates. Clio evidence-hook v6 replays to 10 observations,
$0.000581, 953 ms, zero safety failures, and passing machine gates. Each is 60%
complete because blinded quality lift and reviewer-effort delta are still absent;
neither grade changes. Receipt absence yields no projection and receipt mutation
fails closed. Next remains two genuine independent returns per prepared surface.
AI evidence routing now uses `ai-agent-evidence-registry.ts` as a versioned leaf
contract for all six readiness surfaces. Readiness file discovery, assignment
bindings, provider verification, reconciliation source/retirement replay, and
audit machine-scorecard loading resolve canonical paths from that registry.
Focused drift tests prove unique surface/provider registrations and forbid copied
provider receipt literals in consumers. Canonical replay remains 4 provider
surfaces, 2 ready assignments, 2 retired behaviors, and 0 proven advantages.
The registry's five governed verification kinds now dispatch through one reusable
`ai-agent-evidence-verifiers.ts` table. Coverage validation fails on a missing,
duplicate, unknown, or unregistered verifier before any readiness projection.
Calliope, Clio hook, reconciliation/retirement, Voyage, and visual receipts retain
their existing exact parsers and fixtures. The CLI no longer owns family-specific
branches, while canonical replay remains 4/2/2/0 and grants no new authority.
Public assignment-receipt replay now joins the shared evidence verifier result.
Both exact assignment/readiness pairs verify canonically. Missing halves withhold
only assignment credit; a packet-hash mutation yields the bounded code
`ASSIGNMENT_RECEIPT_INVALID` while leaving provider evidence independently valid.
The CLI's separate assignment loop and binding export are removed. Readiness
remains 4 provider surfaces, 2 assignments, 2 retirements, and 0 proven benefits.
Unified readiness v3 now exposes privacy-safe `blockerCodes` per surface. The
canonical artifact contains six empty arrays. Mutation coverage proves an invalid
Clio assignment keeps provider evidence true, assignment readiness false, and
publishes only `ASSIGNMENT_RECEIPT_INVALID` plus generic prose—never packet hashes,
paths, parser text, candidate content, or reviewer metadata. The schema version
advances from v2 to v3; canonical counts and false authority remain unchanged.
Public review completion is now fail-closed: readiness semantically replays the
aggregate receipt and surface binding instead of trusting a filename. Forged
math, counts, gates, hashes, schema, or authority produce only
`REVIEW_RECEIPT_INVALID`; provider and assignment evidence remain independent.
The next external action is unchanged: obtain two genuine blinded returns for
Calliope v10 and Clio evidence-hook v6.
The public review verifier now binds completion to the full registered handoff:
verified assignment, exact readiness packet hash, surface, two-reviewer count, and
observation count. Readiness receipts themselves fail closed on scoring-dimension
or false-authority drift. This removes the remaining path for a structurally valid
but unrelated aggregate to receive completion credit.
Promotion evaluation now shares this verifier instead of calling the receipt
parser directly. The CLI rejects a wrong-packet aggregate before reading it as
subjective evidence or writing a candidate-comparison artifact. This closes the
downstream bypass while preserving private response-file handling.
Reviewer ergonomics now addresses the 10-pair navigation burden without reducing
evidence depth. Every pair announces complete/incomplete state, and a keyboard-
focusable control advances cyclically to unfinished pairs. New private Calliope
and Clio forms are bound by versioned v2 public assignment receipts. Browser
automation could not open the local `file://` artifact under its URL policy;
static accessibility contracts and end-to-end prepare/assign/aggregate tests pass.