← Documentation home

Canonical Markdown source · Oct 20, 2018

Risk Register

risk-register.md · 884 lines · SHA-256 eba6154c1dcc

Current risk posture

| Severity | Gap | Impact | Status |

|---|---|---|---|

| High | Complex provider and pipeline boundaries | Potential contract drift if adapter boundaries weaken | Closed — proven by RSI-1 on 2026-06-09 |

| Medium | UI journey automation scope | Hidden regressions outside current smoke paths | Closed — proven by RSI-2 on 2026-06-09 |

| Low | Single file and process complexity | Long-tail maintainability cost | Closed — proven by RSI-3 action 9 on 2026-06-09 |

| Medium | Chat grounding and citation enforcement | Ungrounded `/api/ai/chat` responses could leak hallucinated claims | Closed — proven by RSI-4 on 2026-06-09 |

| Medium | AI evidence freshness and citation drift | AI outputs can become stale or under-cited as retrieval/model behavior shifts | Closed — proven by RSI-5 Action 2 on 2026-06-09 |

| Medium | AI eval drift baselines omit citation freshness | Model/prompt regression gates could miss stale-evidence degradation | Closed — proven by RSI-6 on 2026-06-09 |

| Medium | AI eval artifact visibility and freshness-aging pressure | CI/pass status can hide stale-evidence aging pressure or make eval artifacts hard to inspect | Closed — proven by RSI-7 on 2026-06-09 |

| Medium | AI eval artifact dashboard visibility | Operators still need an in-app view of latest/trend eval artifacts and aging alerts | Closed — proven by RSI-8 on 2026-06-09 |

| Medium | AI eval artifact diff review speed | Operators can miss meaningful latest-vs-previous eval changes when raw artifacts are inspected manually | Closed — proven by RSI-9 on 2026-06-09 |

| Medium | AI eval diff prioritization | Raw deltas can obscure review priority when metric drift and evidence-aging pressure compete | Closed — proven by RSI-10 on 2026-06-09 |

| Medium | AI eval priority discoverability | Review severity can drift from policy or remain hidden unless operators open raw artifacts | Closed — proven by RSI-11 on 2026-06-09 |

| Medium | AI eval agent-summary reliability | Agents need validated policy and compact JSON/history instead of scraping UI or raw artifacts | Closed — proven by RSI-12 on 2026-06-09 |

| Medium | AI eval contract and CI annotation reliability | Agents and PR reviewers need schema-backed summary payloads, configurable history windows, and CI-visible priority warnings | Closed — proven by RSI-13 on 2026-06-09 |

| Medium | AI eval version and retention hygiene | Agents need explicit summary-version negotiation and artifact retention must not accumulate orphaned run JSON | Closed — proven by RSI-14 on 2026-06-09 |

| Medium | AI eval migration and reporting visibility | Future summary migrations, retention dry-runs, and CI annotation/pruning status need explicit reviewer evidence | Closed — proven by RSI-15 on 2026-06-09 |

| Medium | AI eval summary artifact/schema visibility | CI summary markdown, agent pruning status, and future migration compatibility must stay contract-visible | Closed — proven by RSI-16 on 2026-06-09 |

| Medium | Visual ETL Mapper AI-assist safety | LLM-assisted mappings could invent unsafe Linked Art paths unless suggestions are contract-validated and review-only | Closed — proven by RSI-17 on 2026-06-09 |

| Medium | Mapper-assist fixture/schema/importability drift | Tricky columns, UI draft import, or OpenAPI response docs could drift after initial mapper-assist launch | Closed — proven by RSI-18 on 2026-06-09 |

| Medium | Mapper-assist provider/browser/request-schema drift | Provider-specific columns, browser import flows, or request-body docs could drift from the mapper-assist contract | Closed — proven by RSI-19 on 2026-06-09 |

| Medium | Mapper-assist near-miss/docs/visual drift | Almost-mappable columns, imported-draft visuals, or human docs examples could drift from safe mapper behavior | Closed — proven by RSI-20 on 2026-06-09 |

| Medium | Mapper-assist layout/example/confidence drift | UI overlap, duplicated docs examples, or low-confidence suggestions could reduce curator trust | Closed — proven by RSI-21 on 2026-06-09 |

| Medium | Public source narrative and trust-page drift | Stale provider stats, untracked prototype assets, or placeholder legal copy could mislead humans and agents | Closed — proven by RSI-22 on 2026-06-09 |

| Medium | Public source summary and trust-smoke drift | Agents or reviewers could miss source/trust regressions if public pages lack contract JSON and screenshot proof | Closed — proven by RSI-23 on 2026-06-09 |

| Medium | Public source contract/artifact drift | Agents could miss public-source schema changes, copied assets could drift, or screenshots could accumulate without latest/previous review context | Closed — proven by RSI-24 on 2026-06-09 |

| Medium | Public trust docs and CI artifact visibility drift | Humans or agents could miss public-source examples, screenshot diffs, or CI artifact links during review | Closed — proven by RSI-25 on 2026-06-09 |

| Medium | Public trust pixel-diff and artifact API drift | Meaningful visual drift or missing CI/API artifact context could escape public trust review | Closed — proven by RSI-26 on 2026-06-09 |

| Medium | Public trust threshold/docs badge drift | Overbroad visual thresholds, undocumented trust examples, or unstable CI summaries could weaken public trust review | Closed — proven by RSI-27 on 2026-06-09 |

| Medium | Public trust policy/API/annotation drift | Hardcoded visual policy, hidden applied thresholds, or silent under-threshold changes could weaken trust review | Closed — proven by RSI-28 on 2026-06-09 |

| Medium | Public trust rationale/schema/summary drift | Reviewer rationale could be hidden, future policy schemas could be consumed unsafely, or CI summaries could obscure severity | Closed — proven by RSI-29 on 2026-06-09 |

| Medium | Public trust ownership/migration/annotation drift | Review ownership could be hidden, future schema migration could lack a fixture, or warning text could drift silently | Closed — proven by RSI-30 on 2026-06-09 |

| Medium | Pre-revenue SaaS path ambiguity | Buyers/operators could mistake a built pilot offer for paid-pilot or in-app billing readiness | Mitigated — `/pilot` now exposes commercial readiness: 0 paid pilots, manual invoice only, and in-app billing not built |

| High | Large API surface drift | Permission, cache, or storage-scope drift could hide across 146 route files, 61 route families, and 288 exported methods | Mitigated — `operations-risk-controls` ties route metrics and first-level route-family classification to auth, org-scope, and storage-write regression guards |

| Medium | Script proliferation ownership drift | 131 package aliases across 42 command namespaces could become an unowned parallel operations app | Mitigated — `operations-risk-controls` requires every package-script namespace to stay classified with owner/pruning guidance and every high-churn evidence alias to stay covered by the evidence ownership registry |

| Medium | Mixed runtime topology parity drift | Vercel, Neon, Render, Redis, optional Solr/GraphDB, and cron drains can diverge from local assumptions | Mitigated — `operations-risk-controls` requires deployment preflight, projection readiness, and disabled-system drill evidence |

| Medium | Evidence artifacts as product state | Readiness UI and launch claims could consume stale, unvalidated, or unretained artifacts | Mitigated — `operations-risk-controls` requires schema, freshness, and latest/previous retention evidence |

| High | Cold record reads and SLO history depth | Prior cold-read p95 misses or clustered samples could overstate performance readiness | Mitigated — `performance-scalability-controls` proves cold-record p95 misses fail and 30 distinct SLO days are required |

| Medium | Projection scaling enablement drift | Solr/GraphDB could be enabled too early or without scheduler readiness | Mitigated — `performance-scalability-controls` exercises projection thresholds and scheduled-drain blockers |

| Medium | Cron drain backlog growth | Bounded Vercel Cron batches could hide sustained backlog that needs dedicated workers | Mitigated — `performance-scalability-controls` exercises worker lag states and documents Render background-worker escalation |

| Medium | External provider/API dependence | Provider latency, rate limits, or schema drift could escape adapter cache/limit assumptions | Mitigated — `performance-scalability-controls` requires provider cache policy, paging limits, and validation-drift evidence |

| High | Secrets hygiene recurrence | Rotated local/Vercel secrets could regress if sensitive `.env*` files become tracked or historical again | Mitigated — `security-reliability-controls` checks git tracking/history, `.gitignore`, and launch preflight secret evidence |

| High | Production test-token exposure | Staging role/test override tokens could accidentally be present in public production | Mitigated — `security-reliability-controls` proves runtime production overrides return null and preflight fails token presence |

| High | Org scope boundary drift | New scoped route families could forget membership-validated storage selection | Mitigated — `security-reliability-controls` joins request-scope hardening to the org-scope route matrix and regression tests |

| Medium | Cron authorization drift | Enabled projection/publication drains could run without `CRON_SECRET` if worker flags expand | Mitigated — `security-reliability-controls` proves outbox and publish cron routes fail closed before work runs |

| High | Complete onboarding coverage drift | Sign-in, org selection, import, review, or publication boundaries could regress independently | Mitigated — `testing-gap-controls` verifies the full onboarding coverage chain stays represented in tests |

| Medium | Disabled feature drill drift | AG2, Solr/GraphDB, projection, cron, or publication paths could be enabled without current staging rehearsal | Mitigated — `testing-gap-controls` executes the disabled-system drill and requires the 9-check non-production rehearsal |

| Medium | TypeScript command confusion | Developers could treat direct `tsc --noEmit` as the release gate instead of the CI-aligned command | Mitigated — `testing-gap-controls` locks `pnpm typecheck` to `pnpm build`, and direct `pnpm typecheck:diagnostic` now converges with a 0-error parity report |

| High | Storage-scope route family coverage drift | Records, annotations, wiki drafts, AI/editorial reads, AgentTasks, or org admin APIs could lose matrix coverage | Mitigated — `testing-gap-controls` requires matrix depth and route-family regression files across the scoped surface |

| High | Local-vs-external evidence collapse | Local green gates could be mistaken for long-window SLO, adoption, paid-pilot, retention, or margin proof | Open strict blocker — `/readiness` now exposes 5 real-world evidence contracts that reject local substitutes, depend on source plus acceptance-ledger rows, and can downgrade any global strict-gate pass until real SLO, adoption, paid-pilot, retention, and margin evidence lands |

| Medium | Organic campaign attribution continuity | UTM context could disappear after the landing-page navigation or invalid landing slugs could split the measured funnel | Mitigated — consented attribution is sanitized and retained only for the browser session; generated landing paths are registry-checked and the four-week packet remains content-addressed and approval-gated |

| High | AI branding without comparative value proof | Deterministic heuristics, wrappers, self-scored fixtures, or unbalanced repeated observations could be presented as agent intelligence while adding cost and review burden without measurable benefit | Open — all 13 surfaces now have explicit baselines, owners, acceptance criteria, commands, blockers, next experiments, scorecards, and dispositions; 0 AI advantages are proven and five model-backed workflows remain constrained. Promotion has hard privacy/security, authorization, tool-discipline, safety, complete case/run balance, duplicate rejection, and within-case repeatability gates; baseline violations and malformed/non-finite metrics fail closed. Version-2 receipts bind exact inputs but remain non-authoritative candidates pending external-evidence acceptance; synthetic canonical receipts are forbidden. Local authenticated dashboard proof passes desktop/mobile semantics, diagnostics, and overflow after two wrapping fixes. Real provider/reviewer observations still require closure. |

| High | Site-wide route coverage drift | New pages, handlers, private operator surfaces, or noindex workflows could bypass browser and authorization audits | Mitigated — the generated inventory classifies all 77 pages and 190 handlers/351 methods, requires zero unassigned files, and binds all 15 protected plus 53 public/noindex pages to explicit audit lanes |

Calliope risk update (2026-08-24): the second held-out trial closes the observed

malformed-rights spend and 12-second latency defects locally (5/5 pre-model

adversarial refusals; 3,920 ms maximum; $0.010505 total provider cost). The AI-

value risk remains open because independent blinded quality and reviewer-effort

evidence is still absent.

Calliope packet/readiness update (2026-08-24): the v3 rerun retains an actual

private A/B packet, separate randomization key, and ten-row worksheet for two

independent reviewers. Five paid representative calls passed cost/latency gates

($0.010575 total; 3,409 ms maximum), and five malformed-rights observations again

refused before spend. Bounded prompts now include source note/URL/rights context

and explicit heritage no-inference rules; truncated or empty responses fail to

fallback. Risk stays open because model drafts are substantially longer than

the deterministic label and genuine quality/reviewer-effort returns are absent.

Calliope hybrid-grounding update (2026-08-25): `claim-evidence-v2` no longer

asks Haiku to perform brittle exact-copy work. The model proposes bounded claims

and source IDs; deterministic tooling selects exact Linked Art evidence lines

and enforces source confinement, minimum lexical support, URL/year checks, and

authority boundaries. v7/v8 now retain partial paid failure telemetry and safe

rejection categories instead of losing failed-session evidence. v9/v10 passed

the same contract across 10 paid calls with 1,991 ms worst latency, $0.00979

maximum trial cost, minimum 0.667 lexical support, full excerpt/source coverage,

and 10/10 rights refusals before spend. Comparative-value risk remains open

because independent blinded quality and reviewer-effort evidence is absent.

Reconciliation-model risk update (2026-08-24): true evidence-equivalent ties

now abstain before spend, model output is candidate-confined, timeout is capped

at five seconds, and exact usage/latency are retained. Comparative-value risk

remains open because no OpenAI credential or independent adjudication is

available; mock-provider boundary tests are not counted as model evidence.

Embeddings risk update (2026-08-24): the deterministic comparator is now lexical

and retrieval-relevant, while Voyage responses are bounded, metered, and checked

for order/dimension drift. Provider egress now fails before AI quota gating and

network access unless the caller explicitly attests public-catalog safety and a

cultural-care review; obvious direct identifiers, private/non-public markers,

and culturally restricted markers override that attestation. Value remains open

without a real Voyage run and independent relevance judgments, and the scanner

is defense in depth rather than a substitute for collection governance.

Configured credentials no longer activate Voyage by themselves: the route stays

deterministic unless the authenticated caller explicitly requests the model, and

an unavailable explicitly requested provider fails closed without AI spend.

Embedding evidence update (2026-08-24): a strict held-out runner now pairs

document/query Voyage calls across five repetitions of representative and

adversarial false-friend cases, records exact pair cost/latency and ranking

aggregates, and separates randomized review keys from aggregate public receipts.

Recognized `voyage-4-lite` responses use the dated official $0.02/M-token list

price without assuming free credits; unknown snapshots retain null cost. No

credential is configured, so no provider receipt or benefit claim exists.

Embedding trial v2 replaces answer-revealing document IDs and context-free review

rows with opaque IDs plus private query/corpus context. BM25 is the strong

deterministic comparator; provider query/document calls run in parallel with

wall-clock pair latency, cross-batch dimension/model checks, nonzero bounded

vectors, reciprocal-rank and stability gates, and retained spend on paid

postflight rejection. Provider and independent reviewer evidence remain absent.

Embedding trial v3 reduces evaluation overfitting and forged-readiness risk. The

registered corpus now balances three representative and three adversarial cases

over thirty observations; telemetry is manifest-derived, the actual $0.01 cost

gate is enforced, and readiness strictly replays the receipt/manifest binding.

The remaining high-risk gap is external: no Voyage-backed observation or two

independent cultural-heritage relevance/effort responses exist.

Visual-similarity evidence update (2026-08-25): the governed packet now uses

opaque case/item identifiers and rank-only candidates so expected labels and

model-specific score ranges are not exposed to reviewers. The receipt digest

binds the nonblank rights basis, and paid postflight rejections preserve model,

latency, and valid provider-reported cost. Risk remains open: repeated ranking

stability could still help a reviewer guess system identity, and no configured

SigLIP run or independent expert response exists.

Visual trial v2 removes a zero-headroom evaluation defect: the earlier metadata

baseline already achieved 10/10, so no model could meet the lift gate. The six-case

replacement has a strong 25/30 metadata baseline, forces a model to reach 30/30,

derives telemetry from the manifest, enforces the registered $0.02 ceiling, and

strictly replays its receipt. Rights literals are no longer trusted alone: 24

official AIC records replay before spend, and any stale/non-public-domain response

blocks all model calls. Remaining risk is the absent SigLIP run and independent

review, not local test readiness.

Visual-similarity risk update (2026-08-24): public-HTTPS preflight, strict

candidate confinement, score validation, timeout, and a discovery-only heritage

boundary now close the observed local failure modes. SigLIP egress also requires

exact rights-review, provider-fetch permission, and cultural-care approval before

AI quota evaluation or network access; the service enforces the same contract for

non-route callers. The risk stays open because no SigLIP service, expert relevance

judgments, infrastructure-cost evidence, or independently verified rights corpus

is currently available.

As with embeddings, configuration alone cannot activate SigLIP; explicit

per-request model opt-in is required, and unavailable opt-in fails closed.

Visual-evidence update (2026-08-24): the service now accepts only non-negative

provider-reported per-ranking cost with an exact supported basis; absent cost

stays null and fails the economic gate. A strict rights-documented held-out

runner compares ten repeated rankings against normalized metadata overlap,

measures false-friend selection and top-one stability, separates private image/

source review packets from aggregate receipts, and denies attribution authority.

The configured service is absent, so the launcher created no evidence artifacts

and the risk remains open pending a provider run and two independent reviews.

Clio social-editor risk update (2026-08-24): deterministic preflight, explicit

public-safe evidence attestation, strict retain/revise semantics, removal of

model self-scores, timeout, and exact token telemetry close the observed local

defects. The real ten-run trial passed cost, latency, safety, and authority gates;

net value remains open until independent blinded preference and reviewer-time

scores are imported.

Clio social-editor evidence update (2026-08-25): the original review packet is

superseded because it used a refusal-style straw baseline, exposed case roles,

omitted grounding context, and did not guarantee balanced label placement. V3

uses a safe deterministic edit, opaque IDs, identical source/evidence context,

exact 5/5 A/B placement, and retained paid-postflight telemetry. Ten genuine

calls passed machine gates; two independent responses remain missing, so benefit

and promotion stay unproven.

The standalone review command now also requires explicit `--use-model` opt-in

before draft file or provider access. Credential presence alone cannot activate

Clio; the governed trial remains the only intentionally model-backed batch lane.

Clio retirement update (2026-08-25): v4/v5 retained paid failures showing that

free-form revisions replaced prohibited certainty with unsupported speculation.

`clio-evidence-v1` now checks only vocabulary introduced beyond the original

draft, and unsafe drafts receive the deterministic safe edit before provider

spend. v6/v7 passed 10 paid safe-draft reviews plus 10 deterministic adversarial

edits, but all 20 blinded candidate pairs were byte-identical to the baseline.

The aggregate equivalence receipt therefore retires the current model behavior

and suppresses pointless human review. Further Clio spend requires a materially

different registered capability that first proves non-identical output.

Agent-trial orchestration update (2026-08-25): one aggregate-only readiness

receipt now maps all five model surfaces to provider configuration, governed

trial, independent review, or comparison. It records configuration presence but

never values, private paths, candidate text, or reviewer identity, and grants no

provider-spend or promotion authority. External providers and genuine reviewers

remain outstanding rather than being obscured by scattered setup instructions.

Calliope reviewer-handoff update (2026-08-25): the assignment command validates

both fixed pseudonymous codes before constructing filenames, writes exactly two

exclusive offline forms, and emits a separate aggregate-only receipt with the

packet hash, two form hashes, counts, and false completion/deployment/publication

authority. Readiness v2 detects that receipt and routes to aggregation; it does

not mistake prepared forms for completed independent review.

Calliope provider-replay update (2026-08-25): unified readiness no longer trusts

the v10 filename. It reconstructs both held-out fixture hashes and strictly parses

the receipt's dataset, recomputed token cost, latency, adversarial rights refusal,

grounding, blocker, and false-authority fields. Forged or drifted evidence cannot

unlock reviewer aggregation. Clio and historical reconciliation provider/retirement

receipts remain the next filename-authority hardening targets.

Reconciliation model retirement update (2026-08-25): the tiebreaker provider

boundary now accepts OpenAI or Anthropic while preserving the same deterministic

equivalence preflight, candidate confinement, exact telemetry, cost ceiling, and

false merge authority. Two Haiku sessions retained 10 paid representative calls

and 10 no-call equivalence abstentions. The deterministic identifier-aware

baseline scored 20/20; Haiku scored 10/20 by abstaining on all paid cases, after

an earlier response also failed schema postflight. The aggregate receipt retires

this behavior and blocks human review/further spend until a future hypothesis

uses adjudicated cases that deterministic evidence genuinely cannot resolve.

The follow-up failure-receipt repair closes the observed evidence-loss path:

paid postflight failures now retain privacy-safe partial provider telemetry and

progress through an exclusive public receipt. Provider/usage failures keep cost

null when telemetry is incomplete, raw error text is never retained, and the

receipt cannot authorize retry or identity merge.

Runtime retirement is now enforced: `/curator/reconciliation` contains no model

adapter activation path and the legacy flag cannot restore one. New trial spend

requires a registered v3 manifest plus a separate digest-bound two-reviewer

adjudication receipt. Eligibility rejects representative cases the

identifier-aware baseline can already solve and adversarial cases that are not

evidence-equivalent, preventing another weak-baseline trial by construction.

Reconciliation retirement-replay update (2026-08-25): both historical provider

receipts now bind the exact frozen cases and pass strict schema, usage, latency,

safety, tool, blocker, and authority parsing. Readiness recomputes the 20-observation

retirement aggregate and requires deterministic correctness and representative

selection to strictly dominate the model before exposing hypothesis revision.

Forged provider or retirement filenames no longer carry authority.

Blind-review ergonomics update (2026-08-24): reviewers no longer need to edit

private JSON by hand. A packet-bound offline HTML form provides labeled native

controls, anchored scores, automatic pair timing, completeness validation, exact

response export, and no network access. Browser proof repaired mobile digest

overflow and verified a 360px layout without console errors. Candidate text and

responses remain private, the key stays separate, and no human result is implied.

Blind-review gate update (2026-08-25): aggregate overall quality can no longer

hide a grounding, citation, calibration, robustness, or cultural-care regression.

The authoritative comparison evaluator now requires the parsed packet-bound

aggregate receipt from two or more independent reviewers, verifies full

observation coverage and surface identity, and rejects worksheet declarations

without that receipt; genuine completed human evidence is still outstanding.

Public receipts retain only per-dimension means/lifts, regressed dimension names,

and non-regression booleans alongside existing hashes and counts. Promotion still

requires separate accepted provider, cost, latency, variance, effort, and safety

evidence; the aggregate itself grants no authority.

Agent execution-loop update (2026-08-24): all five `/api/agents/run` workflows

now persist one shared telemetry receipt inside their organization-scoped,

approval-required `AgentTask`. It records actual tools, execution kind, latency,

requested/executed provider state, returned model/tokens, zero-cost fallback,

Haiku cost estimates, source-record scope, HTTP(S) Linked Open Data IDs, and

citation URLs. Unknown model pricing stays `null`. This closes invisible tool/

grounding/spend drift, but comparative aggregates and observed human reviewer

time remain open and the receipt grants no publication or promotion authority.

Reconciliation trial update (2026-08-24): GPT-4.1 mini tiebreaks now reject

inconsistent token totals, price only recognized returned snapshots under a

dated official basis, and retain `null` for unknown-model cost. A strict held-out

two-case/five-repeat harness candidate-confines decisions, checks adjudicated

selection and adversarial abstention, alternates blind labels, separates its key,

and exposes only aggregate false-authority receipts. Provider evidence is absent

because `OPENAI_API_KEY` is not configured; synthetic tests are not benefit proof.

The workbench now starts a review timer on receipt of a run and stores only a

pseudonymous, bounded scorecard in a separate append-only organization-scoped

file. Exact task/telemetry digests, strict no-free-text parsing, required privacy

declarations, duplicate rejection, and false collection/publication/promotion

authority reduce review-custody risk. They do not prove reviewer savings until

real independent reviews are returned and aggregated against the baseline.

Blind-review custody update (2026-08-24): the former baseline/model-labeled

score sheet is no longer the subjective judgment surface. A private packet-bound

A/B worksheet with a separately held key now requires two distinct pseudonymous

independent reviewers, exact candidate replay, eight complete quality dimensions,

preferences, and measured effort. Public readiness/aggregate artifacts exclude

candidate text, labels, reviewer codes, identities, and keys. The risk remains

open because genuine reviewers have not returned completed worksheets.

Use `docs/closeout-notes.md`(docs/closeout-notes.md) for the ready-to-fill one-click RSI closeout block each cycle.

Remediation slice queue

```text

## [ ] RSI-5 closeout (ready-to-fill)

```

  • [x] RSI-1: Provider/pipeline boundary drift hardening (High-severity remediation) — closed (proven)
  • Owner: Platform
  • Scope: keep adapter boundaries enforceable by contract tests + deterministic governance checks.
  • Evidence before close-out (must all pass before slice status can move to done):
  • `tests/contracts/provider-boundary-contracts.test.ts` passes and has no unexpected allowed-import exceptions.
  • `src/adapters/*.ts` contains no new cross-adapter imports beyond `provider-interface`, `adapter-utils`, or `expansion-provider`, with any change to this exception set documented in this register.
  • At least one fresh `pnpm test` run includes the new boundary contract test.
  • This row is reflected in the same-cycle `CLAUDE.md` close-out row and synchronized to `README.md` and `docs/roadmap.md` before the next RSI expansion scope.
  • Closed proof (2026-06-09):
  • `pnpm test`, `pnpm lint`, and `pnpm build` all passed in the same cycle.
  • Full close-out sync completed in `CLAUDE.md`, `README.md`, and `docs/roadmap.md`.
  • [x] RSI-2: UI journey automation breadth (Medium-severity remediation) — closed (proven)
  • Owner: Product + Platform
  • Scope: remove hidden journey blind spots by broadening journey smoke automation to role+provider matrix plus entity-role read-path assertions.
  • Evidence before close-out:
  • Smoke probe includes all matrix scenarios (`public` and `researcher` × `met` and `getty`) in `scripts/smoke-explore-import-matrix.ts`.
  • Probe includes `/api/objects/[id]`, `/api/works/[id]`, `/api/agents/[id]`, `/api/places/[id]`, and `/api/sets/[id]` happy/404 route checks for imported records.
  • Slice references `pnpm smoke:explore:matrix` and shares one-liner proof path in `package.json`.
  • Close-out row is added and this risk row is synchronized to `README.md` and `docs/roadmap.md`.
  • Closed proof (2026-06-09):
  • Proof references now include `scripts/smoke-explore-import-matrix.ts` and `package.json` (`pnpm smoke:explore:matrix`) covering role/provider matrix plus route-assertion checks.
  • `docs/roadmap.md` and `README.md` updated to mark the journey automation gap as closed with RSI-2 proof.
  • `CLAUDE.md` close-out log row added.
  • [x] RSI-3: Single-file and process-complexity reduction (Low-severity remediation) — closed (proven)
  • Owner: Product + Platform
  • Scope: reduce long-tail maintenance cost by splitting oversized files, documenting process ownership, and making future complexity boundaries explicit.
  • Status update (2026-06-09): Action 1 complete — candidate-file complexity inventory captured with owners and decomposition targets.
  • Status update (2026-06-09): Action 2 complete — publish-queue worker refactor slice landed and proven in `src/services/publish-queue-*.ts` with full gate verification.
  • Status update (2026-06-09): Action 3 complete — `src/services/issues.ts` split into `src/services/issues/{cache.ts,github.ts,analysis.ts,types.ts}` and proven with full gate evidence.
  • Status update (2026-06-09): Action 4 complete — `src/services/outbox.ts` split into focused modules under `src/services/outbox/`.
  • Status update (2026-06-09): Action 5 complete — `src/services/reconciliation.ts` split into focused modules with preserved behavior and passing gate suite.
  • Status update (2026-06-09): Action 6 complete — `src/services/wiki-publish.ts` split into focused modules with behavior preserved and full gate evidence (`pnpm test`, `pnpm lint`, and `pnpm build`).
  • Evidence required before close-out:
  • Identify and decompose top 8 file classes with sustained complexity symptoms (e.g., route handlers > ~400 LOC, scripts with multi-step orchestration, services with mixed responsibilities).
  • Add a short decomposition plan and owner map in the risk row before implementation.
  • New or refactored files must keep `pnpm test` and `pnpm lint` green with existing RSI checks.
  • Add a close-out row in `CLAUDE.md` plus updates in `docs/roadmap.md` and `README.md`.
  • Candidate inventory (Action 1; baseline from code-size scan):
  • `src/services/publish-queue-worker.ts` (679 LOC) — Owner: Product + Platform; target: split worker orchestration, provider adapters, and persistence transaction helpers.
  • `src/services/issues.ts` (670 LOC) — Owner: Product + Platform; target: split cache hydration, GitHub event fetch, and SSE/event-bus publishing modules.
  • `src/services/outbox.ts` (669 LOC) — Owner: Platform; target: split status/query read-models, projector/retry orchestration, and DLQ policy helpers.
  • `src/services/reconciliation.ts` (603 LOC) — Owner: Product + Platform; target: split candidate scoring, queue orchestration, and merge policy enforcement.
  • `scripts/authority-cache-refresh.ts` (483 LOC) — Owner: Platform; target: split CLI orchestration, external authority fetch, and cache persistence/error retry stages.
  • `src/services/ai-layer.ts` (468 LOC) — Owner: Product + Platform; target: split prompt/command validation, route-level orchestration, and response shaping/metrics.
  • Proof plan for next close-out:
  • Action-1 inventory + owner map is now captured in this register.
  • Action-2 completion proof (2026-06-09):
  • `pnpm test`, `pnpm lint`, and `pnpm build` pass after the publish-queue refactor slice landed.
  • `tests/services/publish-queue-worker.test.ts` proves queue orchestration behavior and daily-cap deferral behavior remains unchanged.
  • `CLAUDE.md`, `README.md`, and `docs/roadmap.md` are updated in the same slice closeout package.
  • Action-3 completion proof (2026-06-09):
  • `tests/services/issues.test.ts`, `tests/api/issues.test.ts`, `tests/api/issues-stream.test.ts`, and `tests/api/issues-webhook.test.ts` pass.
  • `pnpm test`, `pnpm lint`, and `pnpm build` pass in the same cycle.
  • `CLAUDE.md`, `README.md`, and `docs/roadmap.md` are updated in the same slice closeout package.
  • Action-4 completion proof (2026-06-09):
  • `src/services/outbox.ts` split into `src/services/outbox/{commands.ts,payload.ts,pool.ts,policy.ts,types.ts}` with all `outbox` behavior preserved.
  • `tests/services/outbox.test.ts`, `tests/services/outbox-projector.test.ts`, and `tests/services/outbox-alerts.test.ts` pass.
  • `tests/services/outbox.test.ts` covers payload and failure-policy behavior; `tests/services/outbox-projector.test.ts` and `tests/services/outbox-alerts.test.ts` remain stable with unchanged exported contracts.
  • `pnpm test`, `pnpm lint`, and `pnpm build` pass in the same cycle.
  • `CLAUDE.md`, `README.md`, and `docs/roadmap.md` are updated in the same slice closeout package.
  • Action-5 completion proof (2026-06-09):
  • `src/services/reconciliation.ts` split into `src/services/reconciliation/{candidates.ts,scoring.ts,service.ts,thresholds.ts,tiebreaker.ts,types.ts,utils.ts}` while keeping public API stable (`ReconciliationCandidate` remains exported).
  • `tests/quality/reconciliation-exhibitions-literature.test.ts` pass with unchanged fixture coverage and output shape validation.
  • `pnpm test`, `pnpm lint`, and `pnpm build` pass in the same cycle.
  • `CLAUDE.md`, `README.md`, and `docs/roadmap.md` are updated in the same slice closeout package.
  • Action-6 completion proof (2026-06-09):
  • `src/services/wiki-publish.ts` split into `src/services/wiki-publish/{client.ts,draft.ts,plan.ts,preflight.ts,publish.ts,types.ts,utils.ts}` with preserved public exports and test behavior.
  • `tests/services/wiki-publish.test.ts` passes (and no wiki publish integration behavior changed).
  • `pnpm test`, `pnpm lint`, and `pnpm build` pass in the same cycle.
  • `CLAUDE.md`, `README.md`, and `docs/roadmap.md` are updated in the same slice closeout package.
  • Action-7 completion proof (2026-06-09):
  • `src/services/monitoring-telemetry.ts` decomposed into `src/services/monitoring-telemetry/{service.ts,uptime.ts,kpis.ts,io.ts,utils.ts,types.ts}` with public API stability and all existing callers preserved (`tests/services/monitoring-telemetry.test.ts`, `scripts/monitoring-telemetry-sync.ts`).
  • `pnpm test`, `pnpm lint`, and `pnpm build` pass in the same cycle (monitoring slice regression check validated via `tests/services/monitoring-telemetry.test.ts`).
  • `CLAUDE.md`, `README.md`, and `docs/roadmap.md` are updated in the same slice closeout package.
  • Action-8 completion proof (2026-06-08):
  • `scripts/authority-cache-refresh.ts` split into `scripts/authority-cache-refresh/{cli.ts,config.ts,io.ts,parser.ts,refresh.ts,report.ts,runner.ts,types.ts}` with behavior preserved from single-file logic.
  • `pnpm test` (full suite), `pnpm lint`, and `pnpm build` pass in the same cycle.
  • `CLAUDE.md`, `README.md`, and `docs/roadmap.md` are updated in the same slice closeout package.
  • Action-9 completion proof (2026-06-09):
  • `src/services/ai-layer.ts` decomposed into `src/services/ai-layer/{build.ts,embed.ts,persist.ts,pool.ts,types.ts,utils.ts,constants.ts,visual.ts}` with public API preserved via facade export.
  • `pnpm test`, `pnpm lint`, and `pnpm build` pass in the same cycle after `src/services/ai-layer/utils.ts` type hardening.
  • `CLAUDE.md`, `README.md`, and `docs/roadmap.md` are updated in the same slice closeout package.
  • [x] RSI-4: Chat grounding and citation enforcement (Medium-severity remediation) — closed (proven)
  • Owner: Platform + AI Reliability
  • Scope: enforce Graph-RAG chat grounding (`/api/ai/chat`) with mandatory sentence-level citations and automatic refusal on insufficient coverage.
  • Evidence before close-out:
  • `src/services/ai-chat.ts` and `app/api/ai/chat/route.ts` implement citation scoring, per-sentence citation grouping, and refusal responses with evidence metadata.
  • `tests/api/ai-chat.test.ts` covers invalid input, refusal path, and successful cited answer behavior.
  • Closed proof (2026-06-09):
  • `tests/api/ai-chat.test.ts` passes in-cycle.
  • `pnpm test`, `pnpm lint`, and `pnpm build` all pass in the same closeout cycle.
  • `CLAUDE.md`, `README.md`, and this roadmap are updated in the same slice closeout package.
  • [x] RSI-5: AI evidence drift + citation freshness (Medium-severity remediation) — closed (proven)
  • Owner: Platform + AI Reliability
  • Status update (2026-06-09): Action 1 complete — query-route citation metadata and refusal fields are implemented in `src/services/ai-query.ts`, with matching assertions in `tests/api/ai-query.test.ts` and `tests/quality/cite-or-refuse-conformance.test.ts`.
  • Status update (2026-06-09): Action 2 complete — shared citation freshness metadata now adds `retrievedAt` and `citationFreshness` diagnostics to `/api/ai/query` and `/api/ai/chat`, refusing stale evidence through policy-backed tests.
  • Scope:
  • Prevent AI output quality regressions by enforcing citation coverage in broader AI workflows beyond `/api/ai/chat`, especially `/api/ai/query`.
  • Ensure every AI answer path emits stable evidence metadata (`entityId`, `propertyPath`, `sourceUrl`, `retrievedAt`) and explicit refusal when coverage or freshness is below policy.
  • Track this as a recurring AI-RSI checkpoint in the same evidence loop as prior RSI slices.
  • Acceptance:
  • `/api/ai/query` returns structured citation metadata on success and refusal paths.
  • Any under-cited output returns refusal state with explicit citation coverage failure reason.
  • Any stale cited output returns refusal state with explicit citation freshness failure reason and `citationFreshness` diagnostics.
  • Query tests prove: invalid input handling, refusal path behavior, and cited success path.
  • `/api/ai/chat` keeps grounded response semantics while gaining the same freshness refusal guard.
  • Evidence plan (before close-out):
  • Add/extend `tests/api/ai-query.test.ts` with cite-or-refuse assertions. Action 1 complete.
  • Add/extend `tests/quality/cite-or-refuse-conformance.test.ts` to compare generated-content and `/api/ai/query` response behavior. Action 1 complete.
  • Add/extend `tests/api/ai-chat.test.ts` with stale-citation refusal assertions. Action 2 complete.
  • Run `pnpm test`, `pnpm lint`, and `pnpm build` in the same closeout cycle. Closeout gate.
  • Mirror closeout evidence in `README.md`, `docs/roadmap.md`, and `CLAUDE.md`. Closeout gate.
  • One-click closeout block template (ready-to-fill):
  • Slice: Platform + AI
  • Severity: Medium
  • Action: RSI-5 / AI evidence drift + citation freshness
  • Date: 2026-06-09
  • Status: DONE
  • Evidence path:
  • [ ] `tests/api/ai-query.test.ts`
  • [ ] `tests/api/ai-chat.test.ts`
  • [ ] `tests/quality/cite-or-refuse-conformance.test.ts`
  • [ ] `pnpm test` (full suite)
  • [ ] `pnpm lint`
  • [ ] `pnpm build`
  • [ ] Closeout row sync: `docs/risk-register.md`
  • [ ] Closeout row sync: `docs/roadmap.md`
  • [ ] Closeout row sync: `README.md`
  • [ ] Closeout log sync: `CLAUDE.md`
  • Compounding proof:
  • [ ] Risk status updated to closed/proven with “proven” date
  • [ ] Explicit what changed / evidence / next action present
  • [ ] No active RSI-5 TODOs from this slice remain
  • [ ] Existing AI-RSI evidence loop (docs/rsi-wiki.md) references the slice outcome
  • Submit now:
  • [ ] Copy/paste this block into `docs/roadmap.md`, `CLAUDE.md`, and `README.md` as final closeout package
  • [x] RSI-6: AI eval drift baselines include citation freshness (Medium-severity remediation) — closed (proven)
  • Owner: Platform + AI Reliability
  • Scope: make AI eval regression baselines compound on RSI-5 by scoring and fail-fast gating citation freshness drift, not only citation structure/accuracy.
  • Acceptance:
  • `EvalScoreThresholds` includes `citationFreshness` for current metrics, baselines, artifact summaries, and drift comparisons.
  • `config/ai-eval-regression-policy.json` defines `citationFreshness`, `maxCitationFreshnessDrop`, and `failFastOnCitationFreshnessDrift`.
  • Golden eval dataset rubric includes `citationFreshnessThreshold = 0.95`.
  • Stale evidence is proven to lower eval freshness score and fail the sample gate.
  • Live eval gate reports zero drift for the current baseline identity.
  • Closed proof (2026-06-09):
  • `src/services/ai-eval-harness.ts` scores freshness from actual citation `retrievedAt` / source timestamps.
  • `src/services/ai-eval-regression.ts` compares `citationFreshnessDrop` and fail-fast policy.
  • `scripts/ai-eval-gate.ts` persists and prints `citationFreshness` metrics/drift.
  • Tests: `tests/services/ai-eval-harness.test.ts`, `tests/services/ai-eval-regression.test.ts`, `tests/services/ai-eval-artifacts.test.ts`, and `tests/quality/ai-eval-golden-dataset.test.ts`.
  • Gate: direct AI eval check passed with `citationFreshness=1` and `citationFreshnessDrop=0`.
  • [x] RSI-7: AI eval summary badges + aging-pressure alerting (Medium-severity remediation) — closed (proven)
  • Owner: Platform + AI Reliability
  • Scope: make eval freshness status visible in CI artifacts and warn when cited evidence ages toward the policy window even while `citationFreshness` remains flat/green.
  • Acceptance:
  • `artifacts/evals/summary.md` renders badges for status, faithfulness, relevance, citation accuracy, `citationFreshness`, and pass rate.
  • Summary includes latest/run/trend artifact links plus citation freshness aging (`oldestAgeDays`, `oldestAgeRatio`, `maxAgeDays`).
  • CI appends the summary to `$GITHUB_STEP_SUMMARY` and uploads `artifacts/evals/`.
  • Trend alert emits `citation_freshness_aging_pressure` when freshness is flat/improved but `oldestAgeRatio` crosses the threshold.
  • Closed proof (2026-06-09):
  • `.github/workflows/ai-eval-gate.yml` publishes and uploads AI eval artifacts.
  • `src/services/ai-eval-artifacts.ts` builds the summary and evaluates freshness-aging pressure.
  • `src/services/ai-eval-harness.ts` records freshness-aging trend fields.
  • `scripts/ai-eval-gate.ts` persists summary path and prints `citationFreshnessAging`.
  • Tests: `tests/services/ai-eval-artifacts.test.ts`.
  • Gate: direct AI eval check passed with `summary=artifacts/evals/summary.md`, `citationFreshness=1`, and `oldestAgeRatio=0`.
  • [x] RSI-8: AI eval artifact dashboard visibility (Medium-severity remediation) — closed (proven)
  • Owner: Platform + AI Reliability
  • Scope: expose latest eval summary, trend index, freshness-aging state, and warnings in a read-only app dashboard.
  • Acceptance:
  • Dashboard reads ignored `artifacts/evals/ai-eval-gate-latest.json`, `artifacts/evals/trend-index.json`, and `artifacts/evals/summary.md` when present.
  • Empty/missing local artifact state is explicit and non-failing.
  • Display model includes latest metrics, `citationFreshness`, freshness-aging warnings, run identity, and artifact timestamps.
  • Parser/render-model tests prove dashboard output matches artifact contents.
  • Closed proof (2026-06-09):
  • `src/services/ai-eval-dashboard.ts` loads latest/trend/summary artifacts with explicit `ready`, `partial`, and `empty` states.
  • `app/(workspace)/ai-evals/page.tsx` renders latest metrics, `citationFreshness`, freshness-aging state, active warnings, run identity, and artifact timestamps.
  • Navigation exposes `/ai-evals` in the workspace sidebar, Workspace menu, and footer.
  • Tests: `tests/services/ai-eval-dashboard.test.ts` and `tests/pages/ai-evals-page.test.ts`.
  • Focused gate: system Node `--test tests/services/ai-eval-dashboard.test.ts tests/pages/ai-evals-page.test.ts` passed 3 tests.
  • [x] RSI-9: latest-vs-previous AI eval artifact diff (Medium-severity remediation) — closed (proven)
  • Owner: Platform + AI Reliability
  • Scope: accelerate eval artifact review by showing latest-vs-previous deltas and “what changed” notes directly in `/ai-evals`.
  • Acceptance:
  • Dashboard model computes metric deltas for faithfulness, relevance, citation accuracy, `citationFreshness`, and pass rate.
  • Dashboard model computes freshness-aging pressure movement between retained runs when both runs have aging metadata.
  • Dashboard notes status changes, prompt-count changes, identity changes, and no-movement cases.
  • Page renders an explicit “need at least two retained eval runs” state when a diff is unavailable.
  • Closed proof (2026-06-09):
  • `src/services/ai-eval-dashboard.ts` computes `AiEvalRunDiff` with metric deltas, aging deltas, and review notes.
  • `app/(workspace)/ai-evals/page.tsx` renders the “Latest vs previous run” and “What changed notes” sections.
  • Tests: `tests/services/ai-eval-dashboard.test.ts` and `tests/pages/ai-evals-page.test.ts`.
  • Focused gate: system Node `--test tests/services/ai-eval-dashboard.test.ts tests/pages/ai-evals-page.test.ts` passed 3 tests.
  • [x] RSI-10: severity-labeled AI eval diff triage (Medium-severity remediation) — closed (proven)
  • Owner: Platform + AI Reliability
  • Scope: prioritize `/ai-evals` latest-vs-previous review with severity labels for metric drift, freshness-aging pressure, and overall diff status.
  • Acceptance:
  • Dashboard model labels each metric delta as `regression`, `watch`, `stable`, or `improved` using explicit metric thresholds.
  • Dashboard model labels freshness-aging pressure movement using explicit aging thresholds.
  • Overall diff severity escalates failed-status regressions, identity-change watch states, metric regressions, and aging regressions before stable/improved labels.
  • Page renders the overall priority label and the applied metric/aging thresholds for fast operator triage.
  • Closed proof (2026-06-09):
  • `src/services/ai-eval-dashboard.ts` defines `AI_EVAL_DIFF_SEVERITY_THRESHOLDS` and classifies `AiEvalRunDiff`, metric deltas, and aging deltas.
  • `app/(workspace)/ai-evals/page.tsx` renders priority labels and threshold pills in “Latest vs previous run.”
  • Tests: `tests/services/ai-eval-dashboard.test.ts` and `tests/pages/ai-evals-page.test.ts`.
  • Gates: focused system Node test passed 3 tests; `pnpm test` passed 787/249; `pnpm lint` passed with one existing warning; production Next build passed via system Node.
  • [x] RSI-11: policy-driven AI eval priority visibility (Medium-severity remediation) — closed (proven)
  • Owner: AI Reliability + Platform + DevEx
  • Scope: keep eval severity thresholds config-driven and surface priority in dashboard and CI summary paths.
  • Acceptance:
  • `config/ai-eval-regression-policy.json` owns `diffSeverityPolicy` thresholds for metric movement and freshness-aging pressure.
  • Dashboard model consumes policy thresholds and computes last-N severity distribution.
  • `/ai-evals` renders `regression`, `watch`, `stable`, and `improved` distribution counts even when raw artifacts are absent.
  • CI summary renders review priority and threshold context when two retained runs are available.
  • Closed proof (2026-06-09):
  • `src/services/ai-eval-severity.ts` centralizes threshold normalization, run diff severity, and distribution counting.
  • `src/services/ai-eval-dashboard.ts` and `src/services/ai-eval-artifacts.ts` consume the shared classifier.
  • `scripts/ai-eval-gate.ts` passes policy thresholds and prints `reviewPriority`.
  • Tests: `tests/services/ai-eval-dashboard.test.ts`, `tests/services/ai-eval-artifacts.test.ts`, and `tests/pages/ai-evals-page.test.ts`.
  • Gates: focused system Node test passed 8 tests; `pnpm test` passed 788/249; `pnpm lint` passed with one existing warning; AI eval gate printed `reviewPriority: stable`; production Next build passed via system Node.
  • [x] RSI-12: AI eval agent-summary reliability (Medium-severity remediation) — closed (proven)
  • Owner: AI Reliability + Platform + Agent Experience
  • Scope: make eval priority safe for agents by validating severity policy, exposing compact JSON, and showing comparison history in the dashboard.
  • Acceptance:
  • Malformed `diffSeverityPolicy` values fail schema/order validation with explicit error messages.
  • `/api/ai-evals/summary` returns agent-ready JSON with review priority, severity distribution, severity history, alerts, and artifact links.
  • `/ai-evals` renders a compact severity sparkline plus comparison history for retained runs.
  • Route supports `GET` + `OPTIONS` with baseline public JSON/CORS behavior.
  • Closed proof (2026-06-09):
  • `src/services/ai-eval-severity.ts` centralizes validation and history symbols.
  • `src/services/ai-eval-dashboard.ts` exposes `loadAiEvalAgentSummary`.
  • `app/api/ai-evals/summary/route.ts` serves the agent summary endpoint.
  • Tests: `tests/services/ai-eval-severity.test.ts`, `tests/api/ai-evals-summary.test.ts`, `tests/services/ai-eval-dashboard.test.ts`, `tests/services/ai-eval-artifacts.test.ts`, and `tests/pages/ai-evals-page.test.ts`.
  • Gates: focused system Node test passed 13 tests; `pnpm test` passed 793/251; `pnpm lint` passed with one existing warning; AI eval gate printed `reviewPriority: stable`; production Next build confirmed `/api/ai-evals/summary`.
  • [x] RSI-13: AI eval contract and CI annotation reliability (Medium-severity remediation) — closed (proven)
  • Owner: AI Reliability + DevEx + Agent Experience
  • Scope: prevent agent/API contract drift and invisible eval priority by adding schema-backed JSON/OpenAPI tests, a config-driven severity-history window, and GitHub PR annotations for `watch`/`regression`.
  • Acceptance:
  • `/api/ai-evals/summary` payload parses against the reusable Zod contract.
  • `/api/openapi` references `AiEvalSummaryResponse` for the route.
  • `severityHistoryPolicy.maxComparisons` drives dashboard and agent summary window metadata.
  • Malformed history-window policy values fail explicitly.
  • CI logs emit a GitHub warning annotation when eval priority is `watch` or `regression`.
  • Closed proof (2026-06-09):
  • `src/contracts/zod/ai-eval-summary.ts`, `src/services/openapi.ts`, `src/services/ai-eval-severity.ts`, `src/services/ai-eval-dashboard.ts`, `src/services/ai-eval-artifacts.ts`, `scripts/ai-eval-gate.ts`, `.github/workflows/ai-eval-gate.yml`, and `config/ai-eval-regression-policy.json`.
  • Tests: `tests/api/ai-evals-summary.test.ts`, `tests/api/openapi.test.ts`, `tests/services/ai-eval-severity.test.ts`, `tests/services/ai-eval-dashboard.test.ts`, `tests/services/ai-eval-artifacts.test.ts`, and `tests/pages/ai-evals-page.test.ts`.
  • Gates: focused system Node test passed 18 tests; `pnpm test` passed 796/251; `pnpm lint` passed with one existing warning; AI eval gate printed `reviewPriority: stable`; production Next build confirmed `/api/ai-evals/summary`.
  • [x] RSI-14: AI eval version and retention hygiene (Medium-severity remediation) — closed (proven)
  • Owner: AI Reliability + DevEx + Platform
  • Scope: make the agent summary contract explicitly versioned, keep CI annotation examples copy/pasteable and snapshot-proven, and prevent stale eval run artifacts from accumulating outside the retained trend window.
  • Acceptance:
  • `/api/ai-evals/summary` returns `schemaVersion: 1`.
  • Unsupported summary schema versions fail contract parsing in tests.
  • EvalOps docs publish exact GitHub annotation examples for `regression` and `watch`.
  • Annotation examples are asserted against formatter snapshots.
  • Retention pruning removes orphaned run JSON while preserving retained run JSON and non-JSON notes.
  • Closed proof (2026-06-09):
  • `src/contracts/zod/ai-eval-summary.ts`, `app/api/ai-evals/summary/route.ts`, `src/services/ai-eval-artifacts.ts`, and `docs/evals/golden-museum-questions.md`.
  • Tests: `tests/api/ai-evals-summary.test.ts`, `tests/api/openapi.test.ts`, and `tests/services/ai-eval-artifacts.test.ts`.
  • Gates: focused system Node test passed 22 tests; `pnpm test` passed 800/251; `pnpm lint` passed with one existing warning; AI eval gate printed `reviewPriority: stable`; production Next build confirmed `/api/ai-evals/summary`.
  • [x] RSI-15: AI eval migration and reporting visibility (Medium-severity remediation) — closed (proven)
  • Owner: AI Reliability + Platform + DevEx
  • Scope: make future summary migrations explicit, make pruning observable in dry-run/delete modes, and put annotation/pruning status directly into CI summary markdown.
  • Acceptance:
  • `schemaVersion: 2` migration notes exist while v2 remains unsupported.
  • Fixture compatibility tests accept `schema-v1-ready.json` and reject `schema-v2-planned.json`.
  • Retention pruning reports `delete` and `dry-run` modes with retained/orphaned/deleted/preserved counts.
  • CI summary markdown includes latest annotation status and retention pruning status.
  • Closed proof (2026-06-09):
  • `src/contracts/zod/ai-eval-summary.ts`, `src/services/ai-eval-artifacts.ts`, `docs/evals/golden-museum-questions.md`, and `tests/fixtures/ai-eval-summary/*`.
  • Tests: `tests/api/ai-evals-summary.test.ts` and `tests/services/ai-eval-artifacts.test.ts`.
  • Gates: focused system Node test passed 24 tests; `pnpm test` passed 802/251; `pnpm lint` passed with one existing warning; AI eval gate printed `reviewPriority: stable`; production Next build confirmed `/api/ai-evals/summary`.
  • [x] RSI-16: AI eval summary artifact/schema visibility (Medium-severity remediation) — closed (proven)
  • Owner: DevEx + Platform + AI Reliability
  • Scope: make CI summary markdown snapshot-verifiable, make latest pruning status agent-readable, and make schema migration compatibility visible through OpenAPI.
  • Acceptance:
  • Generated `artifacts/evals/summary.md` matches `tests/fixtures/ai-eval-summary/summary-snapshot.md`.
  • `/api/ai-evals/summary` includes latest `summary.retentionPruneReport`.
  • `/api/openapi` includes the AI eval summary schema migration compatibility table.
  • Closed proof (2026-06-09):
  • `src/services/ai-eval-artifacts.ts`, `src/services/ai-eval-dashboard.ts`, `src/contracts/zod/ai-eval-summary.ts`, `src/services/openapi.ts`, `docs/evals/golden-museum-questions.md`, and `tests/fixtures/ai-eval-summary/summary-snapshot.md`.
  • Tests: `tests/api/ai-evals-summary.test.ts`, `tests/api/openapi.test.ts`, `tests/services/ai-eval-artifacts.test.ts`, and `tests/services/ai-eval-dashboard.test.ts`.
  • Gates: focused system Node test passed 25 tests; `pnpm test` passed 803/251; `pnpm lint` passed with one existing warning; AI eval gate printed `reviewPriority: stable`; production Next build confirmed `/api/ai-evals/summary`.
  • [x] RSI-17: Visual ETL Mapper AI-assist safety (Medium-severity remediation) — closed (proven)
  • Owner: AI Reliability + Curator Experience + Platform
  • Scope: complete C4 LLM-assisted mapping without allowing automated ingestion activation from unreviewed suggestions.
  • Acceptance:
  • `/api/ai/mapping-assist` returns a contract-valid draft `MappingTemplate`.
  • Unknown columns are emitted as diagnostics, not hallucinated mappings.
  • `/etl/mapper` exposes the assist action as curator-review-only.
  • Closed proof (2026-06-09):
  • `src/services/mapping-assist.ts`, `src/contracts/zod/mapping-assist.ts`, `app/api/ai/mapping-assist/route.ts`, and `src/components/etl-mapper-workbench.tsx`.
  • Tests: `tests/services/mapping-assist.test.ts`, `tests/api/ai-mapping-assist.test.ts`, `tests/components/etl-mapper-config.test.ts`, and `tests/api/openapi.test.ts`.
  • Gates: focused system Node test passed 7 mapper-assist tests plus 2 OpenAPI tests; `pnpm test` passed 809/253; `pnpm lint` passed with one existing warning; AI eval gate printed `reviewPriority: stable`; production Next build confirmed `/api/ai/mapping-assist`.
  • [x] RSI-18: mapper-assist fixture/schema/importability hardening (Medium-severity remediation) — closed (proven)
  • Owner: AI Reliability + Curator UX + Platform
  • Scope: harden mapper assist after launch with golden fixtures, importable draft rendering, and OpenAPI contract visibility.
  • Acceptance:
  • Tricky rights/credit/sensitive columns stay unmapped and cannot create hallucinated target paths.
  • Returned suggestions can be converted into ReactFlow draft nodes/edges for curator review.
  • `/api/openapi` exposes and references `MappingAssistResponse`.
  • Closed proof (2026-06-09):
  • `tests/fixtures/mapping-assist/tricky-columns.json`, `src/services/mapping-assist.ts`, `src/utils/etl-mapper-assist.ts`, `src/components/etl-mapper-workbench.tsx`, `src/contracts/zod/mapping-assist.ts`, and `src/services/openapi.ts`.
  • Tests: `tests/services/mapping-assist.test.ts`, `tests/utils/etl-mapper-assist.test.ts`, `tests/components/etl-mapper-config.test.ts`, `tests/api/openapi.test.ts`, and `tests/api/ai-mapping-assist.test.ts`.
  • Gates: focused system Node test passed 11 tests; `pnpm test` passed 811/254; `pnpm lint` passed with one existing warning; AI eval gate printed `reviewPriority: stable`; production Next build confirmed `/api/ai/mapping-assist`.
  • [x] RSI-19: provider-family/browser/request-schema mapper hardening (Medium-severity remediation) — closed (proven)
  • Owner: AI Reliability + Curator UX + Platform
  • Scope: make mapper assist safer across provider families, browser interaction, and request contract discovery.
  • Acceptance:
  • Met/Getty/Rijks provider-family fixtures map only expected columns to approved Linked Art target paths.
  • `/etl/mapper` has a browser-level assist/import smoke command discoverable as `pnpm smoke:etl:mapper-assist`.
  • `/api/openapi` exposes `MappingAssistRequest` and references it as the mapping-assist POST request body.
  • Closed proof (2026-06-09):
  • `tests/fixtures/mapping-assist/provider-families.json`, `scripts/smoke-etl-mapper-assist.ts`, `package.json`, `src/contracts/zod/mapping-assist.ts`, and `src/services/openapi.ts`.
  • Tests: `tests/services/mapping-assist.test.ts`, `tests/scripts/etl-mapper-smoke-script.test.ts`, and `tests/api/openapi.test.ts`.
  • Gates: focused system Node test passed 7 tests; `pnpm smoke:etl:mapper-assist` passed against `http://localhost:3001/en`; `pnpm test` passed 813/255; `pnpm lint` passed with one existing warning; AI eval gate printed `reviewPriority: stable`; production Next build confirmed `/api/ai/mapping-assist`.
  • [x] RSI-20: negative mapper fixtures, visual screenshot, and API-doc examples (Medium-severity remediation) — closed (proven)
  • Owner: AI Reliability + Curator UX + Platform
  • Scope: harden near-miss column refusal, imported-draft visual regression evidence, and human/agent examples in `/api/docs`.
  • Acceptance:
  • Almost-mappable provider columns containing rights, credit, restriction, sensitivity, donor, or flag wording stay unmapped.
  • Browser smoke imports the assist draft and writes a stable screenshot artifact at `artifacts/smoke/etl-mapper-assist-imported.png`.
  • `/api/docs` includes concrete request and response examples for `/api/ai/mapping-assist`.
  • Closed proof (2026-06-09):
  • `tests/fixtures/mapping-assist/negative-provider-families.json`, `src/services/mapping-assist.ts`, `scripts/smoke-etl-mapper-assist.ts`, and `app/api/docs/route.ts`.
  • Tests: `tests/services/mapping-assist.test.ts`, `tests/scripts/etl-mapper-smoke-script.test.ts`, and `tests/api/docs.test.ts`.
  • Gates: focused system Node test passed 11 tests; `pnpm smoke:etl:mapper-assist` passed and wrote `artifacts/smoke/etl-mapper-assist-imported.png`; `pnpm test` passed 814/255; `pnpm lint` passed with one existing warning; AI eval gate printed `reviewPriority: stable`; production Next build passed.
  • [x] RSI-21: mapper layout, OpenAPI-sourced docs examples, and confidence policy (Medium-severity remediation) — closed (proven)
  • Owner: Curator UX + Platform + AI Reliability
  • Scope: improve curator trust by fixing mapper action layout, making docs examples single-source, and gating low-confidence suggestions.
  • Acceptance:
  • Refreshed browser screenshot shows mapper actions in a non-overlapping wrapped row.
  • `/api/docs` request/response examples are generated from `/api/openapi` examples.
  • Low-confidence accession/place/description patterns stay diagnostics-only.
  • Closed proof (2026-06-09):
  • `app/globals.css`, `app/api/docs/route.ts`, `src/contracts/zod/mapping-assist.ts`, `src/services/openapi.ts`, and `src/services/mapping-assist.ts`.
  • Tests: `tests/components/etl-mapper-config.test.ts`, `tests/api/docs.test.ts`, and `tests/services/mapping-assist.test.ts`.
  • Gates: focused system Node test passed 15 tests; `pnpm smoke:etl:mapper-assist` refreshed `artifacts/smoke/etl-mapper-assist-imported.png`; `pnpm test` passed 817/255; `pnpm lint` passed with one existing warning; AI eval gate printed `reviewPriority: stable`; production Next build passed.
  • [x] RSI-22: public source narrative and trust uplift (Medium-severity remediation) — closed (proven)
  • Owner: Platform + Curator UX + Trust
  • Scope: import only useful sibling-repo source-network ideas into current production surfaces: registry-backed `/datasets`, reviewed local metadata image, asset provenance tracking, footer trust links, and rewritten Contact/Privacy/Terms pages.
  • Acceptance:
  • `/datasets` and homepage stats derive provider counts from `getProviderCapabilities`, not copied constants.
  • OpenGraph/Twitter metadata points at a reviewed local image.
  • Footer exposes Contact, Privacy, Terms, and asset provenance links.
  • Imported sibling assets have SHA-256 provenance and placeholder thumbnails/legal copy are excluded.
  • Closed proof (2026-06-09):
  • `src/services/public-source-narrative.ts`, `src/services/site-metadata.ts`, `app/datasets/page.tsx`, `app/contact/page.tsx`, `app/privacy/page.tsx`, `app/terms/page.tsx`, `docs/asset-provenance.md`, and `public/images/meta-museum-art/*`.
  • Tests: `tests/services/public-source-narrative.test.ts` and `tests/pages/public-source-pages.test.ts`.
  • Gates: focused RSI-22 tests passed 6/2; `pnpm test` passed 823/257; `pnpm lint` passed with one existing warning; `pnpm build` passed; screenshot captured at `artifacts/smoke/datasets-page.png`.
  • [x] RSI-23: public-source agent API and trust smoke hardening (Medium-severity remediation) — closed (proven)
  • Owner: Platform + Trust + Curator UX
  • Scope: make the public source/trust uplift machine-readable and visually checked by adding a summary API, license-review status enum, and browser smoke screenshots for all public trust pages.
  • Acceptance:
  • `/api/public-sources/summary` returns schema-versioned JSON that parses against `publicSourcesSummaryResponseSchema`.
  • Asset provenance rows use the explicit `reviewed-local-use` / `needs-license-review` / `rejected-placeholder` enum, and unknown values fail tests.
  • `pnpm smoke:public-trust` captures `/datasets`, `/contact`, `/privacy`, and `/terms` screenshots.
  • The nested `C:\Projects\metamuseum\meta-museum-art` prototype copy is removed after all images are preserved under `public/images/meta-museum-art/` and inventoried in `docs/asset-provenance.md`.
  • Closed proof (2026-06-09):
  • `app/api/public-sources/summary/route.ts`, `src/contracts/zod/public-sources-summary.ts`, `src/services/public-source-narrative.ts`, `scripts/smoke-public-trust-pages.ts`, `docs/asset-provenance.md`, and `package.json`.
  • Tests: `tests/api/public-sources-summary.test.ts`, `tests/services/public-source-narrative.test.ts`, and `tests/scripts/public-trust-smoke-script.test.ts`.
  • Gates: focused RSI-23 tests passed 6/3; `pnpm smoke:public-trust` passed against `http://localhost:3001`; `pnpm test` passed 827/259; `pnpm lint` passed with one existing warning; `pnpm build` passed.
  • [x] RSI-24: OpenAPI, checksum drift, and screenshot retention hardening (Medium-severity remediation) — closed (proven)
  • Owner: Platform + Trust + Curator UX
  • Scope: make the public source summary discoverable through OpenAPI, prove imported visual assets have not drifted from provenance checksums, and convert public trust screenshots into retained latest/previous review artifacts.
  • Acceptance:
  • `/api/openapi` exposes `PublicSourcesSummaryResponse` and references it from `/api/public-sources/summary`.
  • `tests/services/public-source-narrative.test.ts` fails when a copied asset hash differs from `docs/asset-provenance.md`.
  • `pnpm smoke:public-trust` writes timestamped screenshot runs plus `latest` copies and prunes old runs beyond `PUBLIC_TRUST_SCREENSHOT_RETENTION_MAX_RUNS`.
  • Closed proof (2026-06-09):
  • `src/services/openapi.ts`, `src/contracts/zod/public-sources-summary.ts`, `src/services/public-trust-smoke-artifacts.ts`, `scripts/smoke-public-trust-pages.ts`, and `docs/asset-provenance.md`.
  • Tests: `tests/api/openapi.test.ts`, `tests/services/public-source-narrative.test.ts`, `tests/services/public-trust-smoke-artifacts.test.ts`, and `tests/scripts/public-trust-smoke-script.test.ts`.
  • Gates: focused RSI-24 tests passed 6/5; `pnpm smoke:public-trust` passed against `http://localhost:3001`; `pnpm test` passed 828/260; `pnpm lint` passed with one existing warning; `pnpm build` passed.
  • [x] RSI-25: Public trust docs, diff metadata, and CI artifact visibility (Medium-severity remediation) — closed (proven)
  • Owner: Platform + Trust + DevEx
  • Scope: keep public-source examples and public trust smoke artifacts review-visible by adding OpenAPI-sourced `/api/docs` payloads, latest-vs-previous screenshot diff metadata, and CI summary/upload links.
  • Acceptance:
  • `/api/docs` renders `/api/public-sources/summary` example payloads from `/api/openapi`.
  • Public trust smoke `summary.json` includes per-screenshot latest-vs-previous checksum/byte diff metadata.
  • CI appends public trust artifact links to `$GITHUB_STEP_SUMMARY` and uploads `public-trust-smoke-artifacts`.
  • Closed proof (2026-06-09):
  • `app/api/docs/route.ts`, `src/services/openapi.ts`, `src/contracts/zod/public-sources-summary.ts`, `src/services/public-trust-smoke-artifacts.ts`, `src/services/public-trust-smoke-ci-summary.ts`, `scripts/public-trust-smoke-ci-summary.ts`, and `.github/workflows/public-trust-smoke.yml`.
  • Tests: `tests/api/docs.test.ts`, `tests/api/openapi.test.ts`, `tests/services/public-trust-smoke-artifacts.test.ts`, `tests/services/public-trust-smoke-ci-summary.test.ts`, `tests/scripts/public-trust-smoke-script.test.ts`, and `tests/scripts/public-trust-ci-workflow.test.ts`.
  • Gates: focused RSI-25 tests passed 12/10; `pnpm smoke:public-trust` passed against temporary `http://localhost:3001`; `pnpm test` passed 830/262; `pnpm lint` passed with one existing warning; `pnpm build` passed.
  • [x] RSI-26: Pixel-diff thresholds, public trust summary API, and main CI artifact links (Medium-severity remediation) — closed (proven)
  • Owner: Trust + DevEx + Platform
  • Scope: harden public trust screenshot review by failing meaningful pixel drift, exposing the latest smoke report to agents, and publishing artifact links from main CI.
  • Acceptance:
  • `pnpm smoke:public-trust` fails when screenshot changed-pixel ratio exceeds `PUBLIC_TRUST_SCREENSHOT_PIXEL_DIFF_THRESHOLD`.
  • `/api/public-trust/summary` returns schema-versioned latest public trust artifact JSON.
  • Main `.github/workflows/ci.yml` appends public trust summary links and uploads `public-trust-smoke-artifacts`.
  • Closed proof (2026-06-09):
  • `src/services/public-trust-smoke-artifacts.ts`, `scripts/smoke-public-trust-pages.ts`, `src/contracts/zod/public-trust-summary.ts`, `src/services/public-trust-summary.ts`, `app/api/public-trust/summary/route.ts`, `src/services/openapi.ts`, and `.github/workflows/ci.yml`.
  • Tests: `tests/services/public-trust-smoke-artifacts.test.ts`, `tests/api/public-trust-summary.test.ts`, `tests/api/openapi.test.ts`, `tests/scripts/public-trust-smoke-script.test.ts`, and `tests/scripts/public-trust-ci-workflow.test.ts`.
  • Gates: focused RSI-26 tests passed 10/7; `pnpm smoke:public-trust` passed with `unchanged=4` and `pixel failures=0`; `pnpm test` passed 834/263; `pnpm lint` passed with one existing warning; `pnpm build` passed.
  • [x] RSI-27: Public trust per-page thresholds, OpenAPI docs example, and retention badge (Medium-severity remediation) — closed (proven)
  • Owner: Trust + Platform + DevEx
  • Scope: make visual drift review more precise with per-page thresholds, make `/api/public-trust/summary` discoverable in `/api/docs`, and snapshot-lock the public trust CI retention badge.
  • Acceptance:
  • `/datasets` uses stricter pixel threshold (`0.005`) than Contact/Privacy/Terms (`0.02`).
  • `/api/docs` renders `/api/public-trust/summary` examples from `/api/openapi`.
  • `renderPublicTrustSmokeCiSummary` output matches `tests/fixtures/public-trust-summary/summary-snapshot.md`.
  • Closed proof (2026-06-09):
  • `src/services/public-trust-smoke-artifacts.ts`, `scripts/smoke-public-trust-pages.ts`, `src/contracts/zod/public-trust-summary.ts`, `src/services/openapi.ts`, `app/api/docs/route.ts`, and `src/services/public-trust-smoke-ci-summary.ts`.
  • Tests: `tests/services/public-trust-smoke-artifacts.test.ts`, `tests/scripts/public-trust-smoke-script.test.ts`, `tests/services/public-trust-smoke-ci-summary.test.ts`, `tests/fixtures/public-trust-summary/summary-snapshot.md`, `tests/api/docs.test.ts`, and `tests/api/openapi.test.ts`.
  • Gates: focused RSI-27 tests passed 13/9; `pnpm smoke:public-trust` passed with per-page thresholds, `unchanged=4`, and `pixel failures=0`; `pnpm test` passed 835/263; `pnpm lint` passed with one existing warning; `pnpm build` passed.
  • [x] RSI-28: JSON public trust threshold policy, agent-visible policy, and CI drift annotations (Medium-severity remediation) — closed (proven)
  • Owner: Trust + Platform + DevEx
  • Scope: move per-page public trust thresholds into JSON policy, expose the applied policy through the agent summary API, and warn reviewers when screenshots change under threshold.
  • Acceptance:
  • `scripts/smoke-public-trust-pages.ts` consumes `config/public-trust-smoke-policy.json` through `loadPublicTrustSmokePolicy`.
  • `/api/public-trust/summary` includes `thresholdPolicy` with default and per-page thresholds.
  • CI summary tooling emits GitHub warning annotations for changed screenshots whose pixel ratio remains under the configured threshold.
  • Closed proof (2026-06-09):
  • `config/public-trust-smoke-policy.json`, `src/services/public-trust-smoke-policy.ts`, `scripts/smoke-public-trust-pages.ts`, `src/services/public-trust-summary.ts`, `src/contracts/zod/public-trust-summary.ts`, and `src/services/public-trust-smoke-ci-summary.ts`.
  • Tests: `tests/services/public-trust-smoke-policy.test.ts`, `tests/scripts/public-trust-smoke-script.test.ts`, `tests/api/public-trust-summary.test.ts`, `tests/services/public-trust-smoke-ci-summary.test.ts`, and `tests/scripts/public-trust-ci-workflow.test.ts`.
  • Gates: focused RSI-28 tests passed 20/12; `pnpm smoke:public-trust` passed with JSON policy thresholds, `unchanged=4`, and `pixel failures=0`; `pnpm test` passed 838/264; `pnpm lint` passed with one existing warning; `pnpm build` passed.
  • [x] RSI-29: Public trust reviewer rationale, schema rejection, and severity-grouped PR summaries (Medium-severity remediation) — closed (proven)
  • Owner: Trust + Platform + DevEx
  • Scope: make public trust review rationale explicit in policy/API payloads, reject unsupported future policy schema versions, and group under-threshold CI drift by route severity.
  • Acceptance:
  • Policy pages include `reviewSeverity`, `reviewerNote`, and `reasonCodes`.
  • `/api/public-trust/summary` exposes reviewer rationale metadata in `thresholdPolicy.pages`.
  • `loadPublicTrustSmokePolicy` rejects unsupported `schemaVersion: 2` fixtures.
  • CI PR summary and warnings group changed-under-threshold screenshots by severity.
  • Closed proof (2026-06-09):
  • `config/public-trust-smoke-policy.json`, `src/services/public-trust-smoke-policy.ts`, `src/services/public-trust-summary.ts`, `src/contracts/zod/public-trust-summary.ts`, `scripts/smoke-public-trust-pages.ts`, and `src/services/public-trust-smoke-ci-summary.ts`.
  • Tests: `tests/services/public-trust-smoke-policy.test.ts`, `tests/api/public-trust-summary.test.ts`, `tests/services/public-trust-smoke-ci-summary.test.ts`, and `tests/fixtures/public-trust-summary/summary-snapshot.md`.
  • Gates: focused RSI-29 tests passed 18/11 plus CI-summary focused retest 2/1; `pnpm exec start-server-and-test "pnpm dev" http://localhost:3000 "pnpm smoke:public-trust"` passed with JSON policy metadata and `pixel failures=0`; `pnpm test` passed 839/264; `pnpm lint` passed with one existing warning; `pnpm build` passed.
  • [x] RSI-30: Owner/reviewer initials, schema v2 fixture, and grouped annotation snapshot (Medium-severity remediation) — closed (proven)
  • Owner: Trust + Platform + DevEx
  • Scope: make review ownership explicit per policy row, preserve a future v2 migration fixture, and snapshot exact grouped warning annotations.
  • Acceptance:
  • Policy pages include `ownerInitials` and `reviewerInitials`.
  • `/api/public-trust/summary` exposes owner/reviewer initials in `thresholdPolicy.pages`.
  • CI summary rows and grouped warning annotations include owner/reviewer initials.
  • `tests/fixtures/public-trust-policy/schema-v2-planned.json` remains rejected until migration support lands.
  • Closed proof (2026-06-09):
  • `config/public-trust-smoke-policy.json`, `src/services/public-trust-smoke-policy.ts`, `src/services/public-trust-summary.ts`, `src/contracts/zod/public-trust-summary.ts`, `src/services/public-trust-smoke-ci-summary.ts`, and `scripts/smoke-public-trust-pages.ts`.
  • Tests: `tests/fixtures/public-trust-policy/schema-v2-planned.json`, `tests/fixtures/public-trust-summary/grouped-annotations-snapshot.txt`, `tests/services/public-trust-smoke-policy.test.ts`, `tests/api/public-trust-summary.test.ts`, and `tests/services/public-trust-smoke-ci-summary.test.ts`.
  • Gates: focused RSI-30 tests passed 15/10; `pnpm exec start-server-and-test "pnpm dev" http://localhost:3000 "pnpm smoke:public-trust"` passed with `unchanged=4` and `pixel failures=0`; first full `pnpm test` exposed a transient `tests/api/artworks/by-id.test.ts` failure that passed isolated retest; final `pnpm test` passed 839/264; `pnpm lint` passed with one existing warning; `pnpm build` passed.

Clio evidence-hook risk update (2026-08-25): the new trial-only capability

removes deterministic safety editing from the proposed model benefit and binds

each generated sentence to explicit evidence IDs plus a factual-token verifier.

Aggregate packet inspection invalidated v3 because all five adversarial model

hooks omitted required negated boundary evidence. Contract v2 now types fact and

boundary evidence, requires every boundary ID, and preserves explicit negation.

V4 confirmed the repair; v5 exposed replayable boundary counts but failed shared

packet-schema interoperation before assignment. V6 emits exact four-field rows,

retains 10/10/10 required-citation, observed-citation, and preserved-negation

counts, and produces ten non-identical pairs at $0.00581 total cost and

1.198-second worst latency. This does not

prove quality: sentence-level confinement can still permit misleading word order,

and no independent review exists. Runtime/deployment/publication authority

remains false; promotion requires two blinded reviewers, protected-dimension

non-regression, at least 10% lift, and no reviewer-effort increase. Strict receipt

replay prevents filename-only provider evidence.

Materiality update: the exact packet/key-bound analyzer reports 10 normalized

non-cosmetic pairs and stable per-case output, but 0.942 token-set overlap and a

5.1% model length increase. Treat these as review-routing diagnostics only;

neither structural difference nor stability establishes user benefit.

Reviewer-handoff update: two v6 forms and their public packet/form-digest receipt

exist and semantically replay; they prove interface readiness only. Genuine

distribution, independence, completed responses, and value remain external.

Current inventory clarification (2026-08-25): the evidence-hook registration

raises the governed inventory to 14 surfaces and six model-backed workflows;

zero AI advantages are proven. Earlier 13-surface/five-model wording records the

pre-registration audit state only.

Controls and mitigation plan

| Risk | Control | Owner | Evidence / check |

|---|---|---|---|

| Adapter and pipeline boundary drift | Static boundary contract test ensures adapter modules do not import each other (except allowed shared interfaces/utils) | Platform | `tests/contracts/provider-boundary-contracts.test.ts` |

| Adapter and pipeline boundary drift | AI-RSI evidence loop captures boundary checks each cycle | Engineering + AI-assisted operator | `docs/rsi-wiki.md`, `README.md`, `docs/roadmap.md`, `CLAUDE.md` |

| Adapter and pipeline boundary drift | Risk review in each close-out row for new provider/pipeline touchpoints | Platform owner | `CLAUDE.md` close-out log + section updates |

| UI journey automation | Role/provider matrix smoke plus entity-role route assertions for imported records | QA + Product | `scripts/smoke-explore-import-matrix.ts`, `package.json` (`smoke:explore:matrix`) |

| Low-severity complexity | Maintainability decomposition plan before future scale drift | Product + Platform | Planned owner list and decomposition evidence in this register + RSI close-out rows |

| AI eval freshness drift | Regression baseline includes `citationFreshness` and fail-fast freshness-drop policy | Platform + AI Reliability | `config/ai-eval-regression-policy.json`, `scripts/ai-eval-gate.ts`, `tests/services/ai-eval-harness.test.ts` |

| AI eval artifact visibility | CI summary badges, artifact upload, and freshness-aging pressure alerts | Platform + AI Reliability | `.github/workflows/ai-eval-gate.yml`, `src/services/ai-eval-artifacts.ts`, `tests/services/ai-eval-artifacts.test.ts` |

| AI eval dashboard visibility | Read-only dashboard for latest/trend eval artifacts and warnings | Platform + AI Reliability | `src/services/ai-eval-dashboard.ts`, `app/(workspace)/ai-evals/page.tsx`, `tests/services/ai-eval-dashboard.test.ts`, `tests/pages/ai-evals-page.test.ts` |

| AI eval artifact diff review | Latest-vs-previous dashboard diff with metric/aging deltas and review notes | Platform + AI Reliability | `src/services/ai-eval-dashboard.ts`, `app/(workspace)/ai-evals/page.tsx`, `tests/services/ai-eval-dashboard.test.ts`, `tests/pages/ai-evals-page.test.ts` |

| AI eval diff prioritization | Severity-labeled diff thresholds for `regression`, `watch`, `stable`, and `improved` triage | Platform + AI Reliability | `src/services/ai-eval-dashboard.ts`, `app/(workspace)/ai-evals/page.tsx`, `tests/services/ai-eval-dashboard.test.ts`, `tests/pages/ai-evals-page.test.ts` |

| AI eval priority discoverability | Config-driven severity policy, dashboard distribution, and CI review-priority summary | AI Reliability + Platform + DevEx | `config/ai-eval-regression-policy.json`, `src/services/ai-eval-severity.ts`, `src/services/ai-eval-artifacts.ts`, `scripts/ai-eval-gate.ts` |

| AI eval agent-summary reliability | Policy validation, agent JSON endpoint, and severity sparkline/history | AI Reliability + Platform + Agent Experience | `src/services/ai-eval-severity.ts`, `app/api/ai-evals/summary/route.ts`, `tests/api/ai-evals-summary.test.ts` |

| AI eval contract and CI annotation reliability | Schema-backed summary payload, configurable history window, and PR-visible priority annotations | AI Reliability + DevEx + Agent Experience | `src/contracts/zod/ai-eval-summary.ts`, `src/services/openapi.ts`, `scripts/ai-eval-gate.ts`, `tests/api/openapi.test.ts` |

| AI eval version and retention hygiene | Versioned summary contract, snapshot-documented annotations, and orphaned run pruning | AI Reliability + DevEx + Platform | `src/contracts/zod/ai-eval-summary.ts`, `src/services/ai-eval-artifacts.ts`, `tests/services/ai-eval-artifacts.test.ts` |

| AI eval migration and reporting visibility | Future schema migration notes, fixture compatibility, dry-run pruning reports, and CI summary status | AI Reliability + Platform + DevEx | `src/contracts/zod/ai-eval-summary.ts`, `src/services/ai-eval-artifacts.ts`, `tests/fixtures/ai-eval-summary/schema-v1-ready.json` |

| AI eval summary artifact/schema visibility | Snapshot-locked CI summary markdown, agent-visible pruning report, and OpenAPI migration compatibility table | DevEx + Platform + AI Reliability | `tests/fixtures/ai-eval-summary/summary-snapshot.md`, `src/services/ai-eval-dashboard.ts`, `src/services/openapi.ts`, `tests/api/openapi.test.ts` |

| Visual ETL Mapper AI-assist safety | Review-only mapping suggestions with contract validation, rationale, standards anchors, and unmapped-column diagnostics | AI Reliability + Curator Experience + Platform | `src/services/mapping-assist.ts`, `app/api/ai/mapping-assist/route.ts`, `tests/services/mapping-assist.test.ts`, `tests/api/ai-mapping-assist.test.ts` |

| Mapper-assist fixture/schema/importability drift | Golden tricky-column fixture, importable ReactFlow draft helper, and OpenAPI response component | AI Reliability + Curator UX + Platform | `tests/fixtures/mapping-assist/tricky-columns.json`, `src/utils/etl-mapper-assist.ts`, `src/contracts/zod/mapping-assist.ts`, `tests/api/openapi.test.ts` |

| Mapper-assist provider/browser/request-schema drift | Provider-family fixtures, browser smoke command, and OpenAPI request-body component | AI Reliability + Curator UX + Platform | `tests/fixtures/mapping-assist/provider-families.json`, `scripts/smoke-etl-mapper-assist.ts`, `src/contracts/zod/mapping-assist.ts`, `tests/api/openapi.test.ts` |

| Mapper-assist near-miss/docs/visual drift | Negative provider fixtures, imported-draft screenshot artifact, and `/api/docs` examples | AI Reliability + Curator UX + Platform | `tests/fixtures/mapping-assist/negative-provider-families.json`, `scripts/smoke-etl-mapper-assist.ts`, `app/api/docs/route.ts`, `tests/api/docs.test.ts` |

| Mapper-assist layout/example/confidence drift | Wrapped mapper actions, OpenAPI-sourced docs examples, and low-confidence diagnostics-only policy | Curator UX + Platform + AI Reliability | `app/globals.css`, `app/api/docs/route.ts`, `src/services/mapping-assist.ts`, `tests/components/etl-mapper-config.test.ts` |

| Public source/trust drift | Registry-backed public narrative, provenance-tracked imported assets, and rewritten trust pages | Platform + Curator UX + Trust | `src/services/public-source-narrative.ts`, `docs/asset-provenance.md`, `tests/pages/public-source-pages.test.ts` |

| Public source/trust drift | Agent summary API, asset license-review enum, and four-page screenshot smoke | Platform + Trust + Curator UX | `app/api/public-sources/summary/route.ts`, `src/contracts/zod/public-sources-summary.ts`, `scripts/smoke-public-trust-pages.ts` |

| Public source/trust drift | OpenAPI schema reference, checksum drift test, and screenshot latest/previous retention pruning | Platform + Trust + Curator UX | `src/services/openapi.ts`, `tests/api/openapi.test.ts`, `src/services/public-trust-smoke-artifacts.ts` |

| Public source/trust drift | OpenAPI-sourced docs examples, screenshot diff metadata, and CI summary/upload links | Platform + Trust + DevEx | `app/api/docs/route.ts`, `src/services/public-trust-smoke-ci-summary.ts`, `.github/workflows/public-trust-smoke.yml` |

| Public source/trust drift | Pixel-diff threshold failure, public trust summary API, and main CI artifact links | Trust + Platform + DevEx | `src/services/public-trust-smoke-artifacts.ts`, `app/api/public-trust/summary/route.ts`, `.github/workflows/ci.yml` |

| Public source/trust drift | Per-page pixel thresholds, OpenAPI-sourced trust summary example, and retention badge snapshot | Trust + Platform + DevEx | `scripts/smoke-public-trust-pages.ts`, `app/api/docs/route.ts`, `tests/fixtures/public-trust-summary/summary-snapshot.md` |

| Public source/trust drift | JSON threshold policy, agent-visible applied thresholds, and CI under-threshold drift warnings | Trust + Platform + DevEx | `config/public-trust-smoke-policy.json`, `app/api/public-trust/summary/route.ts`, `src/services/public-trust-smoke-ci-summary.ts` |

| Public source/trust drift | Reviewer rationale metadata, schema-version rejection, and severity-grouped PR summaries | Trust + Platform + DevEx | `config/public-trust-smoke-policy.json`, `src/services/public-trust-smoke-policy.ts`, `src/services/public-trust-smoke-ci-summary.ts` |

| Public source/trust drift | Owner/reviewer initials, future v2 migration fixture, and snapshot-locked warning annotations | Trust + Platform + DevEx | `config/public-trust-smoke-policy.json`, `tests/fixtures/public-trust-policy/schema-v2-planned.json`, `tests/fixtures/public-trust-summary/grouped-annotations-snapshot.txt` |

| Pre-revenue SaaS path ambiguity | Public commercial-readiness ledger and manual invoice gate prevent revenue/billing overclaims | Product + Platform | `src/services/pilot-offer.ts`, `app/pilot/page.tsx`, `tests/services/pilot-offer.test.ts`, `tests/pages/public-source-pages.test.ts`, `docs/ops/managed-linked-art-pilot-runbook.md` |

| Large API surface drift | Operations risk report measures route files, route families, exported methods, scoped-storage routes, and mutating routes; every first-level API route family must stay classified with owner and review lanes | Platform + DevEx | `src/services/operations-risk-controls.ts`, `tests/services/operations-risk-controls.test.ts`, `tests/quality/org-scope-route-matrix.test.ts`, `tests/quality/route-storage-write-guard.test.ts` |

| Script proliferation ownership drift | Operations risk report joins script/package counts to the evidence ownership registry, fails on unclassified package-script namespaces, and fails on unowned high-churn aliases | Platform + Operators | `src/services/operations-risk-controls.ts`, `src/services/evidence-script-ownership.ts`, `tests/services/evidence-script-ownership.test.ts`, `docs/ops/evidence-script-ownership.md` |

| Mixed runtime topology parity drift | Operations risk report requires deployment docs, deployment preflight, projection readiness, and disabled-system drill coverage for Vercel/Neon/Render/Redis/Cron/Solr/GraphDB | Platform + Operations | `src/services/operations-risk-controls.ts`, `docs/deployment.md`, `docs/ops/deployment-preflight.md`, `src/services/deployment-preflight.ts`, `src/services/staging-disabled-systems-drill.ts` |

| Evidence artifacts as product state | Operations risk report requires review-goals schema validation, readiness source age display, public trust retention, and migration fixtures | Platform + Trust + Operators | `src/services/operations-risk-controls.ts`, `docs/schemas/review-goals.schema.json`, `src/services/operator-readiness-dashboard.ts`, `src/services/public-trust-smoke-artifacts.ts` |

| Cold record reads and SLO history depth | Performance scale report synthetically proves `apiColdRecord` p95 misses fail and clustered samples do not satisfy 30-day SLO depth | Platform + Reliability | `src/services/performance-scalability-controls.ts`, `src/services/long-term-evidence-runway.ts`, `scripts/k6-slo.js`, `tests/services/performance-scalability-controls.test.ts` |

| Projection scaling enablement drift | Performance scale report exercises Solr/GraphDB enable thresholds and scheduled-drain blockers before target flags are allowed | Platform + Search/Graph | `src/services/performance-scalability-controls.ts`, `src/services/projection-scale-readiness.ts`, `scripts/projection-readiness.ts`, `tests/services/projection-scale-readiness.test.ts` |

| Cron drain backlog growth | Performance scale report exercises `/workers` lag states and keeps sustained backlog escalation tied to Render background workers | Platform + Operations | `src/services/performance-scalability-controls.ts`, `src/services/worker-status.ts`, `scripts/outbox-alert-check.ts`, `docs/deployment.md` |

| External provider/API dependence | Performance scale report requires adapter cache/rate-limit policy, provider request limits, authority refresh, and validation-drift checks | Platform + Provider Reliability | `src/services/performance-scalability-controls.ts`, `src/adapters/provider-interface.ts`, `src/services/validation-drift.ts`, `scripts/authority-cache-refresh.ts` |

| Generative prose with decorative or mismatched citations | Calliope requires strict claim-to-source JSON and validates URLs, years, and authority language against each claim's cited evidence; paid postflight failures retain spend telemetry and fall back locally | AI Reliability + Editorial Research | `src/services/grounded-draft-contract.ts`, `tests/services/grounded-draft-contract.test.ts`, `artifacts/ai-agent-value/trials/calliope-curatorial-drafting-2026-08-24-v4.json` |

| Static agent tool lists mistaken for observed execution | Clio, Mercator, Janus, Themis, and Calliope emit execution-time tool status, bounded metrics, and evidence hashes; blocked precondition paths claim no execution | AI Reliability + Agent Experience | `src/services/agents.ts`, `tests/services/agent-run-telemetry.test.ts`, `src/components/agents-workbench.tsx` |

| Reconciliation model compared against a weak baseline or charged for evidence-equivalent ties | Retired v2 evidence is historical only; new spend requires a registered/adjudicated v3 representative that the identifier-aware baseline cannot resolve, while evidence-equivalent adversarial rows abstain before provider access | AI Reliability + Data Curation | `src/services/reconciliation-tiebreaker-trial.ts`, `scripts/reconciliation-tiebreaker-trial.ts`, `tests/services/reconciliation-tiebreaker-trial.test.ts` |

| Secrets hygiene recurrence | Security reliability report checks tracked sensitive env files, sensitive env-file history, `.gitignore`, and production preflight secret-history evidence | Platform + Security | `src/services/security-reliability-controls.ts`, `src/services/deployment-preflight.ts`, `tests/services/security-reliability-controls.test.ts`, `tests/services/deployment-preflight.test.ts` |

| Production test-token exposure | Security reliability report proves test role overrides are disabled in production and production preflight fails token presence | Platform + Security | `src/services/security-reliability-controls.ts`, `src/auth/test-role-override.ts`, `tests/auth/test-role-override.test.ts`, `tests/services/deployment-preflight.test.ts` |

| Org scope boundary drift | Security reliability report requires membership-validated active-org selection plus org-scope matrix depth across scoped route families | Platform + Tenant Safety | `src/services/security-reliability-controls.ts`, `src/services/request-storage-scope.ts`, `src/services/org-scope-route-matrix.ts`, `tests/api/scoped-read-routes.test.ts` |

| Cron authorization drift | Security reliability report proves enabled outbox and publish cron drains return `cron_secret_missing` without `CRON_SECRET` | Platform + Operations | `src/services/security-reliability-controls.ts`, `src/services/cron-workers.ts`, `app/api/cron/outbox/route.ts`, `app/api/cron/publish/route.ts` |

| Complete onboarding coverage drift | Testing gap report preserves sign-in, active-org selection, import, AgentTask review, wiki review/publish, and publish-queue boundaries | Product + Platform + QA | `src/services/testing-gap-controls.ts`, `tests/services/testing-gap-controls.test.ts`, `tests/pages/protected-operator-pages.test.ts`, `tests/api/wiki-drafts/flow.test.ts` |

| Disabled feature drill drift | Testing gap report executes the staging disabled-system drill and requires AG2/projection/cron/publication rehearsal coverage | Platform + Operations | `src/services/testing-gap-controls.ts`, `src/services/staging-disabled-systems-drill.ts`, `tests/scripts/staging-disabled-systems-drill.test.ts` |

| TypeScript command confusion | Testing gap report locks release typechecking to `pnpm typecheck`/`pnpm build`; `typecheck:diagnostic` now converges and `typecheck:diagnostic:report` emits JSON/Markdown parity artifacts for future direct-tsc drift | Platform + DevEx | `src/services/testing-gap-controls.ts`, `src/services/typescript-diagnostic-debt.ts`, `scripts/typescript-diagnostic-debt.ts`, `docs/development/typescript-command-contract.md`, `package.json` |

| Storage-scope route family coverage drift | Testing gap report requires records, annotations, wiki, AI/editorial, AgentTask, and org-admin regression evidence | Platform + Tenant Safety + QA | `src/services/testing-gap-controls.ts`, `src/services/org-scope-route-matrix.ts`, `tests/api/scoped-ai-editorial-read-routes.test.ts`, `tests/api/org-tenants-routes.test.ts` |

| Local-vs-external evidence collapse | Testing gap report requires five strict external evidence contracts in the real-world lane, each with command/artifact targets, rejected local substitutes, and acceptance-ledger source links; readiness contracts are evaluated from their own source/acceptance statuses and downgrade overbroad global strict-gate claims, so local gates cannot satisfy SLO, adoption, paid-pilot, retention, or margin proof | Platform + Operators + Trust | `src/services/testing-gap-controls.ts`, `src/services/review-goals.ts`, `src/services/operator-readiness-dashboard.ts`, `tests/services/operator-readiness-dashboard.test.ts`, `tests/pages/operator-readiness-page.test.ts` |

Next review

Machine-scorecard projection update (2026-08-25): the canonical audit no longer

leaves verified Calliope and Clio cost, latency, safety, and repetition evidence

null. A shared loader replays each public receipt against exact fixtures and its

strict trial parser before projecting bounded metrics. Missing receipts receive

no credit; forged cost or fixture drift fails the audit. The resulting 60%

completeness is explicitly machine evidence only: quality lift and reviewer

effort remain null, both grades remain unproven, and no deployment/publication

authority changes.

Evidence-registry update (2026-08-25): duplicated provider receipt paths across

readiness, its CLI, and audit projection are removed. One versioned leaf registry

owns all six surface routes, including reconciliation source receipts and review/

retirement bindings. Consumer-scanning tests prevent literal reintroduction and

unique registration tests prevent ambiguous routing. Existing semantic parsers,

readiness counts, grades, and false authority remain unchanged.

Verifier-dispatch update (2026-08-25): provider-specific branches have moved out

of the readiness CLI into an exhaustive typed table. Runtime coverage rejects

missing, duplicate, unknown, or orphaned verification kinds, and only successfully

parsed provider/retirement paths enter readiness. The five existing strict parser

families and canonical 4/2/2/0 result are unchanged; this closes silent dispatch

drift without claiming model benefit.

Assignment-isolation update (2026-08-25): exact assignment/readiness replay now

runs inside the shared verifier result rather than a separate CLI loop. Missing

bindings withhold assignment credit; malformed bindings emit only surface ID and

`ASSIGNMENT_RECEIPT_INVALID`. Provider/retirement evidence remains independently

available, and no candidate text, reviewer code, private path, or parser detail is

retained. Both canonical assignments still verify; human responses remain absent.

Readiness-diagnostic update (2026-08-25): v3 adds per-surface `blockerCodes`.

Canonical rows are empty. A malformed assignment exposes only

`ASSIGNMENT_RECEIPT_INVALID` and generic prose while provider evidence remains

true and assignment readiness false. Privacy tests reject packet hashes, paths,

parser messages, candidate data, and reviewer metadata. The version bump avoids

silently changing the v2 artifact contract; authority and grades stay false/

unchanged.

  • Weekly RSI review pass
  • Expand this register when new provider/pipeline slices are introduced

Review-receipt verification update (2026-08-25): independent review completion

now requires strict semantic replay and registered-surface binding. Filename-only,

schema-drifted, arithmetically inconsistent, or authority-expanding receipts are

withheld with only `REVIEW_RECEIPT_INVALID`; provider and assignment evidence are

preserved. Reviewer identity and independence remain external obligations.

Review-chain binding update (2026-08-25): valid aggregate structure alone is no

longer sufficient. Completion requires a verified registered assignment and exact

agreement across packet hash, surface, reviewer count, and observation count.

Readiness scoring dimensions and authority are also exact-schema replayed.

Promotion-chain parity update (2026-08-25): the comparison CLI no longer parses

review receipts in isolation. It shares exact assignment/readiness chain replay,

and a wrong-packet receipt cannot produce a candidate-comparison artifact.

Blind-review navigation update (2026-08-25): long packets now expose per-pair

completion and keyboard-focusable next-incomplete navigation. The form retains no

browser draft, adds no network access or score defaults, and v2 assignment hashes

bind the exact improved forms. Local `file://` browser QA was policy-blocked;

static accessibility and full CLI workflow tests remain the verification source.