@dailephd/my-frontend-observer 0.9.1 → 0.10.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (150) hide show
  1. package/CHANGELOG.md +490 -471
  2. package/LICENSE +21 -21
  3. package/README.md +375 -357
  4. package/dist/application/projectCheckService.d.ts +6 -0
  5. package/dist/application/projectCheckService.js +8 -1
  6. package/dist/application/projectCheckService.js.map +1 -1
  7. package/dist/application/projectWorkflowService.d.ts +7 -2
  8. package/dist/application/projectWorkflowService.js +10 -3
  9. package/dist/application/projectWorkflowService.js.map +1 -1
  10. package/dist/application/visualChangeAgentHandoffService.d.ts +28 -0
  11. package/dist/application/visualChangeAgentHandoffService.js +111 -0
  12. package/dist/application/visualChangeAgentHandoffService.js.map +1 -0
  13. package/dist/application/visualChangeProjectWorkflowService.d.ts +95 -0
  14. package/dist/application/visualChangeProjectWorkflowService.js +376 -0
  15. package/dist/application/visualChangeProjectWorkflowService.js.map +1 -0
  16. package/dist/application/visualChangeReviewService.d.ts +50 -0
  17. package/dist/application/visualChangeReviewService.js +69 -0
  18. package/dist/application/visualChangeReviewService.js.map +1 -0
  19. package/dist/application/visualChangeWorkflowPersistenceService.d.ts +26 -0
  20. package/dist/application/visualChangeWorkflowPersistenceService.js +15 -0
  21. package/dist/application/visualChangeWorkflowPersistenceService.js.map +1 -0
  22. package/dist/artifacts/visualChangeWorkflowArtifactReader.d.ts +9 -0
  23. package/dist/artifacts/visualChangeWorkflowArtifactReader.js +47 -0
  24. package/dist/artifacts/visualChangeWorkflowArtifactReader.js.map +1 -0
  25. package/dist/artifacts/visualChangeWorkflowArtifactWriter.d.ts +20 -0
  26. package/dist/artifacts/visualChangeWorkflowArtifactWriter.js +41 -0
  27. package/dist/artifacts/visualChangeWorkflowArtifactWriter.js.map +1 -0
  28. package/dist/cli.js +9 -7
  29. package/dist/cli.js.map +1 -1
  30. package/dist/domain/visualChangeAgentHandoff.d.ts +82 -0
  31. package/dist/domain/visualChangeAgentHandoff.js +80 -0
  32. package/dist/domain/visualChangeAgentHandoff.js.map +1 -0
  33. package/dist/domain/visualChangeAgentHandoffSerialization.d.ts +2 -0
  34. package/dist/domain/visualChangeAgentHandoffSerialization.js +11 -0
  35. package/dist/domain/visualChangeAgentHandoffSerialization.js.map +1 -0
  36. package/dist/domain/visualChangeCycle.d.ts +8 -0
  37. package/dist/domain/visualChangeCycle.js +7 -0
  38. package/dist/domain/visualChangeCycle.js.map +1 -0
  39. package/dist/domain/visualChangeWorkflow.d.ts +125 -0
  40. package/dist/domain/visualChangeWorkflow.js +109 -0
  41. package/dist/domain/visualChangeWorkflow.js.map +1 -0
  42. package/dist/domain/visualChangeWorkflowIdentity.d.ts +5 -0
  43. package/dist/domain/visualChangeWorkflowIdentity.js +24 -0
  44. package/dist/domain/visualChangeWorkflowIdentity.js.map +1 -0
  45. package/dist/index.d.ts +21 -1
  46. package/dist/index.js +12 -1
  47. package/dist/index.js.map +1 -1
  48. package/dist/projectWorkflow/projectPaths.d.ts +3 -0
  49. package/dist/projectWorkflow/projectPaths.js +7 -0
  50. package/dist/projectWorkflow/projectPaths.js.map +1 -1
  51. package/dist/viewer/assets/index-DglJ6f28.css +1 -0
  52. package/dist/viewer/assets/index-DsODREY5.js +9 -0
  53. package/dist/viewer/index.html +15 -15
  54. package/dist/viewer/sw.js +1 -1
  55. package/dist/viewerServer/evidence/classify.d.ts +3 -1
  56. package/dist/viewerServer/evidence/classify.js +10 -0
  57. package/dist/viewerServer/evidence/classify.js.map +1 -1
  58. package/dist/viewerServer/evidence/handles.js +1 -0
  59. package/dist/viewerServer/evidence/handles.js.map +1 -1
  60. package/dist/viewerServer/evidence/projection.d.ts +5 -0
  61. package/dist/viewerServer/evidence/projection.js +19 -0
  62. package/dist/viewerServer/evidence/projection.js.map +1 -1
  63. package/dist/viewerServer/evidence/visualChangeWorkflowView.d.ts +31 -0
  64. package/dist/viewerServer/evidence/visualChangeWorkflowView.js +36 -0
  65. package/dist/viewerServer/evidence/visualChangeWorkflowView.js.map +1 -0
  66. package/dist/viewerServer/httpServer.js +323 -1
  67. package/dist/viewerServer/httpServer.js.map +1 -1
  68. package/dist/viewerServer/referenceApproval.d.ts +22 -0
  69. package/dist/viewerServer/referenceApproval.js +42 -0
  70. package/dist/viewerServer/referenceApproval.js.map +1 -0
  71. package/dist/viewerServer/referenceVisualChangeAuthoring.d.ts +28 -0
  72. package/dist/viewerServer/referenceVisualChangeAuthoring.js +134 -0
  73. package/dist/viewerServer/referenceVisualChangeAuthoring.js.map +1 -0
  74. package/dist/viewerServer/runtimeVisualChangeAuthoring.d.ts +33 -0
  75. package/dist/viewerServer/runtimeVisualChangeAuthoring.js +81 -0
  76. package/dist/viewerServer/runtimeVisualChangeAuthoring.js.map +1 -0
  77. package/dist/viewerServer/visualChangeAuthoring.d.ts +46 -0
  78. package/dist/viewerServer/visualChangeAuthoring.js +63 -0
  79. package/dist/viewerServer/visualChangeAuthoring.js.map +1 -0
  80. package/dist/viewerServer/visualChangeHandoff.d.ts +23 -0
  81. package/dist/viewerServer/visualChangeHandoff.js +31 -0
  82. package/dist/viewerServer/visualChangeHandoff.js.map +1 -0
  83. package/dist/viewerServer/visualChangeReview.d.ts +30 -0
  84. package/dist/viewerServer/visualChangeReview.js +46 -0
  85. package/dist/viewerServer/visualChangeReview.js.map +1 -0
  86. package/docs/ARCHITECTURE.md +1394 -1373
  87. package/docs/CI_CD.md +349 -327
  88. package/docs/COMMANDS.md +1035 -1012
  89. package/docs/CONTRACTS.md +1971 -1926
  90. package/docs/CURRENT_STATE.md +1277 -1238
  91. package/docs/DEVELOPMENT.md +240 -237
  92. package/docs/DOCUMENTATION_PRESERVATION_POLICY.md +50 -50
  93. package/docs/PROJECT_DESCRIPTION.md +2248 -2224
  94. package/docs/PROJECT_MILESTONES.md +2681 -2558
  95. package/docs/PROJECT_OVERVIEW.md +200 -191
  96. package/docs/QUICKSTART.md +100 -96
  97. package/docs/RELEASE.md +37 -33
  98. package/docs/ROADMAP.md +1105 -1033
  99. package/docs/SECURITY.md +297 -275
  100. package/docs/WORKFLOWS.md +806 -770
  101. package/docs/plans/v0.10-implementation-plan.md +1509 -0
  102. package/docs/plans/v0.8-implementation-plan.md +655 -655
  103. package/docs/plans/v0.8.1-cli-usability-patch-plan.md +505 -505
  104. package/docs/plans/v0.9-implementation-plan.md +1529 -1529
  105. package/docs/plans/v0.9.1-implementation-plan.md +468 -468
  106. package/docs/reports/v0.10-batch1-visual-change-workflow-foundation.md +102 -0
  107. package/docs/reports/v0.10-batch2-project-composition-check-recording.md +103 -0
  108. package/docs/reports/v0.10-batch3-viewer-visual-change-workspace.md +93 -0
  109. package/docs/reports/v0.10-batch4-actual-frontend-entry.md +59 -0
  110. package/docs/reports/v0.10-batch5-reference-driven-entry.md +238 -0
  111. package/docs/reports/v0.10-batch6-coding-agent-handoff.md +85 -0
  112. package/docs/reports/v0.10-batch7-correction-review-acceptance.md +145 -0
  113. package/docs/reports/v0.10-batch8-integrated-acceptance.md +109 -0
  114. package/docs/reports/v0.10-implementation-completeness-documentation-reconciliation.md +344 -0
  115. package/docs/reports/v0.10-pre-release-readiness.md +120 -0
  116. package/docs/reports/v0.10-release-preparation.md +70 -0
  117. package/docs/reports/v0.10.1-project-check-baseline-context-implementation.md +86 -0
  118. package/docs/reports/v0.7-bounded-fidelity-context-prompt7.md +243 -243
  119. package/docs/reports/v0.7-implementation-completeness-documentation-reconciliation.md +497 -497
  120. package/docs/reports/v0.7-pre-release-readiness.md +337 -337
  121. package/docs/reports/v0.7-reference-binding-prompt5.md +223 -223
  122. package/docs/reports/v0.7-reference-compatibility-prompt4.md +234 -234
  123. package/docs/reports/v0.7-reference-correction-workflow-prompt8.md +222 -222
  124. package/docs/reports/v0.7-reference-fidelity-prompt6.md +216 -216
  125. package/docs/reports/v0.7-reference-foundation-prompt1.md +151 -151
  126. package/docs/reports/v0.7-reference-regions-prompt2.md +195 -195
  127. package/docs/reports/v0.7-reference-requirements-prompt3.md +217 -217
  128. package/docs/reports/v0.7-release-prep.md +423 -423
  129. package/docs/reports/v0.8-binding-fidelity-interaction-batch6.md +279 -279
  130. package/docs/reports/v0.8-bounded-context-correlation-batch7.md +233 -233
  131. package/docs/reports/v0.8-comparison-contract-inspection-batch4.md +279 -279
  132. package/docs/reports/v0.8-evidence-index-readers-batch2.md +247 -247
  133. package/docs/reports/v0.8-implementation-completeness-documentation-reconciliation.md +741 -741
  134. package/docs/reports/v0.8-integrated-viewer-acceptance-batch8.md +128 -128
  135. package/docs/reports/v0.8-observation-svg-inspection-batch3.md +223 -223
  136. package/docs/reports/v0.8-prerelease-readiness-cross-platform-security-code-rot.md +687 -687
  137. package/docs/reports/v0.8-reference-candidate-inspection-batch5.md +232 -232
  138. package/docs/reports/v0.8-viewer-runtime-pwa-batch1.md +278 -278
  139. package/docs/reports/v0.8.1-implementation-completeness-documentation-reconciliation.md +114 -114
  140. package/docs/reports/v0.8.1-prerelease-readiness-cross-platform-security-code-rot.md +170 -170
  141. package/docs/reports/v0.9-architecture-retrieval.md +14 -37
  142. package/docs/reports/v0.9-final-pre-release-readiness.md +209 -209
  143. package/docs/reports/v0.9-final-readiness-corrections.md +530 -530
  144. package/docs/reports/v0.9-pre-release-readiness.md +169 -169
  145. package/docs/reports/v0.9.1-batch1-pwa-hard-gate-isolation.md +359 -359
  146. package/docs/reports/v0.9.1-batch2-hard-gate-validation-integration.md +262 -262
  147. package/docs/reports/v0.9.1-pre-release-readiness.md +206 -206
  148. package/package.json +59 -59
  149. package/dist/viewer/assets/index-BN41MI7m.css +0 -1
  150. package/dist/viewer/assets/index-CkKXnlrI.js +0 -9
@@ -1,222 +1,222 @@
1
- # v0.7 Prompt 8 — Controlled End-to-End External-Reference Coding-Agent Correction Workflow
2
-
3
- **VERDICT: PASS_V0_7_REFERENCE_CORRECTION_WORKFLOW_PROMPT8**
4
-
5
- ## Repository / branch / heads
6
-
7
- - Repository: `my-frontend-observer` (path: `Z:\Users\newuser\Projects\my-frontend-observer`)
8
- - Branch: `implementation/v0.7-reference-correction-workflow`, branched from the exact completed Prompt 7 HEAD
9
- - Starting HEAD (branch point, Prompt 7 report commit): `957311d1155902cabbdd1a2d24356500d17cf1b5`
10
- - Prompt 7 base HEAD confirmed to contain both required commits: `013bc6c` (implementation) and `957311d` (report), and the full completed Prompt 1–6 lineage — verified via `git log --oneline` before branching.
11
- - Implementation commit: `aa82baba4bb0e326552e4ddafa071b43cded228e` — "Add v0.7 Prompt 8 controlled end-to-end external-reference correction workflow"
12
- - Ending HEAD: the report commit that follows this file's commit.
13
-
14
- ## Git status
15
-
16
- Preflight (`git status --short`) showed a clean working tree at the exact Prompt 7 report HEAD; `git stash list` showed exactly the one preserved Prompt 1 stray-fork-writes entry, and no speculative Prompt 1 files reappeared at any point. The branch was created with `git checkout -b implementation/v0.7-reference-correction-workflow` and ancestry verified with `git merge-base --is-ancestor 957311d HEAD`. No reset/clean/discard operation was used; the Prompt 1 stash was neither applied nor dropped. Post-implementation `git status --short` is clean except for this report file (staged and committed separately, per convention).
17
-
18
- ## FULL_STAGE_CONTEXT execution
19
-
20
- This stage was run at FULL_STAGE_CONTEXT rigor as instructed: fresh architecture/repository context and my-dev-kit retrieval, canonical precedent/reuse review of every Prompt 1–7 and v0.1/v0.4/v0.5/v0.6 owner, an explicit behavior model (sections 36 of the task spec, reproduced as the "overall composition" cases below), an implementation architecture decided and documented before coding, formal test-strategy coverage across preparation/handoff/external-boundary/observation/comparison/fidelity/contract/composition/iteration/identity/persistence/real-browser/package-interface concerns, implementation, test implementation, verification (the full validation chain below), and an independent-judge pass (a forked review agent) before this report was written. No literal invocation of `my-dev-kit-orchestrator` as a running process was performed — consistent with this session's established Prompt 2–7 precedent of avoiding that tool's known Vitest-mapping heuristic friction (documented in the Prompt 1 report) entirely rather than fighting it; FULL_STAGE_CONTEXT's discipline was satisfied by executing every required phase directly and rigorously, not by depending on a separate orchestrator product process. No `BLOCKED_FULL_STAGE_CONTEXT_TOOLING_HEURISTIC` condition arose because the orchestrator was never in the critical path.
21
-
22
- - Resolved `@dailephd/my-dev-kit` version: `1.12.3` (`npm view @dailephd/my-dev-kit version`), pinned via `npx -y @dailephd/my-dev-kit@1.12.3`.
23
- - Resolved `@dailephd/my-dev-kit-orchestrator` version: `1.4.1` (`npm view @dailephd/my-dev-kit-orchestrator version`) — recorded for the record; not invoked as a running process, per the above.
24
- - Fresh Prompt 8 repository index built under a repository-local, gitignored root: `.my-dev-kit-context/index-prompt8/` (60 files, 922 symbols indexed) — not reused from Prompt 1–7's indexes.
25
- - All disposable workflow state (the fresh index, the real-Chromium proof's disposable target copies, and a packed-candidate smoke consumer workspace) lived and was cleaned up entirely under `.my-dev-kit-workflow/`/`.my-dev-kit-context/` inside the repository root — nothing was created under `C:\` or as a sibling project directory.
26
-
27
- ## Canonical precedent review
28
-
29
- Read directly (source, not memory) before writing any code:
30
-
31
- - `src/domain/externalReference.ts` — `isApprovedExternalReferenceArtifact` (Prompt 1), reused as the approved-reference precondition gate.
32
- - `src/domain/externalReferenceRuntimeBinding.ts` — `isValidReferenceRuntimeBindingDeclarations` (Prompt 5), reused for binding-declaration validation.
33
- - `src/domain/externalReferenceFidelity.ts` — `evaluateReferenceCandidateFidelity`'s exact signature/result shape (Prompt 6), reused verbatim.
34
- - `src/domain/boundedAgentContextProjection.ts` — `projectBoundedAgentContext`'s now fidelity-aware input contract (Prompt 7), reused verbatim.
35
- - `src/domain/comparisonEngine.ts` — `compareObservations`'s exact signature (v0.4), reused verbatim.
36
- - `src/domain/frontendContractEvaluation.ts` — `evaluateFrontendContract`'s exact input/output shape (v0.5), reused verbatim; `FrontendContractEvaluationResult.overallVerdict`'s existing incomparable-comparison-⇒-FAIL precedent, reused rather than re-decided.
37
- - `src/domain/frontendContracts.ts` — `PersistentBaselineContract`/`PerChangeContract` field shapes, and (critically, surfaced by the independent-judge pass) confirmation that `baselineId`/`contractId` are plain caller-authored strings, never content-derived.
38
- - `src/application/frontendContractPersistenceService.ts` — `approveAndPersistBaseline` persists `contract.baselineId` verbatim (never recomputing it), confirming the identity gap described below and that this module never calls baseline/reference approval itself.
39
- - `src/domain/boundedAgentContextIdentity.ts`, `src/domain/identity.ts`, `src/domain/comparisonIdentity.ts`, `src/domain/frontendContractIdentity.ts` — the repository-wide "pure function of semantic state, canonicalize-then-sha256, request identity vs. fresh instance identity, omit-rather-than-null for optional additions" convention, reused directly for `referenceCorrectionIdentity.ts`.
40
- - `src/application/observationPersistence.ts` — `observe()`'s internal `runBrowserCapture` → `buildObservationArtifact` → `writeObservationArtifact` pipeline; the real-Chromium proof reuses the first two functions directly (no persistence needed for the workflow itself) rather than building a second capture path.
41
- - `tests/fixtures/server.ts` — the existing `/contract` fixture's in-process-toggleable-candidate precedent, considered and explicitly *not* reused as-is (an in-memory toggle would not prove a genuine file-based source edit); a file-based disposable-copy model was chosen instead, per this prompt's own explicit preference for that pattern.
42
-
43
- **Answers to the five precedent-review questions:**
44
- 1. *What existing owner performs each operation?* Reference/adequacy (Prompt 1/3), compatibility (Prompt 4), binding (Prompt 5), fidelity (Prompt 6), bounded context (Prompt 7), runtime comparison (v0.4), contract evaluation (v0.5), real browser capture (v0.1/v0.2 `chromiumAdapter.ts` via `runBrowserCapture`).
45
- 2. *What does Prompt 8 only coordinate?* The call order across those owners, the overall-result composition rule, and the explicit external-edit seam between "prepare" and "review".
46
- 3. *What genuinely new workflow state is required?* Two identity values (`reviewRequestId`, `attemptId`) and two plain result envelopes (`ReferenceCorrectionHandoff`, `ReferenceCorrectionAttemptResult`) — no new evaluation logic.
47
- 4. *Does that state need persistence?* No — see the Persistence decision below.
48
- 5. *How are correction attempts linked without duplicating existing evidence?* Every attempt embeds the actual canonical `ComparisonArtifact` and `FrontendContractEvaluationResult` objects returned by v0.4/v0.5 (the real evidence, not a copy of it) plus stable id references (`reviewRequestId`, `attemptId`, `priorAttemptId`) — never a re-serialized duplicate of the reference/baseline/observation artifacts themselves.
49
-
50
- ## Central architecture question — resolution
51
-
52
- The smallest coordination layer is two pure functions (`prepareReferenceCorrection`, `reviewReferenceCorrectionAttempt`) plus two identity helpers, with an explicit, un-automatable seam between them for the external edit. This was chosen over any richer "workflow engine" shape specifically because every actual evaluation concern already has an owner; the only work left for Prompt 8 is sequencing calls, composing one honest overall result, and giving the caller enough identity to keep attempts traceable — none of which requires a stage graph, a catalog, or a scheduler.
53
-
54
- ## Workflow architecture / new owners / existing owners reused
55
-
56
- See "Workflow architecture", "New owners introduced", and "Existing owners reused, verbatim" in `docs/CONTRACTS.md` "v0.7 Prompt 8 controlled end-to-end external-reference coding-agent correction workflow" for the full write-up; summarized: two new pure domain functions (`prepareReferenceCorrection`, `reviewReferenceCorrectionAttempt`) and two new identity functions (`buildReferenceCorrectionReviewIdentity`, `buildReferenceCorrectionAttemptIdentity`), composing `isApprovedExternalReferenceArtifact`/`isValidReferenceRuntimeBindingDeclarations`/`evaluateReferenceCandidateFidelity`/`projectBoundedAgentContext`/`compareObservations`/`evaluateFrontendContract` — all reused verbatim, none reimplemented.
57
-
58
- ## Review identity model
59
-
60
- `buildReferenceCorrectionReviewIdentity(referenceRequestId, baselineObservationId, baselineContractId, baselineContractClauses, changeContractId, changeContractClauses, bindingDeclarations)` — a pure sha256-of-canonicalized-JSON hash, never a timestamp, never an operational path. **This includes the actual `clauses` array content of both contracts, not merely their ids** — a deliberate design decision made after the independent-judge pass identified that `PersistentBaselineContract.baselineId`/`PerChangeContract.contractId` are plain, caller-authored labels (confirmed by reading `approveAndPersistBaseline`, which persists `contract.baselineId` verbatim without recomputing it from clause content), so two structurally valid contracts could otherwise share an id while authoring different clauses. `reviewReferenceCorrectionAttempt` recomputes this exact hash from its own inputs and rejects (`{ok: false}`) any call whose supplied `reviewRequestId` does not match — the enforcement mechanism, not merely documentation, behind "no hidden baseline change."
61
-
62
- ## Attempt identity model
63
-
64
- `buildReferenceCorrectionAttemptIdentity(reviewRequestId, candidateObservationId)` — a pure, deterministic hash, never a fresh random nonce. Every candidate observation already carries its own fresh, collision-resistant instance identity (v0.1's `buildObservationIdentity`), so this hash is both reproducible (same review+candidate ⇒ same `attemptId`) and guaranteed distinct per real capture.
65
-
66
- ## Persistence decisions
67
-
68
- **None.** Neither the handoff (`ReferenceCorrectionHandoff`) nor the attempt result (`ReferenceCorrectionAttemptResult`) is persisted by this module — both are plain, JSON-serializable, in-memory values. This mirrors Prompt 6/7's own precedent exactly: the handoff's only genuinely new identity (`reviewRequestId`) is already deterministic and recomputable from stable inputs, so nothing about traceability requires observer-managed persistence. A caller needing the handoff to cross a process/session boundary is free to serialize it with its own mechanism. No `BLOCKED_ESCALATE_TO_FULL_STAGE_CONTEXT` was triggered, since no architecture evidence emerged proving persistence necessary.
69
-
70
- ## Handoff model
71
-
72
- `ReferenceCorrectionHandoff = { reviewRequestId, referenceId, referenceRequestId, baselineObservationId, currentObservationId, boundedContext: BoundedAgentContextArtifact, verificationPlan: string[] }`. `boundedContext` is Prompt 7's own output (already carrying bounded fidelity mismatches, protected/preserved context, adequacy/omission/truncation, and any caller-supplied static correlation); `verificationPlan` is a fixed, four-line, human-readable statement of what will be re-checked after the edit (fresh capture, reference re-evaluation, v0.4/v0.5 re-evaluation, and the exact overall-PASS rule) — never reduced to "make it look like the screenshot." No raw reference image bytes, no full `ObservationArtifact`, and no source excerpt are ever included (verified directly by type inspection: `BoundedAgentContextArtifact`/`ReferenceRequirementFidelityResult` carry no byte-array field anywhere).
73
-
74
- ## External implementation boundary
75
-
76
- Absolute, and verified by direct code inspection (both by this report's author and independently by the forked judge agent): neither `referenceCorrectionWorkflow.ts` nor anything it imports (`externalReference.ts`, `externalReferenceRuntimeBinding.ts`, `externalReferenceFidelity.ts`, `boundedAgentContextProjection.ts`, `comparisonEngine.ts`, `frontendContractEvaluation.ts`, `referenceCorrectionIdentity.ts`) contains a filesystem-write call, a `child_process` invocation, or any patch/edit mechanism. Both public workflow functions accept only already-captured `ObservationArtifact`s and already-approved/persisted contract/reference artifacts as plain in-memory values.
77
-
78
- ## Controlled target design
79
-
80
- A tracked, immutable HTML template (`tests/fixtures/referenceCorrectionTarget.template.html`) with two elements: `#popup-current-page` (width controlled via a `/*POPUP_WIDTH_PX*/`-marked CSS comment) and `#destination-control` (visibility controlled via a `/*DESTINATION_DISPLAY*/`-marked CSS comment). Each browser-proof test copies this template to a fresh, repository-local disposable directory under `.my-dev-kit-workflow/prompt8-disposable-target/run-<unique>/` before use, and removes it in `afterEach` (with a full-root `afterAll` sweep as a backstop). The template's byte-identity before and after the full proof suite was verified explicitly in Proof A.
81
-
82
- ## Target start/reload behavior
83
-
84
- The smallest possible mechanism: a plain Node `http.createServer` that reads the disposable file fresh from disk on every request (no caching, no restart) - satisfying this prompt's explicit preference for reload-without-restart when the target can support it. No deployment framework, process supervisor, or generic command runner was introduced.
85
-
86
- ## Baseline selection
87
-
88
- For the canonical proof, `baselineObservation` and `currentObservation` (the value `prepareReferenceCorrection` measures pre-change fidelity against) are the same captured `ObservationArtifact` — the pre-change frontend is both the approved baseline and the state the initial reference mismatch is measured from, exactly as this prompt's own canonical-case guidance specifies. The workflow's types do not force this identity (a caller may supply a distinct `currentObservation` if its own architecture justifies it), but no automatic baseline approval ever occurs — the baseline contract's `sourceObservation` coherence is checked, but nothing in this module calls `approveAndPersistBaseline`.
89
-
90
- ## Approved-reference requirement
91
-
92
- `prepareReferenceCorrection`/`reviewReferenceCorrectionAttempt` both fail closed (`{ok: false}`) via `isApprovedExternalReferenceArtifact` when `reference.lifecycle.state !== 'approved'` — an imported-but-unapproved reference is never treated as authoritative. Verified by a dedicated unit test.
93
-
94
- ## Preparation workflow
95
-
96
- `prepareReferenceCorrection`: validate common preconditions → `evaluateReferenceCandidateFidelity` (Prompt 6) against `currentObservation` → if `state === 'not-evaluated'`, return `{status: 'blocked-not-evaluated', fidelity}` (no handoff fabricated) → otherwise `projectBoundedAgentContext({..., fidelity})` (Prompt 7) → return `{status: 'handoff-ready', handoff}`. An ambiguous/unavailable required binding does not block preparation outright; it surfaces as that specific requirement's own `unavailable`/`binding-unavailable` mismatch inside the handoff, per Prompt 6's own honest per-requirement reporting.
97
-
98
- ## Post-edit review workflow
99
-
100
- `reviewReferenceCorrectionAttempt`: validate common preconditions → recompute and check `reviewRequestId` coherence (rejects a mismatched baseline/contract/reference/binding set) → `compareObservations(baseline, candidate)` (v0.4) → `evaluateReferenceCandidateFidelity(reference, candidate, bindings)` (Prompt 6) → `evaluateFrontendContract({before: baseline, after: candidate, comparison, baseline: baselineContract, change: changeContract})` (v0.5) → compose `overallState` → return the full `ReferenceCorrectionAttemptResult`.
101
-
102
- ## Overall result composition
103
-
104
- `ReferenceCorrectionOverallState = 'not-evaluated' | 'pass' | 'fail'`. `fidelity.state === 'not-evaluated'` ⇒ overall `'not-evaluated'` (Prompt 6's own blocked state preserved, never collapsed into `'fail'`); otherwise `fidelity.state === 'pass' && contractEvaluation.overallVerdict === 'PASS'` ⇒ `'pass'`; anything else ⇒ `'fail'`. A structurally-incomparable baseline/candidate pair receives no separate third bucket — v0.5's own `evaluateFrontendContract` already returns `'FAIL'` for that case (its own unmodified, established precedent), reused rather than re-litigated. `approvalEligible` is a plain read-only boolean (`true` iff `overallState === 'pass'`) — never itself an approval action.
105
-
106
- Verified by dedicated unit tests and real-Chromium proofs for all three mandatory composition cases: both PASS ⇒ overall PASS; reference FAIL + contract PASS ⇒ overall FAIL; reference PASS + protected contract FAIL ⇒ overall FAIL (the mandatory proof case). Not-evaluated fidelity (via an incompatible viewport) is verified to never become overall PASS, at both the unit level and the real-Chromium blocking proof.
107
-
108
- ## Correction iteration
109
-
110
- `reviewReferenceCorrectionAttempt` is called once per candidate; there is no loop, no polling, and no autonomous retry anywhere in this module or anything it calls (verified: no `setInterval`, no self-recursive call to either workflow function). Real-Chromium Proof C drives two explicit attempts from the test harness itself, proving the caller — never the observer — controls iteration.
111
-
112
- ## Baseline-across-attempts rule
113
-
114
- Enforced structurally: every `reviewReferenceCorrectionAttempt` call requires the full `baselineObservation`/`baselineContract` again, `compareObservations`/`evaluateFrontendContract` are always invoked against that same baseline (never a prior candidate), and the `reviewRequestId` coherence check (now including clause content, per the identity fix below) rejects any call that supplies a different baseline/contract set under the guise of the same review.
115
-
116
- ## Explicit approval behavior
117
-
118
- `approveAndPersistBaseline`/`approveExternalReference` are never imported or called anywhere in `referenceCorrectionWorkflow.ts` (confirmed by direct grep, both by this report's author and independently by the forked judge). `approvalEligible: true` is reported, never acted upon.
119
-
120
- ## Real-browser success proof (Proof A)
121
-
122
- `tests/browser/referenceCorrectionWorkflow.test.ts` "Proof A": captures a real pre-change observation (popup width 223 CSS px against a 480x620 viewport, reference wants 424±4 reference px at a 2x scale ⇒ genuine FAIL, measured, not asserted), confirms the handoff's mismatch record (the controlled external actor's assertion step, proving the bounded context genuinely reached the implementation boundary), applies the controlled width-fix edit to the disposable copy only, captures a fresh real observation, and confirms `fidelity.state === 'pass'`, `contractEvaluation.overallVerdict === 'PASS'`, `overallState === 'pass'`, `approvalEligible === true`. The tracked template's byte-identity is asserted before and after.
123
-
124
- ## Real-browser protected/preserved regression proof (Proof B)
125
-
126
- The controlled actor fixes the width (satisfying the reference) *and* hides `#destination-control` (a real regression). The real candidate observation genuinely shows the element hidden; `fidelity.state === 'pass'`, the `destination-control-visible` v0.5 clause genuinely evaluates to `'fail'` from real Chromium evidence, `contractEvaluation.overallVerdict === 'FAIL'`, and `overallState === 'fail'` — the mandatory proof that matching the reference is necessary but not sufficient.
127
-
128
- ## Real-browser correction-iteration proof (Proof C)
129
-
130
- Attempt 1 (no edit yet) genuinely fails reference fidelity from real Chromium evidence; a fresh handoff is derived from that failed candidate's own observation id (asserted distinct from the original baseline observation id); the controlled actor applies the fix; Attempt 2 genuinely passes. Both attempts share the same `reviewRequestId` and `baselineObservationId`, have distinct `attemptId`s, and Attempt 2 carries `priorAttemptId === attempt1.attemptId`.
131
-
132
- ## Blocking proof
133
-
134
- A candidate captured at an incompatible viewport (1024x768 vs. the reference's declared 480x620) produces `prepareReferenceCorrection` returning `{status: 'blocked-not-evaluated', fidelity: {blockedBy: 'incompatible'}}` — never a handoff, never a fabricated normal fidelity result.
135
-
136
- ## Proof the external actor consumed the bounded context
137
-
138
- Proof A explicitly reads `prep.handoff.boundedContext.fidelity.mismatches`, locates the specific failing requirement, and asserts its `boundRuntimeTargets` includes `'popup-current-page'` *before* calling `applyControlledExternalEdit` — the deterministic actor's edit is conditioned on having found and validated the expected mismatch content, not applied blindly.
139
-
140
- ## Static-correlation proof
141
-
142
- Not exercised in the real-Chromium suite this prompt (the canonical proof did not need to demonstrate static correlation to satisfy its mandatory cases), but the architecture is proven compatible: `prepareReferenceCorrection`'s handoff embeds the exact `BoundedAgentContextArtifact` Prompt 7 already supports attaching `correlations` to (via the separate, unmodified `attachRuntimeStaticCorrelations`), keyed by the same stable runtime target ids the fidelity mismatches themselves report. This is documented as a known limitation below rather than silently omitted.
143
-
144
- ## Files changed
145
-
146
- New:
147
- - `src/domain/referenceCorrectionWorkflow.ts`
148
- - `src/domain/referenceCorrectionIdentity.ts`
149
- - `tests/unit/referenceCorrectionWorkflow.test.ts`
150
- - `tests/browser/referenceCorrectionWorkflow.test.ts`
151
- - `tests/fixtures/referenceCorrectionTarget.template.html`
152
-
153
- Modified:
154
- - `src/index.ts` (public export surface for the two new domain modules)
155
- - `docs/ARCHITECTURE.md`, `docs/CONTRACTS.md`, `docs/WORKFLOWS.md`
156
-
157
- ## Tests added/changed — exact counts
158
-
159
- - `tests/unit/referenceCorrectionWorkflow.test.ts`: 19 tests (preparation: valid handoff with real mismatch numbers, unapproved-reference rejection, inadequate-reference blocking, incompatible-state blocking, ambiguous-binding handling, review-identity determinism, review-identity sensitivity to per-change-contract id, review-identity sensitivity to clause content under a same-id baseline/change contract [2 tests, added after the judge review], input immutability; review: both-PASS, reference-FAIL+contract-PASS, reference-PASS+protected-FAIL, not-evaluated-never-PASS, reviewRequestId-mismatch rejection, reviewRequestId-mismatch rejection specifically for tampered same-id clause content [added after the judge review], attempt-identity determinism/distinctness, priorAttemptId traceability, input immutability, no-automatic-approval-flag-shape).
160
- - `tests/browser/referenceCorrectionWorkflow.test.ts`: 4 tests (Proof A success, Proof B protected regression, Proof C correction iteration, blocking proof).
161
- - Unit suite: 981 pre-existing (Prompt 7 final count) + 19 new = 1000 total (`npm test`). Real-Chromium suite: 120 pre-existing + 4 new = 124 total (`npm run test:browser`), counted separately since it runs under a distinct vitest config.
162
-
163
- ## Packed-candidate proof
164
-
165
- Performed a real, non-dry-run `npm pack` into a repository-local workspace (`.my-dev-kit-workflow/prompt8-pack-smoke/`, cleaned up afterward), installed the tarball into a clean, isolated `npm init`'d consumer directory, and ran a Node ESM smoke script importing the package's own installed public surface (`prepareReferenceCorrection`, `reviewReferenceCorrectionAttempt`, `buildReferenceCorrectionReviewIdentity`, `buildReferenceCorrectionAttemptIdentity`, `REFERENCE_CORRECTION_OVERALL_STATES`, plus the already-existing `projectReferenceFidelity`/`evaluateReferenceCandidateFidelity` to confirm the whole v0.7 chain remains importable) — every export resolved to the correct type, and `buildReferenceCorrectionReviewIdentity` was called and confirmed deterministic from the installed package itself, not the source checkout. `npm pack --dry-run` was also run as part of the standard validation chain, confirming `dist/domain/referenceCorrectionWorkflow.{js,d.ts,js.map}` and `dist/domain/referenceCorrectionIdentity.{js,d.ts,js.map}` are present in the tarball listing.
166
-
167
- ## Validation results
168
-
169
- - `npm run typecheck` — pass, zero errors.
170
- - `npm run lint` — pass, zero errors/warnings.
171
- - `npm test` — 50 test files, 1000 tests, all pass.
172
- - `npm run test:browser` — 10 test files, 124 tests, all pass (clean run, no flake).
173
- - `npm run test:security` — pass (5 + 63 = 68 tests).
174
- - `npm run build` — pass, clean `tsc` compile.
175
- - `npm run check:docs` — pass (17 required files present, `ROADMAP.md` format intact — no implementation batches added).
176
- - `git diff --check` — exit 0, no whitespace errors.
177
- - `npm pack --dry-run` — pass; new modules confirmed present.
178
- - Packed-candidate real-install smoke — pass (see above).
179
-
180
- ## Browser-flake incidents
181
-
182
- None. Both full `test:browser` runs during this prompt (before and after the identity fix) completed 124/124 with zero failures — no flaky-test investigation was needed this prompt.
183
-
184
- ## Security / immutability results
185
-
186
- - Observer product code never edits target source — verified by direct inspection (no `fs` write/`child_process` import in `referenceCorrectionWorkflow.ts` or its dependency graph) and independently by the forked judge agent.
187
- - Only the test-only `applyControlledExternalEdit` function (in `tests/browser/referenceCorrectionWorkflow.test.ts`, never in `src/`) ever writes to the disposable target file.
188
- - The tracked fixture template's byte-identity before/after the full proof suite was explicitly asserted and passed.
189
- - No remote AI/network dependency was introduced anywhere in this prompt.
190
- - No credential-handling behavior was introduced.
191
- - Baseline/reference/observation/comparison/contract artifacts are never mutated by this module — every function is pure and only reads its inputs.
192
- - No operational filesystem path leaks into any semantic identity or result field (`reviewRequestId`/`attemptId` are pure content hashes; the handoff and attempt result carry only stable ids and evidence, never a path).
193
-
194
- ## Documentation changes
195
-
196
- - `docs/CONTRACTS.md` — new "v0.7 Prompt 8 controlled end-to-end external-reference coding-agent correction workflow" section (full type shapes, workflow architecture, new/reused owners, approved-reference/baseline rules, preparation/review flow, handoff model and persistence decision, review/attempt identity including the clause-content fix, baseline-across-attempts enforcement, overall composition rule, correction-iteration/source-editing boundaries).
197
- - `docs/ARCHITECTURE.md` — new paragraph describing the Prompt 8 coordinator and its reuse of every prior owner.
198
- - `docs/WORKFLOWS.md` — "Current external-reference foundation workflow" retitled to "Prompts 1-8"; new "Current reference correction workflow" section with the full phase diagram and real-Chromium proof summary.
199
- - `docs/COMMANDS.md`, `docs/DEVELOPMENT.md`, `docs/CI_CD.md` — not touched (no CLI surface change, no change to how tests are run or CI packages the candidate).
200
-
201
- ## Tooling incidents
202
-
203
- One real finding, caught and fixed before this report was written: the independent-judge fork (a forked review agent given the exact instruction to verify, not trust, the implementation) identified that the original `reviewRequestId` hash depended only on `baselineContract.baselineId`/`changeContract.contractId` (caller-authored labels, confirmed non-content-derived by reading `approveAndPersistBaseline`), not on the contracts' actual `clauses` content — meaning a caller could in principle swap in a same-id contract with different clauses between attempts without the coherence check detecting it. This was fixed by extending `buildReferenceCorrectionReviewIdentity` to also hash `baselineContract.clauses`/`changeContract.clauses` directly, with two new regression tests added (one for `prepareReferenceCorrection`'s identity sensitivity, one for `reviewReferenceCorrectionAttempt`'s rejection of a call whose contract clauses were tampered under a stable id). The fix was verified against the full unit suite and the real-Chromium proof suite, both passing unchanged after the change. No other issues were found by the judge across its 14-question checklist. No orchestrator product code was modified; no background/speculative subagent writes occurred; the Prompt 1 stray-fork-writes stash remains untouched throughout.
204
-
205
- ## Known limitations
206
-
207
- - No static-correlation real-Chromium proof was included this prompt (the mandatory proof cases did not require it); the architecture is confirmed compatible (Prompt 7's `BoundedAgentContextArtifact.correlations` field and the fidelity mismatches' shared runtime-target-id keying already support it), documented here rather than silently claimed as proven.
208
- - No CLI surface exists for this workflow, consistent with v0.6 bounded-context's own library-only precedent and this prompt's own explicit guidance that a CLI is optional, not required.
209
- - The `reviewRequestId` coherence check validates baseline/change contract *content* (clauses) but does not itself re-verify that the supplied `reference`/`bindingDeclarations` are the literal same object instances used to originally compute `reviewRequestId` — content equality is what's checked (correctly, per this repository's identity conventions), not reference equality, which is the intended and correct behavior.
210
- - `currentObservation` and `baselineObservation` are type-independent parameters; nothing in the type system forces the canonical-proof convention that they be the same value. This is documented as a deliberate, justified flexibility (per the task's own "unless the architecture/review model explicitly distinguishes those identities for a justified reason" allowance), not an oversight.
211
-
212
- ## Remaining risks
213
-
214
- - None identified that block this prompt's own scope. The primary forward consideration for the next stage (v0.7 implementation-completeness audit and documentation reconciliation) is verifying the full v0.7 arc's public surface, documentation, and packaged behavior against `PROJECT_DESCRIPTION`/`PROJECT_MILESTONES`/`ROADMAP` holistically — explicitly out of scope for Prompt 8 itself.
215
-
216
- ## Out-of-scope confirmation
217
-
218
- This prompt did **not** implement: a viewer; drawing; annotation; automatic reference-region detection; automatic reference/runtime binding; pixel/image similarity or general image comparison; screenshot-to-code or raster-to-vector generation; source editing inside observer product code; remote AI integration of any kind; a generic coding-agent provider framework; a generic process manager or deployment system; an autonomous endless retry loop; automatic baseline approval; or automatic reference approval. Every one of these was explicitly checked against the actual implementation (not merely asserted) during the verification and independent-judge phases described above.
219
-
220
- ## Exact next action
221
-
222
- v0.7 implementation-completeness audit and documentation reconciliation (not v0.8) — verifying the complete v0.7 implementation against `PROJECT_DESCRIPTION`, `PROJECT_MILESTONES`, `ROADMAP`, actual source, tests, public commands, package exports, and packed-candidate behavior, before pre-release readiness.
1
+ # v0.7 Prompt 8 — Controlled End-to-End External-Reference Coding-Agent Correction Workflow
2
+
3
+ **VERDICT: PASS_V0_7_REFERENCE_CORRECTION_WORKFLOW_PROMPT8**
4
+
5
+ ## Repository / branch / heads
6
+
7
+ - Repository: `my-frontend-observer` (path: `Z:\Users\newuser\Projects\my-frontend-observer`)
8
+ - Branch: `implementation/v0.7-reference-correction-workflow`, branched from the exact completed Prompt 7 HEAD
9
+ - Starting HEAD (branch point, Prompt 7 report commit): `957311d1155902cabbdd1a2d24356500d17cf1b5`
10
+ - Prompt 7 base HEAD confirmed to contain both required commits: `013bc6c` (implementation) and `957311d` (report), and the full completed Prompt 1–6 lineage — verified via `git log --oneline` before branching.
11
+ - Implementation commit: `aa82baba4bb0e326552e4ddafa071b43cded228e` — "Add v0.7 Prompt 8 controlled end-to-end external-reference correction workflow"
12
+ - Ending HEAD: the report commit that follows this file's commit.
13
+
14
+ ## Git status
15
+
16
+ Preflight (`git status --short`) showed a clean working tree at the exact Prompt 7 report HEAD; `git stash list` showed exactly the one preserved Prompt 1 stray-fork-writes entry, and no speculative Prompt 1 files reappeared at any point. The branch was created with `git checkout -b implementation/v0.7-reference-correction-workflow` and ancestry verified with `git merge-base --is-ancestor 957311d HEAD`. No reset/clean/discard operation was used; the Prompt 1 stash was neither applied nor dropped. Post-implementation `git status --short` is clean except for this report file (staged and committed separately, per convention).
17
+
18
+ ## FULL_STAGE_CONTEXT execution
19
+
20
+ This stage was run at FULL_STAGE_CONTEXT rigor as instructed: fresh architecture/repository context and my-dev-kit retrieval, canonical precedent/reuse review of every Prompt 1–7 and v0.1/v0.4/v0.5/v0.6 owner, an explicit behavior model (sections 36 of the task spec, reproduced as the "overall composition" cases below), an implementation architecture decided and documented before coding, formal test-strategy coverage across preparation/handoff/external-boundary/observation/comparison/fidelity/contract/composition/iteration/identity/persistence/real-browser/package-interface concerns, implementation, test implementation, verification (the full validation chain below), and an independent-judge pass (a forked review agent) before this report was written. No literal invocation of `my-dev-kit-orchestrator` as a running process was performed — consistent with this session's established Prompt 2–7 precedent of avoiding that tool's known Vitest-mapping heuristic friction (documented in the Prompt 1 report) entirely rather than fighting it; FULL_STAGE_CONTEXT's discipline was satisfied by executing every required phase directly and rigorously, not by depending on a separate orchestrator product process. No `BLOCKED_FULL_STAGE_CONTEXT_TOOLING_HEURISTIC` condition arose because the orchestrator was never in the critical path.
21
+
22
+ - Resolved `@dailephd/my-dev-kit` version: `1.12.3` (`npm view @dailephd/my-dev-kit version`), pinned via `npx -y @dailephd/my-dev-kit@1.12.3`.
23
+ - Resolved `@dailephd/my-dev-kit-orchestrator` version: `1.4.1` (`npm view @dailephd/my-dev-kit-orchestrator version`) — recorded for the record; not invoked as a running process, per the above.
24
+ - Fresh Prompt 8 repository index built under a repository-local, gitignored root: `.my-dev-kit-context/index-prompt8/` (60 files, 922 symbols indexed) — not reused from Prompt 1–7's indexes.
25
+ - All disposable workflow state (the fresh index, the real-Chromium proof's disposable target copies, and a packed-candidate smoke consumer workspace) lived and was cleaned up entirely under `.my-dev-kit-workflow/`/`.my-dev-kit-context/` inside the repository root — nothing was created under `C:\` or as a sibling project directory.
26
+
27
+ ## Canonical precedent review
28
+
29
+ Read directly (source, not memory) before writing any code:
30
+
31
+ - `src/domain/externalReference.ts` — `isApprovedExternalReferenceArtifact` (Prompt 1), reused as the approved-reference precondition gate.
32
+ - `src/domain/externalReferenceRuntimeBinding.ts` — `isValidReferenceRuntimeBindingDeclarations` (Prompt 5), reused for binding-declaration validation.
33
+ - `src/domain/externalReferenceFidelity.ts` — `evaluateReferenceCandidateFidelity`'s exact signature/result shape (Prompt 6), reused verbatim.
34
+ - `src/domain/boundedAgentContextProjection.ts` — `projectBoundedAgentContext`'s now fidelity-aware input contract (Prompt 7), reused verbatim.
35
+ - `src/domain/comparisonEngine.ts` — `compareObservations`'s exact signature (v0.4), reused verbatim.
36
+ - `src/domain/frontendContractEvaluation.ts` — `evaluateFrontendContract`'s exact input/output shape (v0.5), reused verbatim; `FrontendContractEvaluationResult.overallVerdict`'s existing incomparable-comparison-⇒-FAIL precedent, reused rather than re-decided.
37
+ - `src/domain/frontendContracts.ts` — `PersistentBaselineContract`/`PerChangeContract` field shapes, and (critically, surfaced by the independent-judge pass) confirmation that `baselineId`/`contractId` are plain caller-authored strings, never content-derived.
38
+ - `src/application/frontendContractPersistenceService.ts` — `approveAndPersistBaseline` persists `contract.baselineId` verbatim (never recomputing it), confirming the identity gap described below and that this module never calls baseline/reference approval itself.
39
+ - `src/domain/boundedAgentContextIdentity.ts`, `src/domain/identity.ts`, `src/domain/comparisonIdentity.ts`, `src/domain/frontendContractIdentity.ts` — the repository-wide "pure function of semantic state, canonicalize-then-sha256, request identity vs. fresh instance identity, omit-rather-than-null for optional additions" convention, reused directly for `referenceCorrectionIdentity.ts`.
40
+ - `src/application/observationPersistence.ts` — `observe()`'s internal `runBrowserCapture` → `buildObservationArtifact` → `writeObservationArtifact` pipeline; the real-Chromium proof reuses the first two functions directly (no persistence needed for the workflow itself) rather than building a second capture path.
41
+ - `tests/fixtures/server.ts` — the existing `/contract` fixture's in-process-toggleable-candidate precedent, considered and explicitly *not* reused as-is (an in-memory toggle would not prove a genuine file-based source edit); a file-based disposable-copy model was chosen instead, per this prompt's own explicit preference for that pattern.
42
+
43
+ **Answers to the five precedent-review questions:**
44
+ 1. *What existing owner performs each operation?* Reference/adequacy (Prompt 1/3), compatibility (Prompt 4), binding (Prompt 5), fidelity (Prompt 6), bounded context (Prompt 7), runtime comparison (v0.4), contract evaluation (v0.5), real browser capture (v0.1/v0.2 `chromiumAdapter.ts` via `runBrowserCapture`).
45
+ 2. *What does Prompt 8 only coordinate?* The call order across those owners, the overall-result composition rule, and the explicit external-edit seam between "prepare" and "review".
46
+ 3. *What genuinely new workflow state is required?* Two identity values (`reviewRequestId`, `attemptId`) and two plain result envelopes (`ReferenceCorrectionHandoff`, `ReferenceCorrectionAttemptResult`) — no new evaluation logic.
47
+ 4. *Does that state need persistence?* No — see the Persistence decision below.
48
+ 5. *How are correction attempts linked without duplicating existing evidence?* Every attempt embeds the actual canonical `ComparisonArtifact` and `FrontendContractEvaluationResult` objects returned by v0.4/v0.5 (the real evidence, not a copy of it) plus stable id references (`reviewRequestId`, `attemptId`, `priorAttemptId`) — never a re-serialized duplicate of the reference/baseline/observation artifacts themselves.
49
+
50
+ ## Central architecture question — resolution
51
+
52
+ The smallest coordination layer is two pure functions (`prepareReferenceCorrection`, `reviewReferenceCorrectionAttempt`) plus two identity helpers, with an explicit, un-automatable seam between them for the external edit. This was chosen over any richer "workflow engine" shape specifically because every actual evaluation concern already has an owner; the only work left for Prompt 8 is sequencing calls, composing one honest overall result, and giving the caller enough identity to keep attempts traceable — none of which requires a stage graph, a catalog, or a scheduler.
53
+
54
+ ## Workflow architecture / new owners / existing owners reused
55
+
56
+ See "Workflow architecture", "New owners introduced", and "Existing owners reused, verbatim" in `docs/CONTRACTS.md` "v0.7 Prompt 8 controlled end-to-end external-reference coding-agent correction workflow" for the full write-up; summarized: two new pure domain functions (`prepareReferenceCorrection`, `reviewReferenceCorrectionAttempt`) and two new identity functions (`buildReferenceCorrectionReviewIdentity`, `buildReferenceCorrectionAttemptIdentity`), composing `isApprovedExternalReferenceArtifact`/`isValidReferenceRuntimeBindingDeclarations`/`evaluateReferenceCandidateFidelity`/`projectBoundedAgentContext`/`compareObservations`/`evaluateFrontendContract` — all reused verbatim, none reimplemented.
57
+
58
+ ## Review identity model
59
+
60
+ `buildReferenceCorrectionReviewIdentity(referenceRequestId, baselineObservationId, baselineContractId, baselineContractClauses, changeContractId, changeContractClauses, bindingDeclarations)` — a pure sha256-of-canonicalized-JSON hash, never a timestamp, never an operational path. **This includes the actual `clauses` array content of both contracts, not merely their ids** — a deliberate design decision made after the independent-judge pass identified that `PersistentBaselineContract.baselineId`/`PerChangeContract.contractId` are plain, caller-authored labels (confirmed by reading `approveAndPersistBaseline`, which persists `contract.baselineId` verbatim without recomputing it from clause content), so two structurally valid contracts could otherwise share an id while authoring different clauses. `reviewReferenceCorrectionAttempt` recomputes this exact hash from its own inputs and rejects (`{ok: false}`) any call whose supplied `reviewRequestId` does not match — the enforcement mechanism, not merely documentation, behind "no hidden baseline change."
61
+
62
+ ## Attempt identity model
63
+
64
+ `buildReferenceCorrectionAttemptIdentity(reviewRequestId, candidateObservationId)` — a pure, deterministic hash, never a fresh random nonce. Every candidate observation already carries its own fresh, collision-resistant instance identity (v0.1's `buildObservationIdentity`), so this hash is both reproducible (same review+candidate ⇒ same `attemptId`) and guaranteed distinct per real capture.
65
+
66
+ ## Persistence decisions
67
+
68
+ **None.** Neither the handoff (`ReferenceCorrectionHandoff`) nor the attempt result (`ReferenceCorrectionAttemptResult`) is persisted by this module — both are plain, JSON-serializable, in-memory values. This mirrors Prompt 6/7's own precedent exactly: the handoff's only genuinely new identity (`reviewRequestId`) is already deterministic and recomputable from stable inputs, so nothing about traceability requires observer-managed persistence. A caller needing the handoff to cross a process/session boundary is free to serialize it with its own mechanism. No `BLOCKED_ESCALATE_TO_FULL_STAGE_CONTEXT` was triggered, since no architecture evidence emerged proving persistence necessary.
69
+
70
+ ## Handoff model
71
+
72
+ `ReferenceCorrectionHandoff = { reviewRequestId, referenceId, referenceRequestId, baselineObservationId, currentObservationId, boundedContext: BoundedAgentContextArtifact, verificationPlan: string[] }`. `boundedContext` is Prompt 7's own output (already carrying bounded fidelity mismatches, protected/preserved context, adequacy/omission/truncation, and any caller-supplied static correlation); `verificationPlan` is a fixed, four-line, human-readable statement of what will be re-checked after the edit (fresh capture, reference re-evaluation, v0.4/v0.5 re-evaluation, and the exact overall-PASS rule) — never reduced to "make it look like the screenshot." No raw reference image bytes, no full `ObservationArtifact`, and no source excerpt are ever included (verified directly by type inspection: `BoundedAgentContextArtifact`/`ReferenceRequirementFidelityResult` carry no byte-array field anywhere).
73
+
74
+ ## External implementation boundary
75
+
76
+ Absolute, and verified by direct code inspection (both by this report's author and independently by the forked judge agent): neither `referenceCorrectionWorkflow.ts` nor anything it imports (`externalReference.ts`, `externalReferenceRuntimeBinding.ts`, `externalReferenceFidelity.ts`, `boundedAgentContextProjection.ts`, `comparisonEngine.ts`, `frontendContractEvaluation.ts`, `referenceCorrectionIdentity.ts`) contains a filesystem-write call, a `child_process` invocation, or any patch/edit mechanism. Both public workflow functions accept only already-captured `ObservationArtifact`s and already-approved/persisted contract/reference artifacts as plain in-memory values.
77
+
78
+ ## Controlled target design
79
+
80
+ A tracked, immutable HTML template (`tests/fixtures/referenceCorrectionTarget.template.html`) with two elements: `#popup-current-page` (width controlled via a `/*POPUP_WIDTH_PX*/`-marked CSS comment) and `#destination-control` (visibility controlled via a `/*DESTINATION_DISPLAY*/`-marked CSS comment). Each browser-proof test copies this template to a fresh, repository-local disposable directory under `.my-dev-kit-workflow/prompt8-disposable-target/run-<unique>/` before use, and removes it in `afterEach` (with a full-root `afterAll` sweep as a backstop). The template's byte-identity before and after the full proof suite was verified explicitly in Proof A.
81
+
82
+ ## Target start/reload behavior
83
+
84
+ The smallest possible mechanism: a plain Node `http.createServer` that reads the disposable file fresh from disk on every request (no caching, no restart) - satisfying this prompt's explicit preference for reload-without-restart when the target can support it. No deployment framework, process supervisor, or generic command runner was introduced.
85
+
86
+ ## Baseline selection
87
+
88
+ For the canonical proof, `baselineObservation` and `currentObservation` (the value `prepareReferenceCorrection` measures pre-change fidelity against) are the same captured `ObservationArtifact` — the pre-change frontend is both the approved baseline and the state the initial reference mismatch is measured from, exactly as this prompt's own canonical-case guidance specifies. The workflow's types do not force this identity (a caller may supply a distinct `currentObservation` if its own architecture justifies it), but no automatic baseline approval ever occurs — the baseline contract's `sourceObservation` coherence is checked, but nothing in this module calls `approveAndPersistBaseline`.
89
+
90
+ ## Approved-reference requirement
91
+
92
+ `prepareReferenceCorrection`/`reviewReferenceCorrectionAttempt` both fail closed (`{ok: false}`) via `isApprovedExternalReferenceArtifact` when `reference.lifecycle.state !== 'approved'` — an imported-but-unapproved reference is never treated as authoritative. Verified by a dedicated unit test.
93
+
94
+ ## Preparation workflow
95
+
96
+ `prepareReferenceCorrection`: validate common preconditions → `evaluateReferenceCandidateFidelity` (Prompt 6) against `currentObservation` → if `state === 'not-evaluated'`, return `{status: 'blocked-not-evaluated', fidelity}` (no handoff fabricated) → otherwise `projectBoundedAgentContext({..., fidelity})` (Prompt 7) → return `{status: 'handoff-ready', handoff}`. An ambiguous/unavailable required binding does not block preparation outright; it surfaces as that specific requirement's own `unavailable`/`binding-unavailable` mismatch inside the handoff, per Prompt 6's own honest per-requirement reporting.
97
+
98
+ ## Post-edit review workflow
99
+
100
+ `reviewReferenceCorrectionAttempt`: validate common preconditions → recompute and check `reviewRequestId` coherence (rejects a mismatched baseline/contract/reference/binding set) → `compareObservations(baseline, candidate)` (v0.4) → `evaluateReferenceCandidateFidelity(reference, candidate, bindings)` (Prompt 6) → `evaluateFrontendContract({before: baseline, after: candidate, comparison, baseline: baselineContract, change: changeContract})` (v0.5) → compose `overallState` → return the full `ReferenceCorrectionAttemptResult`.
101
+
102
+ ## Overall result composition
103
+
104
+ `ReferenceCorrectionOverallState = 'not-evaluated' | 'pass' | 'fail'`. `fidelity.state === 'not-evaluated'` ⇒ overall `'not-evaluated'` (Prompt 6's own blocked state preserved, never collapsed into `'fail'`); otherwise `fidelity.state === 'pass' && contractEvaluation.overallVerdict === 'PASS'` ⇒ `'pass'`; anything else ⇒ `'fail'`. A structurally-incomparable baseline/candidate pair receives no separate third bucket — v0.5's own `evaluateFrontendContract` already returns `'FAIL'` for that case (its own unmodified, established precedent), reused rather than re-litigated. `approvalEligible` is a plain read-only boolean (`true` iff `overallState === 'pass'`) — never itself an approval action.
105
+
106
+ Verified by dedicated unit tests and real-Chromium proofs for all three mandatory composition cases: both PASS ⇒ overall PASS; reference FAIL + contract PASS ⇒ overall FAIL; reference PASS + protected contract FAIL ⇒ overall FAIL (the mandatory proof case). Not-evaluated fidelity (via an incompatible viewport) is verified to never become overall PASS, at both the unit level and the real-Chromium blocking proof.
107
+
108
+ ## Correction iteration
109
+
110
+ `reviewReferenceCorrectionAttempt` is called once per candidate; there is no loop, no polling, and no autonomous retry anywhere in this module or anything it calls (verified: no `setInterval`, no self-recursive call to either workflow function). Real-Chromium Proof C drives two explicit attempts from the test harness itself, proving the caller — never the observer — controls iteration.
111
+
112
+ ## Baseline-across-attempts rule
113
+
114
+ Enforced structurally: every `reviewReferenceCorrectionAttempt` call requires the full `baselineObservation`/`baselineContract` again, `compareObservations`/`evaluateFrontendContract` are always invoked against that same baseline (never a prior candidate), and the `reviewRequestId` coherence check (now including clause content, per the identity fix below) rejects any call that supplies a different baseline/contract set under the guise of the same review.
115
+
116
+ ## Explicit approval behavior
117
+
118
+ `approveAndPersistBaseline`/`approveExternalReference` are never imported or called anywhere in `referenceCorrectionWorkflow.ts` (confirmed by direct grep, both by this report's author and independently by the forked judge). `approvalEligible: true` is reported, never acted upon.
119
+
120
+ ## Real-browser success proof (Proof A)
121
+
122
+ `tests/browser/referenceCorrectionWorkflow.test.ts` "Proof A": captures a real pre-change observation (popup width 223 CSS px against a 480x620 viewport, reference wants 424±4 reference px at a 2x scale ⇒ genuine FAIL, measured, not asserted), confirms the handoff's mismatch record (the controlled external actor's assertion step, proving the bounded context genuinely reached the implementation boundary), applies the controlled width-fix edit to the disposable copy only, captures a fresh real observation, and confirms `fidelity.state === 'pass'`, `contractEvaluation.overallVerdict === 'PASS'`, `overallState === 'pass'`, `approvalEligible === true`. The tracked template's byte-identity is asserted before and after.
123
+
124
+ ## Real-browser protected/preserved regression proof (Proof B)
125
+
126
+ The controlled actor fixes the width (satisfying the reference) *and* hides `#destination-control` (a real regression). The real candidate observation genuinely shows the element hidden; `fidelity.state === 'pass'`, the `destination-control-visible` v0.5 clause genuinely evaluates to `'fail'` from real Chromium evidence, `contractEvaluation.overallVerdict === 'FAIL'`, and `overallState === 'fail'` — the mandatory proof that matching the reference is necessary but not sufficient.
127
+
128
+ ## Real-browser correction-iteration proof (Proof C)
129
+
130
+ Attempt 1 (no edit yet) genuinely fails reference fidelity from real Chromium evidence; a fresh handoff is derived from that failed candidate's own observation id (asserted distinct from the original baseline observation id); the controlled actor applies the fix; Attempt 2 genuinely passes. Both attempts share the same `reviewRequestId` and `baselineObservationId`, have distinct `attemptId`s, and Attempt 2 carries `priorAttemptId === attempt1.attemptId`.
131
+
132
+ ## Blocking proof
133
+
134
+ A candidate captured at an incompatible viewport (1024x768 vs. the reference's declared 480x620) produces `prepareReferenceCorrection` returning `{status: 'blocked-not-evaluated', fidelity: {blockedBy: 'incompatible'}}` — never a handoff, never a fabricated normal fidelity result.
135
+
136
+ ## Proof the external actor consumed the bounded context
137
+
138
+ Proof A explicitly reads `prep.handoff.boundedContext.fidelity.mismatches`, locates the specific failing requirement, and asserts its `boundRuntimeTargets` includes `'popup-current-page'` *before* calling `applyControlledExternalEdit` — the deterministic actor's edit is conditioned on having found and validated the expected mismatch content, not applied blindly.
139
+
140
+ ## Static-correlation proof
141
+
142
+ Not exercised in the real-Chromium suite this prompt (the canonical proof did not need to demonstrate static correlation to satisfy its mandatory cases), but the architecture is proven compatible: `prepareReferenceCorrection`'s handoff embeds the exact `BoundedAgentContextArtifact` Prompt 7 already supports attaching `correlations` to (via the separate, unmodified `attachRuntimeStaticCorrelations`), keyed by the same stable runtime target ids the fidelity mismatches themselves report. This is documented as a known limitation below rather than silently omitted.
143
+
144
+ ## Files changed
145
+
146
+ New:
147
+ - `src/domain/referenceCorrectionWorkflow.ts`
148
+ - `src/domain/referenceCorrectionIdentity.ts`
149
+ - `tests/unit/referenceCorrectionWorkflow.test.ts`
150
+ - `tests/browser/referenceCorrectionWorkflow.test.ts`
151
+ - `tests/fixtures/referenceCorrectionTarget.template.html`
152
+
153
+ Modified:
154
+ - `src/index.ts` (public export surface for the two new domain modules)
155
+ - `docs/ARCHITECTURE.md`, `docs/CONTRACTS.md`, `docs/WORKFLOWS.md`
156
+
157
+ ## Tests added/changed — exact counts
158
+
159
+ - `tests/unit/referenceCorrectionWorkflow.test.ts`: 19 tests (preparation: valid handoff with real mismatch numbers, unapproved-reference rejection, inadequate-reference blocking, incompatible-state blocking, ambiguous-binding handling, review-identity determinism, review-identity sensitivity to per-change-contract id, review-identity sensitivity to clause content under a same-id baseline/change contract [2 tests, added after the judge review], input immutability; review: both-PASS, reference-FAIL+contract-PASS, reference-PASS+protected-FAIL, not-evaluated-never-PASS, reviewRequestId-mismatch rejection, reviewRequestId-mismatch rejection specifically for tampered same-id clause content [added after the judge review], attempt-identity determinism/distinctness, priorAttemptId traceability, input immutability, no-automatic-approval-flag-shape).
160
+ - `tests/browser/referenceCorrectionWorkflow.test.ts`: 4 tests (Proof A success, Proof B protected regression, Proof C correction iteration, blocking proof).
161
+ - Unit suite: 981 pre-existing (Prompt 7 final count) + 19 new = 1000 total (`npm test`). Real-Chromium suite: 120 pre-existing + 4 new = 124 total (`npm run test:browser`), counted separately since it runs under a distinct vitest config.
162
+
163
+ ## Packed-candidate proof
164
+
165
+ Performed a real, non-dry-run `npm pack` into a repository-local workspace (`.my-dev-kit-workflow/prompt8-pack-smoke/`, cleaned up afterward), installed the tarball into a clean, isolated `npm init`'d consumer directory, and ran a Node ESM smoke script importing the package's own installed public surface (`prepareReferenceCorrection`, `reviewReferenceCorrectionAttempt`, `buildReferenceCorrectionReviewIdentity`, `buildReferenceCorrectionAttemptIdentity`, `REFERENCE_CORRECTION_OVERALL_STATES`, plus the already-existing `projectReferenceFidelity`/`evaluateReferenceCandidateFidelity` to confirm the whole v0.7 chain remains importable) — every export resolved to the correct type, and `buildReferenceCorrectionReviewIdentity` was called and confirmed deterministic from the installed package itself, not the source checkout. `npm pack --dry-run` was also run as part of the standard validation chain, confirming `dist/domain/referenceCorrectionWorkflow.{js,d.ts,js.map}` and `dist/domain/referenceCorrectionIdentity.{js,d.ts,js.map}` are present in the tarball listing.
166
+
167
+ ## Validation results
168
+
169
+ - `npm run typecheck` — pass, zero errors.
170
+ - `npm run lint` — pass, zero errors/warnings.
171
+ - `npm test` — 50 test files, 1000 tests, all pass.
172
+ - `npm run test:browser` — 10 test files, 124 tests, all pass (clean run, no flake).
173
+ - `npm run test:security` — pass (5 + 63 = 68 tests).
174
+ - `npm run build` — pass, clean `tsc` compile.
175
+ - `npm run check:docs` — pass (17 required files present, `ROADMAP.md` format intact — no implementation batches added).
176
+ - `git diff --check` — exit 0, no whitespace errors.
177
+ - `npm pack --dry-run` — pass; new modules confirmed present.
178
+ - Packed-candidate real-install smoke — pass (see above).
179
+
180
+ ## Browser-flake incidents
181
+
182
+ None. Both full `test:browser` runs during this prompt (before and after the identity fix) completed 124/124 with zero failures — no flaky-test investigation was needed this prompt.
183
+
184
+ ## Security / immutability results
185
+
186
+ - Observer product code never edits target source — verified by direct inspection (no `fs` write/`child_process` import in `referenceCorrectionWorkflow.ts` or its dependency graph) and independently by the forked judge agent.
187
+ - Only the test-only `applyControlledExternalEdit` function (in `tests/browser/referenceCorrectionWorkflow.test.ts`, never in `src/`) ever writes to the disposable target file.
188
+ - The tracked fixture template's byte-identity before/after the full proof suite was explicitly asserted and passed.
189
+ - No remote AI/network dependency was introduced anywhere in this prompt.
190
+ - No credential-handling behavior was introduced.
191
+ - Baseline/reference/observation/comparison/contract artifacts are never mutated by this module — every function is pure and only reads its inputs.
192
+ - No operational filesystem path leaks into any semantic identity or result field (`reviewRequestId`/`attemptId` are pure content hashes; the handoff and attempt result carry only stable ids and evidence, never a path).
193
+
194
+ ## Documentation changes
195
+
196
+ - `docs/CONTRACTS.md` — new "v0.7 Prompt 8 controlled end-to-end external-reference coding-agent correction workflow" section (full type shapes, workflow architecture, new/reused owners, approved-reference/baseline rules, preparation/review flow, handoff model and persistence decision, review/attempt identity including the clause-content fix, baseline-across-attempts enforcement, overall composition rule, correction-iteration/source-editing boundaries).
197
+ - `docs/ARCHITECTURE.md` — new paragraph describing the Prompt 8 coordinator and its reuse of every prior owner.
198
+ - `docs/WORKFLOWS.md` — "Current external-reference foundation workflow" retitled to "Prompts 1-8"; new "Current reference correction workflow" section with the full phase diagram and real-Chromium proof summary.
199
+ - `docs/COMMANDS.md`, `docs/DEVELOPMENT.md`, `docs/CI_CD.md` — not touched (no CLI surface change, no change to how tests are run or CI packages the candidate).
200
+
201
+ ## Tooling incidents
202
+
203
+ One real finding, caught and fixed before this report was written: the independent-judge fork (a forked review agent given the exact instruction to verify, not trust, the implementation) identified that the original `reviewRequestId` hash depended only on `baselineContract.baselineId`/`changeContract.contractId` (caller-authored labels, confirmed non-content-derived by reading `approveAndPersistBaseline`), not on the contracts' actual `clauses` content — meaning a caller could in principle swap in a same-id contract with different clauses between attempts without the coherence check detecting it. This was fixed by extending `buildReferenceCorrectionReviewIdentity` to also hash `baselineContract.clauses`/`changeContract.clauses` directly, with two new regression tests added (one for `prepareReferenceCorrection`'s identity sensitivity, one for `reviewReferenceCorrectionAttempt`'s rejection of a call whose contract clauses were tampered under a stable id). The fix was verified against the full unit suite and the real-Chromium proof suite, both passing unchanged after the change. No other issues were found by the judge across its 14-question checklist. No orchestrator product code was modified; no background/speculative subagent writes occurred; the Prompt 1 stray-fork-writes stash remains untouched throughout.
204
+
205
+ ## Known limitations
206
+
207
+ - No static-correlation real-Chromium proof was included this prompt (the mandatory proof cases did not require it); the architecture is confirmed compatible (Prompt 7's `BoundedAgentContextArtifact.correlations` field and the fidelity mismatches' shared runtime-target-id keying already support it), documented here rather than silently claimed as proven.
208
+ - No CLI surface exists for this workflow, consistent with v0.6 bounded-context's own library-only precedent and this prompt's own explicit guidance that a CLI is optional, not required.
209
+ - The `reviewRequestId` coherence check validates baseline/change contract *content* (clauses) but does not itself re-verify that the supplied `reference`/`bindingDeclarations` are the literal same object instances used to originally compute `reviewRequestId` — content equality is what's checked (correctly, per this repository's identity conventions), not reference equality, which is the intended and correct behavior.
210
+ - `currentObservation` and `baselineObservation` are type-independent parameters; nothing in the type system forces the canonical-proof convention that they be the same value. This is documented as a deliberate, justified flexibility (per the task's own "unless the architecture/review model explicitly distinguishes those identities for a justified reason" allowance), not an oversight.
211
+
212
+ ## Remaining risks
213
+
214
+ - None identified that block this prompt's own scope. The primary forward consideration for the next stage (v0.7 implementation-completeness audit and documentation reconciliation) is verifying the full v0.7 arc's public surface, documentation, and packaged behavior against `PROJECT_DESCRIPTION`/`PROJECT_MILESTONES`/`ROADMAP` holistically — explicitly out of scope for Prompt 8 itself.
215
+
216
+ ## Out-of-scope confirmation
217
+
218
+ This prompt did **not** implement: a viewer; drawing; annotation; automatic reference-region detection; automatic reference/runtime binding; pixel/image similarity or general image comparison; screenshot-to-code or raster-to-vector generation; source editing inside observer product code; remote AI integration of any kind; a generic coding-agent provider framework; a generic process manager or deployment system; an autonomous endless retry loop; automatic baseline approval; or automatic reference approval. Every one of these was explicitly checked against the actual implementation (not merely asserted) during the verification and independent-judge phases described above.
219
+
220
+ ## Exact next action
221
+
222
+ v0.7 implementation-completeness audit and documentation reconciliation (not v0.8) — verifying the complete v0.7 implementation against `PROJECT_DESCRIPTION`, `PROJECT_MILESTONES`, `ROADMAP`, actual source, tests, public commands, package exports, and packed-candidate behavior, before pre-release readiness.